Raft Consensus: How Distributed Servers Agree
Imagine you have three servers working together:
1Server 1 Server 2 Server 32 │ │ │3 └────────────┼───────────┘4 │5 Shared applicationAll three servers need to agree on the same information.
For example, they might need to agree on:
"The current leader is Server 1."
But what happens if Server 1 suddenly crashes?
Who becomes the new leader? How do the other servers agree on the change?
This is where consensus algorithms come in.
One of the most popular consensus algorithms is Raft.
What Is Raft?
Raft is a consensus algorithm that allows a group of servers to agree on the same state, even when some servers fail.
It is commonly used in distributed systems where consistency is important.
The basic idea is simple:
Multiple servers work together, but they behave as if there is one reliable source of truth.
Raft does this using three main concepts:
- Leader
- Followers
- Log replication
Let's understand them with a simple example.
Leader and Followers
Imagine we have three Raft servers:
1 Raft Cluster2 3 ┌───────────┐4 │ Server 1 │5 │ Leader │6 └─────┬─────┘7 │8 ┌─────┴─────┐9 ▼ ▼10 ┌──────────┐ ┌──────────┐11 │ Server 2 │ │ Server 3 │12 │ Follower │ │ Follower │13 └──────────┘ └──────────┘At any given time, one server acts as the leader.
The other servers are followers.
The leader is responsible for handling changes and making sure those changes are replicated to the followers.
For example, suppose an application wants to change:
1/config/port = 8080The request goes to the leader.
The leader then replicates that change to the followers.
How Is the Leader Chosen?
When a Raft cluster starts, there is initially no leader.
The servers wait for a short, randomized amount of time.
Eventually, one server's timer expires first.
That server becomes a candidate and asks the other servers to vote for it.
For example:
1Server 1 → "I want to become leader. Will you vote for me?"2 3Server 2 → "Yes"4 5Server 3 → "Yes"Server 1 now has a majority of votes, so it becomes the leader.
1Server 1 → Leader2Server 2 → Follower3Server 3 → FollowerThis process is called a leader election.
What Happens When the Leader Crashes?
This is where Raft becomes really useful.
Suppose Server 1 is the leader:
1 Server 12 Leader3 X4 CRASHED5 6Server 2 Server 37Follower FollowerThe followers stop receiving messages from the leader.
After waiting for their election timeout, one of them starts an election.
For example:
1Server 2 → "I want to become leader."2 3Server 3 → "I vote for Server 2."Server 2 receives a majority of the votes.
Now:
1Server 1 → DOWN2 3Server 2 → Leader4Server 3 → FollowerThe system has automatically recovered from the leader failure.
What Is a Majority?
Raft relies heavily on the idea of a majority, also called a quorum.
For three servers:
13 servers2Majority = 2For five servers:
15 servers2Majority = 3The cluster can continue making progress as long as a majority of servers are available.
For example, with three servers:
1Server 1 → DOWN2Server 2 → UP3Server 3 → UP4 52 out of 3 are available6→ Majority exists7→ Cluster can continueBut if two servers fail:
1Server 1 → DOWN2Server 2 → DOWN3Server 3 → UP4 51 out of 3 are available6→ No majority7→ Cluster cannot safely commit new changesThis is important because Raft would rather stop making changes than allow different parts of the cluster to disagree.
How Does Raft Replicate Data?
Now let's say the leader receives a request:
1SET /app/version = 2The leader first adds the operation to its log.
1Leader2 3Log:41. SET /app/version = 152. SET /app/version = 2It then sends the new log entry to the followers.
1 Leader2 │3 SET /app/version = 24 ┌────┴────┐5 ▼ ▼6 Server 2 Server 3The followers add the same entry to their logs.
Once a majority of servers have stored the entry, the leader can commit it.
1Server 1 → Entry stored2Server 2 → Entry stored3Server 3 → Entry stored4 5Majority reached6 ↓7 COMMITTEDThe committed change becomes part of the cluster's agreed state.
Why Use a Log?
The log is one of the most important parts of Raft.
Instead of simply saying:
1Current state = XRaft keeps a history of operations:
11. Create user22. Update configuration33. Start service44. Change leader information55. Update configurationEach server tries to maintain the same sequence of operations.
If the logs are consistent, the servers can arrive at the same state by applying those operations in the same order.
Think of it like a recipe.
If three people follow the same recipe, in the same order, they should end up with the same result.
What If a Follower Misses an Update?
Suppose the leader sends an update to Server 2, but Server 2 temporarily loses its network connection.
1Leader2 │3 ├──────────X────────── Server 24 │5 └───────────────────── Server 3Server 3 receives the update, but Server 2 doesn't.
Later, Server 2 reconnects.
The leader can send the missing log entries to Server 2.
1Server 2:2 3Before:41. A52. B6 7After synchronization:81. A92. B103. C114. DThis allows the follower to catch up with the rest of the cluster.
What Happens During a Network Failure?
Now imagine the network splits the cluster.
1 Network failure2 3 Server 1 │ Server 2 Server 34 │5 XServer 1 is separated from Servers 2 and 3.
Servers 2 and 3 still have a majority:
1Server 2 + Server 3 = 2/3They can continue operating and elect a leader.
Server 1, however, cannot safely commit new changes because it does not have a majority.
This prevents two isolated parts of the cluster from independently making decisions that could later conflict.
When the network connection comes back, the servers synchronize again.
Why Is Raft Called a Consensus Algorithm?
The word consensus simply means agreement.
Imagine three friends trying to decide where to eat.
If everyone chooses a different restaurant, there is no consensus.
But if two out of three agree on one restaurant, they have a majority decision.
Raft applies a similar idea to computers, but with strict rules for elections, logs, replication, and failures.
The goal is to make sure that the cluster agrees on a consistent sequence of operations.
Raft in the Real World
Raft isn't usually something you interact with directly.
Instead, it is used inside distributed systems.
For example, etcd uses Raft to replicate its data between nodes.
A simplified view looks like this:
1 Application2 │3 ▼4 etcd5 │6 ┌───────┴───────┐7 ▼ ▼8 Raft Leader Followers9 │ │10 └───────┬───────┘11 ▼12 Replicated stateWhen an etcd cluster needs to agree on a change, Raft handles the leader election and replication underneath.
This is one reason understanding Raft makes it much easier to understand how systems like etcd work.
Conclusion
Raft solves a very important problem in distributed systems:
How can multiple servers agree on the same state when servers and networks can fail?
It does this using a relatively simple set of ideas:
- Elect a leader.
- Let the leader coordinate changes.
- Replicate changes through a log.
- Commit changes only after reaching a majority.
- Elect a new leader when the current leader fails.
- Synchronize followers that fall behind.
Once you understand these ideas, distributed systems like etcd become much easier to visualize.
Behind what looks like a simple key-value operation, there can be multiple servers communicating, voting, replicating logs, and making sure everyone agrees on what happened.