etcd: The Small Database That Keeps Distributed Systems in Sync
Imagine you have a system running on three different servers.
One server knows that a service is running. Another server needs to know the same thing. A third server also needs that information. Now imagine one of the servers suddenly crashes.
How do all the servers agree on what is true?
This is the kind of problem etcd is designed to solve.
etcd is a distributed key-value store designed to hold important data used for coordination in distributed systems. It is strongly consistent and uses the Raft consensus algorithm to keep multiple servers in agreement.
What Is etcd?
At its simplest, etcd stores data as key-value pairs.
For example:
1/service/api/port → 80802/service/api/status → running3/database/primary → server-2You can think of it like a small, highly reliable dictionary shared between multiple machines.
But etcd is different from a normal key-value database.
Its main purpose is not to store millions of user records or large files. Instead, it stores small but important pieces of information that distributed systems need to agree on. The etcd documentation describes it as a coordination service designed for relatively small amounts of data.
Why Do We Need etcd?
Consider a simple application with three servers:
1 ┌─────────────┐2 │ Application │3 └──────┬──────┘4 │5 ┌─────────┼─────────┐6 ▼ ▼ ▼7 Server 1 Server 2 Server 3Suppose Server 1 is the leader of your application.
The other servers need to know: "Who is the current leader?"
You could store this information in a normal database.
But now another problem appears.
What happens if the database itself goes down?
You have simply moved the problem somewhere else.
etcd solves this by running as a cluster of multiple nodes.
1 etcd cluster2 ┌─────────────┐3 │ Leader 1 │4 └─────┬───────┘5 │6 ┌────────────────┴──────────────┐7 ▼ ▼8 ┌──────────────────┐ ┌─────────────────┐9 │ Node 2 Follower │ │ Node 3 Follower │10 └──────────────────┘ └─────────────────┘The nodes communicate with each other and maintain the same state.
If one node fails, the remaining nodes can continue operating as long as they still have a majority, or quorum.
The Important Part: Raft
This is where etcd becomes interesting.
etcd uses an algorithm called Raft to keep its nodes synchronized. Raft provides leader election and replicated logs so that multiple machines can maintain the same state.
Imagine three etcd nodes. When a change happens, the leader receives the request and replicates that change to the other nodes.
For example:
1Client2 │3 │ PUT /database/primary = server-24 ▼5Node 1 Leader6 │7 ├──────────► Node 28 │ Follower9 │10 └──────────► Node 311 FollowerOnce enough members have accepted the change, it can be committed.
This means the cluster doesn't simply trust one machine. Multiple machines participate in maintaining the state.
What Happens When the Leader Dies?
This is one of the most useful parts of etcd.
Suppose we start with:
1Node 1 → Leader2Node 2 → Follower3Node 3 → FollowerNow Node 1 suddenly crashes.
The remaining nodes notice that the leader is no longer responding.
1Node 1 → DOWN2Node 2 → ?3Node 3 → ?Node 2 and Node 3 can hold an election.
One of them becomes the new leader:
1Node 1 → DOWN2Node 2 → Leader3Node 3 → FollowerThe application does not need to manually choose the new leader.
This is one of the reasons etcd is useful for distributed systems.
What Can You Store in etcd?
You can store simple values such as:
1/config/database/host → db.example.com2/config/database/port → 54323/service/api/instance-1 → 10.0.0.104/service/api/instance-2 → 10.0.0.115/leader → instance-2The keys can be organized using paths, making the data feel somewhat like a directory structure.
For example:
1/config2 ├── database3 │4 ├── host5 │6 └── port7 │8 └── application9 ├── timeout10 └── workersThis makes etcd useful for storing configuration and service-related information.
Watching for Changes
One particularly useful feature is watching.
Imagine an application wants to know whenever its configuration changes.
Instead of constantly asking:
1"Did the configuration change?"2"Did the configuration change?"3"Did the configuration change?"the application can watch a key.
For example:
1/config/database/hostIf the value changes:
1db-old.example.com2 ↓3db-new.example.cometcd can notify the application that something changed.
This allows applications to react to configuration or state changes without continuously polling the database.
etcd and Kubernetes
One of the most well-known users of etcd is Kubernetes.
A Kubernetes cluster has a huge amount of information that needs to be coordinated.
For example:
- Which pods exist?
- Which nodes are available?
- What deployments should be running?
- What configuration has been applied?
- What is the current state of the cluster?
Kubernetes uses etcd as the backing store for its cluster state.
You can think of it roughly like this:
1 Kubernetes2 │3 ▼4 ┌─────────┐5 │ etcd │6 └────┬────┘7 │8 ┌──────────┼──────────┐9 ▼ ▼ ▼10 Nodes Pods ServicesThe important idea is that etcd provides Kubernetes with a consistent source of truth for cluster state.
How Does an Application Talk to etcd?
etcd provides a gRPC-based API, and the project also provides etcdctl, a command-line client. The standard ports are 2379 for client communication and 2380 for communication between etcd members.
For example, you can store a value:
1etcdctl put mykey "hello"And retrieve it:
1etcdctl get mykeyThe result is simply:
1mykeyhelloIt looks simple because the basic operation really is simple.
The complexity is hidden underneath, in the distributed consensus and replication system.
Is etcd a Normal Database?
Not really.
You can think of etcd as a database because it stores persistent data, but it is designed for a different job.
A traditional application database might contain:
1Users2Orders3Products4Payments5Messagesetcd is better suited for things like:
1Who is the leader?2What configuration should I use?3Which service instances are available?4What is the current cluster state?It is designed for coordination rather than large-scale application data storage.
Why Not Just Use Redis?
Redis is excellent for caching, fast data access, queues, and many other workloads.
But etcd has a different focus.
The key difference is strong consistency and distributed consensus.
When several machines need to agree on an important piece of state, etcd's Raft-based design makes it particularly useful.
This doesn't mean etcd is a replacement for Redis or PostgreSQL. They solve different problems.
Conclusion
etcd may look like a simple key-value store from the outside, but underneath it is solving one of the hardest problems in distributed systems:
How can multiple machines agree on the same information even when some machines fail?
By combining a simple key-value interface with Raft consensus, replication, leader election, and watch functionality, etcd provides a reliable coordination layer for distributed applications.
And that's why, despite its simple API, etcd is such an important piece of infrastructure behind systems like Kubernetes.