Wise Hustlers — Digital Product & App Development Studio Logo
Get Consultation
By Wise Hustler Admin•9/25/2026•10 min read

Read Replicas and Eventual Consistency: What Breaks When You Add Them Without Thinking

Read Replicas and Eventual Consistency: What Breaks When You Add Them Without Thinking

# Read Replicas and Eventual Consistency: What Breaks When You Add Them Without Thinking

TL;DR: Database read replicas scale read throughput, not correctness — the moment you point any reads at a replica, you've traded strong consistency for eventual consistency, and the first thing that breaks is usually a user who writes something and doesn't see it on the very next page load. That's fixable with sticky routing, LSN-based read fences, or (in PostgreSQL 19, currently in beta) a native WAIT FOR LSN command — but only if you know to look for it.

Why teams reach for read replicas

The pitch is simple: writes go to a primary, reads fan out across one or more replicas, and your database stops being the bottleneck on a read-heavy app (dashboards, feeds, search-adjacent list views). Every major managed database — Amazon RDS, Aurora, Cloud SQL, Azure Database for PostgreSQL/MySQL — makes this a checkbox. It genuinely works for horizontal read scaling.

What the checkbox doesn't tell you is that replication is asynchronous by default on essentially every mainstream engine (PostgreSQL streaming replication, MySQL async/semi-sync replication, Aurora's storage-layer replication). Asynchronous replication means the primary acknowledges a commit before every replica has applied it. There is a real, measurable gap between "committed on the primary" and "visible on the replica," and application code that ignores that gap will eventually serve stale data.

How the lag actually happens

On PostgreSQL, a committed transaction on the primary moves through several stages before a replica can serve it to a query: the WAL record is written locally, the WAL sender streams it to the standby, the standby writes it to its own WAL, flushes it to disk, and only then replays it into the actual data files that queries read from (Netdata's replication lag guide). Each stage has independent latency, and under load — a burst of writes, a long-running query holding back replay, network jitter — the gap between "flushed" and "replayed" can widen from milliseconds to seconds or, in degraded cases, minutes.

AWS surfaces this directly as the ReplicaLag CloudWatch metric on RDS: for MySQL/MariaDB it mirrors Seconds_Behind_Master, and for PostgreSQL it's computed from pg_last_xact_replay_timestamp() against wall-clock time (AWS RDS replication monitoring docs). A healthy replica usually looks fine on that graph most of the time, and that "usually" is exactly what lulls teams into treating replica reads as interchangeable with primary reads.

The concrete bug: read-after-write on a replica

Here's the scenario that actually bites teams, almost verbatim from how it shows up in bug trackers:

1. A user submits a form — say, updating their billing address — via POST /api/account, which writes to the primary.

2. The API responds 200 OK.

3. The client immediately re-fetches the account via GET /api/account, which your load balancer or ORM routes to a read replica for scaling.

4. The replica hasn't replayed the write yet. The response shows the old address. The user assumes the save failed and re-submits, or worse, support gets a ticket that "the save button is broken."

This isn't hypothetical. GitLab hit a version of it in production: a change that switched a background worker's data-consistency setting from always (primary) to sticky (replica-preferred) caused the worker to read stale data from replicas that hadn't caught up, producing StuckImportJob errors that the investigating engineer estimated at around 1.2% of cases — after reverting to always, the problem "almost disappeared" (GitLab issue #364799). Same root cause as the billing-address example, just triggered by a routing policy change instead of a page reload.

A minimal reproduction with a naive round-robin read pool:

// naive setup: writes to primary, reads round-robin across replicas
async function updateAddress(userId, address) {
  await primaryPool.query(
    'UPDATE accounts SET address = $1 WHERE id = $2',
    [address, userId]
  );
}

async function getAccount(userId) {
  // BUG: this might hit a replica that hasn't replayed the UPDATE yet
  const replica = pickReplica(); // round-robin over replica pool
  const { rows } = await replica.query(
    'SELECT * FROM accounts WHERE id = $1',
    [userId]
  );
  return rows[0];
}

Nothing here is "wrong" in isolation — it's exactly the read/write split every tutorial recommends. The bug only exists in the gap between the two calls, and it's invisible in dev (single database, zero lag) and often invisible in low-traffic staging too. It shows up under real production load, which is precisely when it's hardest to debug.

Fixing it: three real strategies

1. Sticky reads to the primary after a write

The cheapest fix: track that a given session/user just wrote, and route their next N seconds of reads (or reads within the same request/session) to the primary instead of a replica.

async function getAccount(userId, session) {
  const target = session.lastWriteAt && (Date.now() - session.lastWriteAt < 5000)
    ? primaryPool
    : pickReplica();
  const { rows } = await target.query(
    'SELECT * FROM accounts WHERE id = $1',
    [userId]
  );
  return rows[0];
}

A refinement is to compare the time since the session's last write against the replica's current lag (for example the ReplicaLag metric) and use a replica only if it is less far behind than that — otherwise go to the primary. It's crude but effective, and it doesn't require touching your replication config. The weakness is that the window is a guess: if lag ever exceeds it, the bug comes back.

2. LSN fencing — wait for the replica to catch up

For workflows where a fixed timeout is too imprecise, you can capture the primary's write position and make the replica prove it has replayed past it before answering. On PostgreSQL, that's historically meant polling:

-- On the primary, right after commit:
SELECT pg_current_wal_flush_lsn();  -- e.g. '0/3000A8F8'

-- On the replica, poll until replay catches up:
SELECT pg_last_wal_replay_lsn() >= '0/3000A8F8'::pg_lsn;

You store the LSN from the write in the session (or return it to the client and pass it back), then either poll the replica or, if your driver supports it, block on it for a bounded timeout before falling back to the primary. This is essentially the same idea as read-your-writes token passing in Dynamo-style stores, applied to a relational replica.

PostgreSQL 19 (Beta 4 released September 24, 2026; not yet a stable release) turns this into a first-class primitive: a WAIT FOR LSN command that blocks a session until the given LSN is written, flushed, or replayed, depending on the mode. You capture the LSN after commit and run this on the replica connection before the read — no polling loop required (PostgreSQL 19 docs: WAIT FOR; dbi-services on PostgreSQL 19's WAIT FOR command):

-- On the replica; errors if the timeout passes first (add NO_THROW to get a status row instead)
WAIT FOR LSN '0/3000A8F8' WITH (MODE 'standby_replay', TIMEOUT '5s');

Because 19 is still in beta, check the final release notes before relying on the exact syntax.

3. Quorum reads (if you're not on Postgres/MySQL-style primary-replica)

In quorum-based systems, you avoid staleness by requiring reads and writes to overlap: if R + W > N (read replicas queried + write replicas acknowledged, out of N total), every read set intersects every write set, guaranteeing at least one up-to-date copy is seen. This is the Dynamo/Cassandra model rather than the Postgres/MySQL primary-replica model, but it's worth knowing if your stack includes both.

StrategyConsistency guaranteeCostBest for
Sticky-to-primaryHeuristic (time-window based)Low — app-level routing logicMost CRUD apps, quick fix
LSN fencing / WAIT FOR LSNExact, bounded by timeoutMedium — needs LSN plumbingCheckout flows, anything audit-sensitive
Quorum reads (R+W>N)Exact, per-requestHigh — architectural choiceDynamo/Cassandra-style stores

Beyond read-after-write: other things that break

  • Monotonic reads — a user can see a write, then on a subsequent request get routed to a more-lagged replica and see an older state than before. Sticky routing by session (not just by recent-write window) avoids this.
  • Cascading/chained replicas — a replica-of-a-replica setup compounds lag additively. If your monitoring only checks lag from the primary to the first tier, you can be blind to seconds of extra staleness further downstream.
  • Failover promoting a lagging replica — if a replica that hasn't fully caught up gets promoted during a primary failure, you can lose the most recent committed writes entirely, not just serve them late. This is a separate failure mode from read staleness and needs its own runbook (synchronous replication or remote_apply for the writes you can't afford to lose).
  • Read replicas silently falling further behind under sustained write bursts — lag isn't a constant; it's a function of current write volume and replica hardware. A replica sized for normal traffic can fall minutes behind during a bulk import or migration, right when your app is under the most scrutiny.

Getting the routing layer, lag monitoring, and failover behavior right across all of this is exactly the kind of infrastructure work that's easy to under-scope — if you're setting up read replicas as part of a broader cloud architecture, our cloud/DevOps team helps design that routing and monitoring layer correctly the first time.

FAQ

Does adding a read replica always mean eventual consistency?

Yes, by default, on every mainstream relational engine — PostgreSQL streaming replication and standard MySQL replication are asynchronous unless you explicitly configure otherwise, trading some write latency for a stronger guarantee. On PostgreSQL, synchronous_commit = remote_apply (with synchronous_standby_names set) makes a commit wait until the synchronous standbys have applied it, so it becomes visible to queries there (PostgreSQL docs). MySQL semi-synchronous replication is weaker: it waits until a replica has received and logged the transaction, not applied it, so it protects against data loss on failover but doesn't by itself fix read-after-write staleness (MySQL docs).

How much replication lag is "normal"?

There's no universal number — measure your own baseline with ReplicaLag on RDS or pg_stat_replication on self-managed PostgreSQL. Lag is workload-dependent — bulk writes, long transactions, or an underpowered replica instance can push it to seconds or minutes, so treat it as a metric to alert on, not a constant to assume.

Can I just always read from the primary to avoid this?

You can, and for low-traffic apps that's often the right call — it's one less failure mode to reason about. Read replicas only pay off once primary read load is actually your bottleneck; adding them earlier than that just adds eventual-consistency bugs for no scaling benefit.

Is eventual consistency the same thing as "the data is wrong"?

No — the data on the replica is correct, just temporarily behind. The bug isn't the lag itself, it's application code that assumes zero lag (e.g., reading your own write on the next request) without accounting for it.

Sources

Related articles