Distributed Locks in Redis with Go
Learn to build a Redis distributed lock in Go that stops two requests from spending the same balance, step by step, and see where a lock alone is not enough.
Most backend services run as several identical copies behind a load balancer, so they can take more traffic and survive a crashed server. Each request reads some data, decides what to do and writes the result back. Inside one process, you'd guard that read-decide-write with a mutex. Across servers there's no shared memory to put a mutex in, so the guard has to live somewhere every copy can reach, such as a database row or a Redis distributed lock.
The gap shows up the first time two requests touch the same record at once. A user taps "Pay" twice on a slow connection, or a mobile client retries a payment the moment the first request times out. Two requests for the same wallet land on two different servers within the same millisecond. Both read a balance of 100, both decide 10 is affordable, and both write back 90. The user spent 20, you recorded 10, and nothing in the logs looks wrong.
The standard answer is a distributed lock: one shared place every server can reach that gives one request at a time exclusive access to a record. A database row lock, etcd or ZooKeeper can play that role; Redis is a popular choice because it's fast and often already in the stack. The naive version, one SetNX ("set if not exists": create the key only if nobody holds it) before the work and one Del (delete the key) after, breaks quietly: a crash leaves a lock that never expires, or a slow request deletes someone else's.
A lock kept in Redis is a lease, a key that expires after a set time whether or not you've finished. Take it with SET NX PX (set the key only if it doesn't exist, with an expiry in milliseconds) and a random token, release it only if the token matches, and keep the work shorter than the TTL (time to live, how long the key survives before Redis deletes it). When money is involved, also make the database refuse a write computed from a stale read; no lease can prove you still hold it. If you own the roadmap, the correctness work is a version column and a conditional UPDATE in the database you already run.
In this tutorial, you'll build a distributed lock on Redis and implement it in Go. You'll reproduce the race, fix it by hand, break the fix with a short TTL, switch to the bsm/redislock library, add fencing tokens that PostgreSQL checks, and load-test the result. Everything ran against Redis 8.10.2 and PostgreSQL 18, with go-redis v9.23.0 and bsm/redislock v0.10.0.
Prerequisites#
To follow along, you need:
- Go 1.26 or later, which go-redis v9.23.0 requires. With the default
GOTOOLCHAIN=auto,go getdownloads a newer toolchain if yours is older. - Docker, to run Redis and PostgreSQL locally.
- Working knowledge of goroutines,
context.Contextand SQL. No prior Redis experience is needed. - Ports 6379 and 5432 free, or adjust the
-pmappings and addresses to match.
Step 1: Reproduce the Race Without a Lock#
This step reproduces the double payment on your machine, so every later fix has a baseline. One wallet balance lives in Redis, and 50 goroutines, each a "Pay" request on a different server, try to spend 10 from it at once. Predict how many will believe they can afford it.
Start Redis in a container and check that it answers:
docker run -d --name redis -p 6379:6379 redis:8.10.2
docker exec redis redis-cli ping
PONG
Create a module and add go-redis, bsm/redislock and pgx, the PostgreSQL driver for Step 5. The second output line only appears if your local Go is older than 1.26:
mkdir walletlock && cd walletlock
go mod init example.com/walletlock
go get github.com/redis/go-redis/v9@v9.23.0 github.com/bsm/redislock@v0.10.0 github.com/jackc/pgx/v5@v5.11.0
go: creating new go.mod: module example.com/walletlock
go: github.com/redis/go-redis/v9@v9.23.0 requires go >= 1.26.0; switching to go1.26.9
go: upgraded go 1.24.4 => 1.26.0
go: added github.com/bsm/redislock v0.10.0
go: added github.com/jackc/pgx/v5 v5.11.0
go: added github.com/redis/go-redis/v9 v9.23.0
Spend is a read-modify-write: it reads the balance, checks it and writes the new value back, so its decision rests on a value that can change before the write lands. The sleep stands in for the real work in between.
// Package wallet holds a demo balance in Redis and spends from it the naive way.
package wallet
import (
"context"
"fmt"
"os"
"sync"
"time"
"github.com/redis/go-redis/v9"
)
const BalanceKey = "wallet:42:balance"
// Connect returns a client for REDIS_ADDR, or localhost:6379 by default.
func Connect() *redis.Client {
addr := os.Getenv("REDIS_ADDR")
if addr == "" {
addr = "localhost:6379"
}
return redis.NewClient(&redis.Options{
Addr: addr,
// Let context deadlines cut socket reads short. Without this,
// go-redis waits out its own 5-second ReadTimeout instead.
ContextTimeoutEnabled: true,
})
}
// Spend is a read-modify-write: read the balance, check it, write it back.
// The sleep stands in for real work between the read and the write
// (validation, a fraud check, a call to another service).
func Spend(ctx context.Context, rdb *redis.Client, amount int64, work time.Duration) (bool, error) {
bal, err := rdb.Get(ctx, BalanceKey).Int64()
if err != nil {
return false, err
}
if bal < amount {
return false, nil
}
time.Sleep(work)
return true, rdb.Set(ctx, BalanceKey, bal-amount, 0).Err()
}
// Hammer starts n goroutines at the same moment and counts how many
// calls to fn reported success.
func Hammer(n int, fn func() (bool, error)) (accepted int) {
var wg sync.WaitGroup
var mu sync.Mutex
start := make(chan struct{})
for range n {
wg.Add(1)
go func() {
defer wg.Done()
<-start
ok, err := fn()
if err != nil {
fmt.Println("error:", err)
return
}
if ok {
mu.Lock()
accepted++
mu.Unlock()
}
}()
}
close(start)
wg.Wait()
return accepted
}
// Check compares the final balance with what the accepted spends imply.
func Check(initial, amount int64, accepted int, balance int64) {
want := initial - amount*int64(accepted)
status := "OK"
if balance != want || balance < 0 {
status = "BROKEN"
}
fmt.Printf("accepted=%d balance=%d expected=%d %s\n", accepted, balance, want, status)
}
Hammer releases every goroutine at once by closing a channel they all wait on. Check is the yardstick for every later step: 10 spends of 10 from 100 must leave 0.
The first program spends with no lock at all:
package main
import (
"context"
"log"
"time"
"example.com/walletlock/wallet"
)
func main() {
ctx := context.Background()
rdb := wallet.Connect()
defer rdb.Close()
if err := rdb.Set(ctx, wallet.BalanceKey, 100, 0).Err(); err != nil {
log.Fatal(err)
}
// 50 requests try to spend 10 each from a balance of 100.
accepted := wallet.Hammer(50, func() (bool, error) {
return wallet.Spend(ctx, rdb, 10, 2*time.Millisecond)
})
balance, _ := rdb.Get(ctx, wallet.BalanceKey).Int64()
wallet.Check(100, 10, accepted, balance)
}
go run ./cmd/step1-race
accepted=50 balance=80 expected=-400 BROKEN
accepted=50 balance=70 expected=-400 BROKEN
accepted=50 balance=60 expected=-400 BROKEN
Three of five runs are shown. All 50 requests were accepted, so 500 was spent from 100, yet the balance only dropped to 60 to 80: the ledger shows 50 approved payments and a wallet charged for two to four. Each goroutine read 100, decided it could pay and wrote its own new balance over the others. That failure is a lost update, and the rest of this tutorial fixes it.
Go's race detector doesn't catch it:
go run -race ./cmd/step1-race
accepted=50 balance=70 expected=-400 BROKEN
There's no WARNING: DATA RACE and the exit code is 0. The race detector watches memory inside one Go process, and this shared state lives in Redis. The reliable test is the one Check does: compare the final balance with what the accepted operations imply.
Step 2: Write a Correct Lock by Hand#
This step builds the lock: a way for one request to announce "I'm working on wallet 42, wait your turn" where every server can see it. Picture a sign on a door: whoever hangs "busy" first gets the room. In Redis the sign is a key, and it needs two properties: it disappears by itself if the holder crashes, and only the holder may take it down.
Open a Redis shell, take the lock as caller A, try it as caller B, and check A's remaining time:
docker exec -it redis redis-cli
127.0.0.1:6379> SET lock:wallet:42 token-a NX PX 30000
OK
127.0.0.1:6379> SET lock:wallet:42 token-b NX PX 30000
(nil)
127.0.0.1:6379> PTTL lock:wallet:42
(integer) 29884
NX sets the key only if it doesn't exist, so B gets (nil). PX 30000 attaches a 30-second expiry in the same command, which makes the lock the lease from the introduction: if the holder crashes, Redis removes the key. The value is a token unique to this caller, as on the official Redis distributed locks page.
Both must go in one command. The common SETNX then EXPIRE is two, and a crash between them leaves a lock that never expires. The SETNX page still documents an older SETNX-based lock only because existing code links to it, and points you to SET with a Lua release instead; the SET page even discourages the single-command lock in favor of Redlock, covered in the safety section.
The token matters on release. If A finishes late, its lease may have passed to B, and deleting the key by name would remove B's lock:
Using just DEL is not safe as a client may remove another client's lock.
Redis runs a Lua script atomically, so a script can compare the token and delete the key with nothing in between. Try a release with the wrong token, then the right one:
127.0.0.1:6379> EVAL 'if redis.call("GET", KEYS[1]) == ARGV[1] then return redis.call("DEL", KEYS[1]) end return 0' 1 lock:wallet:42 token-b
(integer) 0
127.0.0.1:6379> EVAL 'if redis.call("GET", KEYS[1]) == ARGV[1] then return redis.call("DEL", KEYS[1]) end return 0' 1 lock:wallet:42 token-a
(integer) 1
B's release returns 0 and leaves A's lock alone; A's returns 1. The key goes in KEYS, as the scripting docs require for every key a script touches. Exit the shell with exit.
In Go, Acquire hangs the sign, retrying until it succeeds or the context ends, and Release is the token-checked compare-and-delete:
// Package lock is a single-node Redis lease: SET NX PX plus a random token.
package lock
import (
"context"
"crypto/rand"
"errors"
"fmt"
mrand "math/rand/v2"
"time"
"github.com/redis/go-redis/v9"
)
var (
ErrNotHeld = errors.New("lock not held (expired or taken by someone else)")
ErrBadTTL = errors.New("lock TTL must be at least 1ms")
)
// Delete the key only if it still holds our token.
var releaseScript = redis.NewScript(`
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("DEL", KEYS[1])
end
return 0`)
type Lock struct {
rdb *redis.Client
Key string
Token string
Fence int64 // set by AcquireFenced, zero otherwise
}
// Acquire retries SET key token NX PX ttl until it wins or ctx ends.
func Acquire(ctx context.Context, rdb *redis.Client, key string, ttl time.Duration) (*Lock, error) {
// A zero or negative TTL would create a lock that never expires.
if ttl < time.Millisecond {
return nil, ErrBadTTL
}
token := rand.Text() // 26 random base32 characters
for {
ok, err := rdb.SetNX(ctx, key, token, ttl).Result()
if err != nil {
return nil, err
}
if ok {
return &Lock{rdb: rdb, Key: key, Token: token}, nil
}
// Busy: wait 5 to 15 ms, randomized so waiters don't retry in lockstep.
wait := 5*time.Millisecond + mrand.N(10*time.Millisecond)
select {
case <-ctx.Done():
return nil, fmt.Errorf("acquire %s: %w", key, ctx.Err())
case <-time.After(wait):
}
}
}
// Release deletes the lock only if we still own it.
func (l *Lock) Release(ctx context.Context) error {
n, err := releaseScript.Run(ctx, l.rdb, []string{l.Key}, l.Token).Int()
if err != nil {
return err
}
if n == 0 {
return ErrNotHeld
}
return nil
}
rand.Text (Go 1.24 and later) returns at least 128 bits of randomness. releaseScript.Run falls back from EVALSHA (run a script Redis has cached, by its hash) to EVAL (send the whole script) when a restart or failover has emptied the volatile script cache. While the lock is busy, Acquire waits a random 5 to 15 ms; without that randomness, waiters retry in lockstep, and Step 4 measures the cost.
A zero TTL means no expiry
In go-redis, SetNX with a TTL of 0 or redis.KeepTTL sends no expiry, so a new lock never expires. An unset config value can do this silently, which is why Acquire rejects anything under 1 ms.
Wrap the spend in the lock and run it:
package main
import (
"context"
"log"
"time"
"example.com/walletlock/lock"
"example.com/walletlock/wallet"
)
const lockKey = "lock:wallet:42"
func main() {
ctx := context.Background()
rdb := wallet.Connect()
defer rdb.Close()
if err := rdb.Set(ctx, wallet.BalanceKey, 100, 0).Err(); err != nil {
log.Fatal(err)
}
accepted := wallet.Hammer(50, func() (bool, error) {
// Give up waiting for the lock after 5 seconds.
actx, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
l, err := lock.Acquire(actx, rdb, lockKey, 2*time.Second)
if err != nil {
return false, err
}
defer func() {
if err := l.Release(ctx); err != nil {
log.Println("release:", err)
}
}()
return wallet.Spend(ctx, rdb, 10, 2*time.Millisecond)
})
balance, _ := rdb.Get(ctx, wallet.BalanceKey).Int64()
wallet.Check(100, 10, accepted, balance)
}
go run ./cmd/step2-lock
accepted=10 balance=0 expected=0 OK
accepted=10 balance=0 expected=0 OK
accepted=10 balance=0 expected=0 OK
Exactly 10 spends were accepted and the balance was 0 on all five runs (three shown). The lock turned the burst into a queue, so the double tap now ends as the user expects: the second request waits, reads 90 and pays from that. Confirm no lock was left behind:
docker exec -it redis redis-cli EXISTS lock:wallet:42
(integer) 0
On Redis 8.4 and later, DELEX does the compare-and-delete natively and the Redis docs recommend it there; go-redis v9.23.0 exposes it as DelExArgs, marked experimental:
n, err := releaseScript.Run(ctx, l.rdb, []string{l.Key}, l.Token).Int()
Works on any Redis that allows scripting, so it's the portable default.
// releaseDelEx is the Redis 8.4+ alternative to the Lua release script.
// go-redis marks DelExArgs as experimental, so its signature may change.
func releaseDelEx(ctx context.Context, rdb *redis.Client, l *lock.Lock) error {
n, err := rdb.DelExArgs(ctx, l.Key, redis.DelExArgs{Mode: "IFEQ", MatchValue: l.Token}).Result()
if err != nil {
return err
}
if n == 0 {
return lock.ErrNotHeld
}
return nil
}
Calling it twice on the same lock prints:
first release: <nil>
second release: lock not held (expired or taken by someone else)
Against Redis 7.4, the same code fails with ERR unknown command 'delex'. Redis Cloud and Redis Software support it; Valkey 9.0 and later uses its own DELIFEQ key value, so check your provider before switching.
Step 3: Handle Work That Outlives the TTL#
This step breaks the lock by letting the work outlast the lease, then fixes it. A slow query or a long garbage-collection pause can stretch work past the TTL, and the lease expires regardless: the sign falls off while you're still inside. The fix is a watchdog, a background goroutine that keeps refreshing the lease while the work runs.
Add a token-checked refresh to lock/lock.go, plus a Watch that refreshes every third of the TTL and cancels a context the moment a refresh fails:
// Extend the TTL only if the key still holds our token.
var refreshScript = redis.NewScript(`
if redis.call("GET", KEYS[1]) == ARGV[1] then
return redis.call("PEXPIRE", KEYS[1], ARGV[2])
end
return 0`)
// Refresh resets the TTL only if we still own the lock.
func (l *Lock) Refresh(ctx context.Context, ttl time.Duration) error {
n, err := refreshScript.Run(ctx, l.rdb, []string{l.Key}, l.Token, ttl.Milliseconds()).Int()
if err != nil {
return err
}
if n == 0 {
return ErrNotHeld
}
return nil
}
// Watch refreshes the lock every ttl/3 and cancels the returned context
// as soon as a refresh fails. Do the protected work with that context.
func (l *Lock) Watch(parent context.Context, ttl time.Duration) (context.Context, context.CancelFunc) {
ctx, cancel := context.WithCancelCause(parent)
go func() {
t := time.NewTicker(ttl / 3)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
// Bound each refresh: a hung call must not outlast the lease.
rctx, rcancel := context.WithTimeout(ctx, ttl/3)
err := l.Refresh(rctx, ttl)
rcancel()
if err != nil {
cancel(fmt.Errorf("lock lost: %w", err))
return
}
}
}
}()
return ctx, func() { cancel(context.Canceled) }
}
Each refresh gets its own TTL/3 deadline, or a refresh against an unresponsive Redis waits out go-redis's default 5-second ReadTimeout. That deadline only works because Connect sets ContextTimeoutEnabled: when I paused the Redis container, a call with a 100 ms deadline took 5 s to fail on a default client and 100 ms with the option on.
Pass the returned context into the protected work so the next call fails once the watchdog sees the lock is gone. It checks only every TTL/3, which narrows the window without closing it.
This program runs 5 callers whose 300 ms of work is three times their 100 ms TTL. Without the watchdog, expect them to overlap:
package main
import (
"context"
"errors"
"flag"
"log"
"sync/atomic"
"time"
"example.com/walletlock/lock"
"example.com/walletlock/wallet"
)
const lockKey = "lock:wallet:42"
func main() {
ttl := flag.Duration("ttl", 100*time.Millisecond, "lock TTL")
work := flag.Duration("work", 300*time.Millisecond, "time spent while holding the lock")
n := flag.Int("n", 5, "concurrent callers")
watchdog := flag.Bool("watchdog", false, "refresh the lock every ttl/3")
flag.Parse()
ctx := context.Background()
rdb := wallet.Connect()
defer rdb.Close()
if err := rdb.Set(ctx, wallet.BalanceKey, 100, 0).Err(); err != nil {
log.Fatal(err)
}
var staleReleases atomic.Int64
start := time.Now()
accepted := wallet.Hammer(*n, func() (bool, error) {
actx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
l, err := lock.Acquire(actx, rdb, lockKey, *ttl)
if err != nil {
return false, err
}
defer func() {
if err := l.Release(ctx); errors.Is(err, lock.ErrNotHeld) {
staleReleases.Add(1)
}
}()
workCtx := ctx
if *watchdog {
var stop context.CancelFunc
workCtx, stop = l.Watch(ctx, *ttl)
defer stop()
}
ok, err := wallet.Spend(workCtx, rdb, 10, *work)
if err == nil && workCtx.Err() != nil {
log.Println(context.Cause(workCtx))
}
return ok, err
})
balance, _ := rdb.Get(ctx, wallet.BalanceKey).Int64()
wallet.Check(100, 10, accepted, balance)
log.Printf("refused stale releases=%d elapsed=%s", staleReleases.Load(), time.Since(start).Round(10*time.Millisecond))
}
Run it without the watchdog, then with it:
go run ./cmd/step3-ttl
go run ./cmd/step3-ttl -watchdog
accepted=5 balance=80 expected=50 BROKEN
2026/10/09 17:17:57 refused stale releases=5 elapsed=740ms
accepted=5 balance=50 expected=50 OK
2026/10/09 17:18:02 refused stale releases=0 elapsed=1.52s
One run of each is shown; all five runs of each gave the same counts.
Without the watchdog, the callers overlapped and the balance lost three of five debits. All 5 releases were refused, because each lock had expired or changed hands, and Step 2's token check kept each caller from deleting someone else's. On call, this looks like lock not held errors and a balance that doesn't match the payments.
With the watchdog, the callers ran one after another (about 1.5 s for 5 jobs of 300 ms) and the balance is correct. But the watchdog runs alongside the work, which is also its weakness:
A watchdog can't save a paused process
The watchdog is a goroutine in the same process, so a stop-the-world pause, a VM migration or a frozen container stops it too. The lease expires and your code resumes believing it still holds the lock. Redis also doesn't use a monotonic clock for TTL expiry, so a wall-clock jump can expire a lock early. Steps 5 and 6 make the database catch what slips through.
Step 4: Use bsm/redislock with Jittered Retries#
This step swaps the hand-written lock for a library and fixes how callers wait. In a real service I'd rather not maintain lock code, so it uses bsm/redislock, a small single-node library on go-redis that you can read in one sitting.
A library also reopens a question the hand-written lock answered quietly: how a caller waits. Exponential backoff waits longer after each failed try, on the same schedule for every caller. Picture 50 people turned away at once and told to come back in 16 ms three times, then 32, then 64, then every 128 ms. Each time, they all return together, one gets in, and the rest are sent away again.
Several of bsm/redislock's v0.10.0 defaults behave differently from what you might assume:
| Behavior in v0.10.0 | What it means for you |
|---|---|
| Retries are off by default | Without a RetryStrategy, Obtain tries once and fails if the lock is busy |
No deadline on ctx means the deadline becomes the TTL |
A caller never waits longer than the TTL unless you pass your own deadline |
| No watchdog | Call Lock.Refresh yourself; the README's example refreshes every TTL/3, like Step 3 |
A wait cut short by the context returns errors.Join(ErrNotObtained, ctx.Err()) |
Compare with errors.Is, never == |
ExponentialBackoff(min, max) has no jitter; it and LimitRetry keep state |
Add jitter yourself. Strategies are safe to share across goroutines (since v0.7.1), but counts carry over, so a shared LimitRetry becomes a global limit; build one per Obtain |
This program obtains the lock 50 times with exponential backoff between 16 ms and 128 ms, plus an optional wrapper that adds jitter, a random spread on each wait:
package main
import (
"context"
"errors"
"flag"
"log"
"math/rand/v2"
"sync/atomic"
"time"
"github.com/bsm/redislock"
"example.com/walletlock/wallet"
)
// fullJitter waits a random time in [1ms, d+1ms), where d is the inner
// strategy's backoff, so callers that failed together don't retry together.
// The 1ms floor matters: redislock treats a backoff under 1ns as "stop
// retrying", so a bare rand.N(d) that returned 0 would end retries silently.
type fullJitter struct{ inner redislock.RetryStrategy }
func (j fullJitter) NextBackoff() time.Duration {
d := j.inner.NextBackoff()
if d <= 0 {
return d // 0 means stop retrying
}
return time.Millisecond + rand.N(d)
}
func main() {
jitter := flag.Bool("jitter", false, "randomize each backoff")
flag.Parse()
ctx := context.Background()
rdb := wallet.Connect()
defer rdb.Close()
locker := redislock.New(rdb)
if err := rdb.Set(ctx, wallet.BalanceKey, 100, 0).Err(); err != nil {
log.Fatal(err)
}
var timedOut, equalsMatched atomic.Int64
start := time.Now()
accepted := wallet.Hammer(50, func() (bool, error) {
// A new strategy per Obtain: the exponential one counts attempts.
var retry redislock.RetryStrategy = redislock.ExponentialBackoff(16*time.Millisecond, 128*time.Millisecond)
if *jitter {
retry = fullJitter{retry}
}
actx, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
l, err := locker.Obtain(actx, "lock:wallet:42", 2*time.Second, &redislock.Options{RetryStrategy: retry})
if errors.Is(err, redislock.ErrNotObtained) {
timedOut.Add(1)
if err == redislock.ErrNotObtained {
equalsMatched.Add(1)
}
return false, nil
} else if err != nil {
return false, err
}
defer l.Release(ctx)
return wallet.Spend(ctx, rdb, 10, 2*time.Millisecond)
})
balance, _ := rdb.Get(ctx, wallet.BalanceKey).Int64()
wallet.Check(100, 10, accepted, balance)
log.Printf("gave up=%d (err == ErrNotObtained matched %d) elapsed=%s",
timedOut.Load(), equalsMatched.Load(), time.Since(start).Round(10*time.Millisecond))
}
Run it without jitter, then with it:
go run ./cmd/step4-redislock
go run ./cmd/step4-redislock -jitter
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:25 gave up=5 (err == ErrNotObtained matched 0) elapsed=5s
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:28 gave up=0 (err == ErrNotObtained matched 0) elapsed=2.91s
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:37 gave up=0 (err == ErrNotObtained matched 0) elapsed=130ms
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:38 gave up=0 (err == ErrNotObtained matched 0) elapsed=180ms
Two of five runs of each are shown; counts and timings vary.
Both versions keep the balance correct, so the difference is availability. Without jitter, 0 to 5 of the 50 callers gave up after waiting 5 seconds. They failed together, backed off on the same schedule and woke together, so each round let through roughly one caller, and some never got a turn before the 5-second deadline. This is a thundering herd: for the user, a "payment failed" screen after a five-second wait.
Full jitter spreads the wake-ups: each caller waits a random time up to its backoff, and all 50 finished in 130 to 180 ms on every run.
The == count shows a separate bug. Every caller that gave up returned an error that errors.Is(err, redislock.ErrNotObtained) matched and == did not: v0.10.0 joins the context error onto ErrNotObtained whenever a retry wait is cut short by the context deadline, including the default one it derives from the TTL. The README's example passes no retry strategy, so its == works there; add a strategy and it silently stops matching.
Step 5: Add Fencing Tokens So PostgreSQL Rejects Stale Writers#
This step makes the database reject a write from a holder that lost its lease without knowing. Say the first tap's request (A) takes the lock, reads 100 and stalls in a garbage-collection pause longer than its lease. The retry (B) takes the lock, pays and finishes. Then A wakes up, still sure it holds the lock.
Think of a deli counter that serves only a ticket higher than the last one it served: ticket 33, back after 34, gets turned away. Martin Kleppmann's answer to the paused holder works the same way. A fencing token is a number that increases with every acquisition and travels with every write; the storage refuses any token lower than the highest it has accepted. With the balance in PostgreSQL, the check happens in the statement that changes it:
Client A wakes up holding a lease that expired. Its write carries token 33, PostgreSQL has already accepted 34, so the write changes nothing.
Start PostgreSQL with a throwaway password for this container only, and wait until pg_isready reports accepting connections:
export POSTGRES_PASSWORD=your_local_password
docker run -d --name postgres -e POSTGRES_PASSWORD="$POSTGRES_PASSWORD" -p 5432:5432 postgres:18
export DATABASE_URL="postgres://postgres:${POSTGRES_PASSWORD}@localhost:5432/postgres"
docker exec postgres pg_isready -U postgres
/var/run/postgresql:5432 - accepting connections
Create the table. fence records the highest token that has written, and the CHECK constraint is a last guard against a negative balance. Then load it:
CREATE TABLE wallets (
id bigint PRIMARY KEY,
balance bigint NOT NULL CHECK (balance >= 0),
fence bigint NOT NULL DEFAULT 0 -- highest fencing token that has written
);
INSERT INTO wallets (id, balance) VALUES (42, 100);
docker exec -i postgres psql -U postgres < schema.sql
docker exec -it postgres psql -U postgres -c 'SELECT * FROM wallets'
CREATE TABLE
INSERT 0 1
id | balance | fence
----+---------+-------
42 | 100 | 0
(1 row)
Redis issues the tokens with INCR, which adds one to a counter and returns the new value. The script takes the lock and increments the counter atomically, so only a caller that got the lock gets a token, and it rejects a TTL under 1 ms (Redis refuses PX 0). Both keys share the hash tag {wallet:42}, because Redis Cluster requires every key a script declares to hash to the same slot and hashes only the text inside the first braces.
package lock
import (
"context"
"crypto/rand"
"errors"
"fmt"
mrand "math/rand/v2"
"time"
"github.com/redis/go-redis/v9"
)
// Take the lock and, only if that worked, bump the fencing counter.
// Both keys share the {wallet:42} hash tag, so they live in the same
// Redis Cluster slot and the script can touch both.
var acquireFencedScript = redis.NewScript(`
if redis.call("SET", KEYS[1], ARGV[1], "NX", "PX", ARGV[2]) then
return redis.call("INCR", KEYS[2])
end
return false`)
// AcquireFenced works like Acquire but also returns a fencing token:
// a number that grows by one every time anyone takes this lock.
func AcquireFenced(ctx context.Context, rdb *redis.Client, key, fenceKey string, ttl time.Duration) (*Lock, error) {
// Under 1ms, PX would be 0, which Redis rejects as an invalid expire time.
if ttl < time.Millisecond {
return nil, ErrBadTTL
}
token := rand.Text()
for {
fence, err := acquireFencedScript.Run(ctx, rdb, []string{key, fenceKey}, token, ttl.Milliseconds()).Int64()
if err == nil {
return &Lock{rdb: rdb, Key: key, Token: token, Fence: fence}, nil
}
if !errors.Is(err, redis.Nil) {
return nil, err
}
wait := 5*time.Millisecond + mrand.N(10*time.Millisecond)
select {
case <-ctx.Done():
return nil, fmt.Errorf("acquire %s: %w", key, ctx.Err())
case <-time.After(wait):
}
}
}
The next program replays the diagram with A pausing 250 ms against a 100 ms TTL, and seeds the counter so the tokens are 33 and 34. With the check, expect A's write to bounce; without it, expect Step 1's bug:
package main
import (
"context"
"flag"
"fmt"
"log"
"os"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"example.com/walletlock/lock"
"example.com/walletlock/wallet"
)
const (
lockKey = "lock:{wallet:42}"
fenceKey = "fence:{wallet:42}"
ttl = 100 * time.Millisecond
)
// The write succeeds only if nobody with a newer token has written since.
const fencedUpdate = `
UPDATE wallets SET balance = $1, fence = $2
WHERE id = 42 AND fence < $2`
const unfencedUpdate = `
UPDATE wallets SET balance = $1, fence = $2
WHERE id = 42`
func main() {
fenced := flag.Bool("fenced", true, "check the fencing token in the UPDATE")
flag.Parse()
ctx := context.Background()
db, err := pgxpool.New(ctx, os.Getenv("DATABASE_URL"))
if err != nil {
log.Fatal(err)
}
defer db.Close()
rdb := wallet.Connect()
defer rdb.Close()
// Fresh state: balance 100, and a fencing counter that hands out 33 next.
db.Exec(ctx, `UPDATE wallets SET balance = 100, fence = 0 WHERE id = 42`)
rdb.Set(ctx, fenceKey, 32, 0)
rdb.Del(ctx, lockKey)
update := unfencedUpdate
if *fenced {
update = fencedUpdate
}
spend := func(name string, l *lock.Lock, read int64) {
tag, err := db.Exec(ctx, update, read-10, l.Fence)
if err != nil {
log.Fatal(err)
}
if tag.RowsAffected() == 0 {
fmt.Printf("%s: write with fence %d REJECTED\n", name, l.Fence)
return
}
fmt.Printf("%s: wrote balance %d with fence %d\n", name, read-10, l.Fence)
}
readBalance := func() (b int64) {
db.QueryRow(ctx, `SELECT balance FROM wallets WHERE id = 42`).Scan(&b)
return b
}
a, err := lock.AcquireFenced(ctx, rdb, lockKey, fenceKey, ttl)
if err != nil {
log.Fatal(err)
}
aRead := readBalance()
fmt.Printf("A: got the lock, fence %d, read balance %d\n", a.Fence, aRead)
fmt.Println("A: paused for 250ms (GC pause, slow network, stopped VM)")
time.Sleep(250 * time.Millisecond)
b, err := lock.AcquireFenced(ctx, rdb, lockKey, fenceKey, ttl)
if err != nil {
log.Fatal(err)
}
bRead := readBalance()
fmt.Printf("B: got the lock, fence %d, read balance %d\n", b.Fence, bRead)
spend("B", b, bRead)
b.Release(ctx)
fmt.Println("A: wakes up and carries on")
spend("A", a, aRead)
if err := a.Release(ctx); err != nil {
fmt.Println("A: release:", err)
}
var balance, fence int64
db.QueryRow(ctx, `SELECT balance, fence FROM wallets WHERE id = 42`).Scan(&balance, &fence)
fmt.Printf("final: balance=%d fence=%d (two spends of 10 were attempted)\n", balance, fence)
}
Run it with the fencing check, then without:
go run ./cmd/step5-fencing
go run ./cmd/step5-fencing -fenced=false
A: got the lock, fence 33, read balance 100
A: paused for 250ms (GC pause, slow network, stopped VM)
B: got the lock, fence 34, read balance 100
B: wrote balance 90 with fence 34
A: wakes up and carries on
A: write with fence 33 REJECTED
A: release: lock not held (expired or taken by someone else)
final: balance=90 fence=34 (two spends of 10 were attempted)
A: got the lock, fence 33, read balance 100
A: paused for 250ms (GC pause, slow network, stopped VM)
B: got the lock, fence 34, read balance 100
B: wrote balance 90 with fence 34
A: wakes up and carries on
A: wrote balance 90 with fence 33
A: release: lock not held (expired or taken by someone else)
final: balance=90 fence=33 (two spends of 10 were attempted)
With the check, A's stale write bounces and only B's payment counts. Without it, both believe they spent 10 and the balance shows one debit: two approved payments, one charge, the Step 1 bug with a lock in place. A's code couldn't have caught it, because A held a valid lease when it read. Only the database, which sees every write in order, could refuse.
The counter lives in Redis, so a failover can lose recent increments and reissue a token. The strict fence < $2 limits the damage: a repeated token can win at most once, and any token at or below the stored fence is refused until the counter passes max(fence). Because the fence key has no TTL, keep allkeys-* eviction policies off any Redis that holds locks.
Fencing also doesn't stop two overlapping holders with different tokens from both writing, which is what Step 6 deals with.
Step 6: Verify Under Load by Checking the Balance#
This step load-tests the whole path with the lock deliberately broken and checks the final balance. It runs the fenced lock, the read from PostgreSQL and the conditional write with 50 callers. With a 1 ms TTL the lock barely holds, so any protection left comes from the database. The -same-version flag adds one more condition to the UPDATE.
package main
import (
"context"
"flag"
"log"
"os"
"sync/atomic"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"example.com/walletlock/lock"
"example.com/walletlock/wallet"
)
func main() {
ttl := flag.Duration("ttl", 2*time.Second, "lock TTL")
sameVersion := flag.Bool("same-version", false, "also require the fence we read to be unchanged")
flag.Parse()
ctx := context.Background()
db, err := pgxpool.New(ctx, os.Getenv("DATABASE_URL"))
if err != nil {
log.Fatal(err)
}
defer db.Close()
rdb := wallet.Connect()
defer rdb.Close()
db.Exec(ctx, `UPDATE wallets SET balance = 100, fence = 0 WHERE id = 42`)
rdb.Del(ctx, "fence:{wallet:42}", "lock:{wallet:42}")
update := `UPDATE wallets SET balance = $1, fence = $2
WHERE id = 42 AND fence < $2`
if *sameVersion {
update += ` AND fence = $3`
}
var rejected atomic.Int64
accepted := wallet.Hammer(50, func() (bool, error) {
actx, cancel := context.WithTimeout(ctx, 5*time.Second)
defer cancel()
l, err := lock.AcquireFenced(actx, rdb, "lock:{wallet:42}", "fence:{wallet:42}", *ttl)
if err != nil {
return false, err
}
defer l.Release(ctx)
var balance, seen int64
err = db.QueryRow(ctx, `SELECT balance, fence FROM wallets WHERE id = 42`).Scan(&balance, &seen)
if err != nil || balance < 10 {
return false, err
}
time.Sleep(2 * time.Millisecond)
args := []any{balance - 10, l.Fence}
if *sameVersion {
args = append(args, seen)
}
tag, err := db.Exec(ctx, update, args...)
if err != nil {
return false, err
}
if tag.RowsAffected() == 0 {
rejected.Add(1)
return false, nil
}
return true, nil
})
var balance int64
db.QueryRow(ctx, `SELECT balance FROM wallets WHERE id = 42`).Scan(&balance)
wallet.Check(100, 10, accepted, balance)
log.Printf("writes rejected by the database=%d", rejected.Load())
}
Run it with a healthy 2-second TTL, then with a 1 ms TTL:
go run ./cmd/step6-verify
go run ./cmd/step6-verify -ttl 1ms
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:39 writes rejected by the database=0
accepted=17 balance=0 expected=-70 BROKEN
2026/10/09 17:18:40 writes rejected by the database=2
accepted=20 balance=0 expected=-100 BROKEN
2026/10/09 17:18:40 writes rejected by the database=1
One 2-second run and two of five 1 ms runs are shown. All five 2-second runs were correct. With a 1 ms TTL, fencing rejected 1 or 2 writes per run and still let 17 to 20 spends through: the token did its job and the balance is still wrong.
The reason is the read. Fencing orders writes but says nothing about what the writer read. With overlapping leases, A (token 33) and B (token 34) can both read 100; A writes 90 first, which is legal, then B writes 90 with a higher token, also legal. B's new balance came from a stale value, and its token couldn't show it.
The fix makes the write depend on what was read. Salvatore Sanfilippo, the creator of Redis, described it in his reply to Kleppmann:
When starting to work with a shared resource, we set its state to "<token>", then we operate the read-modify-write only if the token is still the same when we write.
Salvatore Sanfilippo, Is Redlock safe?
In SQL, the write also requires the fence value it read to be unchanged. This is check-and-set: read a version, do the work, and write only if the version is still the one you read. Because every successful write raises fence, a version never repeats, so a value that changed and changed back can't fool the check:
UPDATE wallets SET balance = $1, fence = $2
WHERE id = 42 AND fence < $2 AND fence = $3
Run the load test with that condition:
go run ./cmd/step6-verify -ttl 1ms -same-version
go run ./cmd/step6-verify -same-version
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:40 writes rejected by the database=9
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:40 writes rejected by the database=11
accepted=10 balance=0 expected=0 OK
2026/10/09 17:18:41 writes rejected by the database=0
Two of five 1 ms runs and one of five 2-second runs are shown; all ten runs came out correct.
Once the version check is in place, the database keeps the balance correct on its own, and the lock's job is to make conflicts rare. With a healthy lease no write was rejected; with a broken one, 9 to 11 of the 50 callers were turned away, each an error the caller can retry, not a wrong balance.
When you're done, remove the containers:
docker rm -f redis postgres
Is a Redis Distributed Lock Safe?#
It depends on what happens when the lock fails. Kleppmann separates an efficiency lock, whose failure means duplicate work, from a correctness lock, whose failure corrupts data. The wallet needs the second kind.
Single-node Redis is a good efficiency lock and an unreliable correctness lock; the Redis docs call a single instance viable "in applications where a race condition from time to time is acceptable". Replication is asynchronous, so a primary can acknowledge your lock and crash before a replica gets it, letting a second client take the lock on the promoted replica: the docs' "SAFETY VIOLATION!". A restart without appendfsync always can lose the key too, and WAIT doesn't make Redis strongly consistent.
Redlock, described on the same Redis locks page, takes the lock on a majority of independent servers to survive losing one. In February 2016, Martin Kleppmann, author of Designing Data-Intensive Applications, published a critique, and Sanfilippo, who designed Redlock, replied the next day. Read both in full; this summary drops most of the nuance.
- Kleppmann judged Redlock against an asynchronous model, where network delays, process pauses and clock errors can be arbitrarily long. Redlock, he argued, assumes they stay small relative to the TTL, and it can't generate fencing tokens, so he called it "neither fish nor fowl" and recommended a consensus system such as ZooKeeper, or at least a database with reasonable transactional guarantees.
- Sanfilippo replied that Redlock needs only a semi-synchronous model, where processes measure time at roughly the same rate, which holds as long as nobody steps the system clock (a slewing NTP daemon adjusts it gradually). He argued its random token supports check-and-set, the idea Step 6 applies, and agreed Redis should use a monotonic clock for expiry; the docs still say it doesn't.
You don't have to settle the argument to build something safe. Once the storage enforces the invariant, a lock that occasionally fails costs retries, so the choice that matters is where that enforcement lives:
| Approach | Use it when | Cost or risk |
|---|---|---|
Single Redis node, SET NX PX plus token |
Efficiency, or correctness with a database guard behind it | Loses locks on failover of a single primary, even with replicas, or on a restart without appendfsync always |
| Redlock, for example redsync v4.18.0 | You want fewer lost locks and can run an odd number of independent primaries (the docs use 5) | Timing assumptions, no fencing tokens, more servers to operate. redsync defaults to an 8 s expiry, 32 tries and a random 50 to 250 ms delay; its README says the algorithm hasn't been formally analyzed |
| Majority across keys, rueidislock (rueidis v1.0.78) | You use the rueidis client and want automatic extension | 5 s key validity, extended every 2.5 s, 2 of 3 keys by default (KeyMajority: 1 for one instance); its context is canceled when the lock is lost; no fencing tokens |
Atomic Lua script, or SET ... IFEQ on Redis 8.4+, on the value itself |
The balance lives in Redis | No lock needed, but the logic must fit in one script, and acknowledged writes can be lost on failover, so the balance is only as durable as Redis replication and persistence |
SELECT ... FOR UPDATE, a version column, a unique constraint or an idempotency key |
The balance lives in PostgreSQL | Usually the right first answer; row locks need a consistent lock order to avoid deadlocks |
etcd concurrency.Mutex |
You need a lock across services and no database can arbitrate | Another distributed system to run; a paused holder can still write late, so keep a check in storage too |
If the balance is already in PostgreSQL, start there. SELECT ... FOR UPDATE serializes the read-modify-write inside a transaction; when one transaction locks several rows, lock them in a consistent order to avoid deadlocks. If the operation fits in one statement, SET balance = balance - 10 WHERE balance >= 10 needs no explicit lock: under READ COMMITTED, a second concurrent UPDATE waits for the first and then re-checks its WHERE clause against the updated row.
Troubleshooting#
| Symptom | Cause | Fix |
|---|---|---|
Bind for 0.0.0.0:6379 failed: port is already allocated |
Another Redis or container already uses the port | Map a different host port, such as -p 6380:6379, and set REDIS_ADDR=localhost:6380 |
A lock key with PTTL of -1 that never goes away |
The lock was set with TTL 0 or redis.KeepTTL (go-redis sends no expiry), or with SETNX then EXPIRE |
Always pass a non-zero TTL in the same SET command; delete the stuck key by hand |
lock not held on release |
The work took longer than the TTL, so the lease expired or passed to another caller | Raise the TTL, add a watchdog, and alert on this error: your lease expired mid-work, so another caller may have overlapped with you |
| Callers time out even though the work is short | Retries without jitter wake in lockstep | Randomize each backoff, as in Step 4 |
| Busy-lock errors logged as Redis failures | err == redislock.ErrNotObtained fails on a joined error |
Use errors.Is(err, redislock.ErrNotObtained) |
ERR unknown command 'delex' (uppercase in redis-cli) |
Redis older than 8.4, or a service without it | Use the Lua release script, or DELIFEQ on Valkey 9.0 and later |
FATAL: the database system is starting up |
PostgreSQL was still initializing when the program connected | Wait until docker exec postgres pg_isready -U postgres reports accepting connections |
| Balance wrong even with fencing | A newer holder read stale data before an older holder wrote | Also require the version you read to be unchanged, as in Step 6 |
Conclusion#
You built a distributed lock on Redis in Go and load-tested it until it broke. A healthy lease kept the balance right; a lease far too short let 17 to 20 spends through where 10 should have, even with fencing, until the write also checked the version it read. A lease holds only while its timing assumptions do: the pause stays shorter than the TTL and the clocks behave. The database sees every write in order and can check each one, so the final check belongs there.
For a team deciding how much to invest, treat a Redis lock as a performance tool and put correctness in the database: a version column, a conditional UPDATE and a load test that checks the final balance. Ask for that load test in code review, and ask what the final balance was.
Next, read the Redis distributed locks page and Kleppmann's How to do distributed locking back to back, then point the Step 6 load test at your own write path with a 1 ms TTL. If the balance survives that, your lock is free to do its real job of keeping contention down.
If your team is building a path where a lock, a reservation or an event moves money between services, I can help design and build it, including the load test that proves the balance holds. You can see how I approach distributed systems, book a free 20-minute intro call, or tell me about your system.
Frequently asked questions
How do you implement a distributed lock in Redis with Go?
Acquire it with a single SET key token NX PX ttl command, where the token is a random value unique to this caller. In go-redis that is SetNX with a non-zero TTL. Release it with a Lua script, or DELEX IFEQ on Redis 8.4 and later, that deletes the key only if it still holds your token. Keep the protected work shorter than the TTL, or refresh the lock while you work.
Why shouldn't you release a Redis lock with DEL?
Because the lock may no longer be yours. If your work outlasted the TTL, the key expired and another caller took the lock. A plain DEL would delete their lock and let a third caller in alongside them. A token-checked release compares the stored value with your token first and does nothing if they differ, so a slow caller can only ever release its own lock.
What is a fencing token?
A number that increases every time anyone acquires the lock, passed along with every write to the protected resource. The storage remembers the highest token it has accepted and rejects writes carrying a lower one. If a client pauses, loses its lease and wakes up later, its write carries an old token and is refused. Redis INCR can generate it, the database must check it, and a read-modify-write should also check the version it read.
Is Redlock safe to use for correctness?
It is disputed. Martin Kleppmann argued that Redlock depends on timing assumptions and cannot generate fencing tokens, so it is too heavy for efficiency locks and not safe enough for correctness. Salvatore Sanfilippo replied that random tokens allow check-and-set and that the timing assumptions are reasonable in practice. Either way, if correctness matters, make the database reject conflicting writes rather than trusting any lock alone.
Should I use a Redis lock or a database lock?
A Redis lock keeps contention down; the database keeps the data correct. If the value lives in PostgreSQL, start with a conditional UPDATE, a version column or SELECT ... FOR UPDATE, which need no extra system. Add a Redis lock in front when you want to stop duplicate work or reduce rejected writes, and keep the database check behind it whenever a wrong result would cost money.
Shahid Yousuf
Senior Software Engineer
I build and consult on software across web, mobile, cloud and AI. Have something in mind? Let's talk.
Hire me Book a free intro call See case studies