Reliability fixes for a sports and gaming platform
For a sports and gaming platform made of many independently deployed services, I diagnosed and fixed data-consistency bugs in the odds, betting and settlement paths.
- Client
- Sports and gaming platform
- Role
- Backend engineer
- Year
- 2026
The challenge#
The platform is made up of many independently deployed services: accounts, live odds aggregation, a feed from a third-party betting exchange, a bookmaking engine, casino, and customer and partner frontends. Several services write market and odds data into a shared Redis store while the betting engine reads it. When those writes and reads disagree, users see markets that should not be there, or miss ones that should.
What I did#
I worked on the bugs that sat between services, where no single team's code looked wrong on its own.
Missing markets. Some match-result markets never appeared. I traced it to the exchange feed attaching those markets to a competition-level placeholder event rather than to the match itself. Our importer could not find that event, saved the markets without one, and a later repair job skipped them forever. I fixed the import so these markets resolve to the right match.
Live events that never cleared. Users' live event lists kept showing finished events. There were two causes: when every bet on a market was cancelled nothing marked it as finished, and when a settlement message was dropped the market stayed open. I fixed cancellation to close the market when its last bet goes, made a new bet reopen it, and added a three-layer guard to the query that builds the list.
Shared data. I documented exactly which keys each service writes and reads when a market is published or withdrawn, along with the rule that readers validate cached data before trusting it, since partially written records can exist for a moment.
How it works#
Services share state through Redis hashes and sorted sets keyed by event and market, plus a relational database for bets and settlement, with AWS EventBridge carrying events between pipelines. Most of the fixes were about making each service agree on that shared contract.
Results#
- The missing match markets now appear for users.
- Finished events no longer linger in live lists.
- The publish and withdraw contract between services is written down, so the next change does not reopen the same bugs.
What I learned#
In a system of many services, the bugs live in the gaps between them. Reading another team's writer alongside your reader, key by key, finds problems that tests on either side never will.