Haskell World logo Haskell WorldWrite code, build worlds
Games

How Online Game Architecture Patterns Are Built to Scale

Scalability in game architecture is not just about handling more players - it is about handling more players without the system's behavior changing in ways that affect fairness, latency, or correctness.

Online game architecture patterns built to scale under load

Scalability in game architecture is not just about handling more players - it is about handling more players without the system's behavior changing in ways that affect fairness, latency, or correctness. A system that scales by degrading gracefully is different from one that scales without degrading at all. Understanding what specific patterns enable which kinds of scalability, and what tradeoffs each pattern carries, is the foundation of designing game backends that hold up under load.

Stateless services and horizontal scaling

The simplest path to horizontal scaling is stateless services. A stateless service handles each request using only the information in that request and shared external state (a database or cache). Any instance can handle any request. Adding capacity means spinning up more instances and routing traffic across them. No coordination between instances is required.

Game APIs that process individual player actions - spin requests, item purchases, profile updates - fit this model well. The player's state lives in the database. The service reads it, applies the action, writes back the result, and returns a response. Whether instance A or instance B processes the request makes no difference to the outcome. A load balancer distributes requests across instances without session affinity, and capacity scales linearly with instance count.

The limitation of stateless services is that they require the shared state - the database - to handle all the load. As instance count grows, database query volume grows proportionally. Eventually the database becomes the bottleneck. The architectural responses to this are caching, read replicas, and horizontal database sharding, each of which introduces new complexity but extends the scalability ceiling.

Event-driven architecture and decoupling

Event-driven architecture separates the production of events from their consumption. When a player completes an action, the service processing that action publishes an event to a message broker. Downstream services - leaderboard updates, notification dispatch, analytics ingestion, social feed generation - subscribe to the event stream and process events independently, at their own pace, without coupling to the original request handler.

This pattern is valuable for game backends because player actions have many downstream effects. A player leveling up might trigger an achievement check, a social feed post, a notification to friends, a leaderboard recalculation, and an analytics event. If the request handler for the level-up action had to wait for all of these to complete synchronously, the response latency would accumulate all of their individual latencies. With event-driven architecture, the request handler publishes a level-up event and returns immediately. Downstream consumers process the event asynchronously.

Kafka has become the standard message broker for high-volume event streams. Its log-based architecture means events are retained for a configurable period, allowing consumers to replay events, catch up after downtime, and process events at different rates without losing data. The retention also provides an audit trail - a complete record of all events that occurred, in order, which is useful for debugging player-reported issues and for compliance purposes.

Functional patterns for scalable state management

Functional programming patterns contribute to scalability in ways that are underappreciated in the game architecture discussion. Immutability makes state safe to share across concurrent processes without locking. Pure functions are easy to parallelize because they have no shared mutable state to coordinate. Event sourcing - representing state as a sequence of immutable events rather than a mutable record - is a functional pattern at the architectural level.

In Haskell specifically, the STM (Software Transactional Memory) library provides composable atomic transactions over shared in-memory state. Game servers with complex shared state - entity positions, world properties, player session data - can use STM to modify multiple pieces of state atomically without explicit lock management. The runtime handles deadlock prevention automatically. This is significantly easier to reason about than manual mutex management, and it scales better because STM transactions compose without requiring global lock ordering analysis.

The CQRS (Command Query Responsibility Segregation) pattern separates write operations from read operations, allowing each to be scaled independently. Writes go through the command path, where they are validated and applied, and events are emitted. Reads go through the query path, which serves pre-computed projections of the current state. Read projections can be cached aggressively because they are updated only when relevant events arrive, not on every query.

Sharding and data partitioning

When a single database instance cannot handle the write load, the data is partitioned across multiple instances. Each instance owns a subset of the data - typically partitioned by player ID or by game world zone - and requests are routed to the correct instance based on the partition key. This is called sharding, and it allows the system to scale writes linearly by adding shards.

The challenge of sharding is cross-shard operations. A query that touches multiple shards - finding all players in a player's friend list, where friends may be on different shards - requires either gathering data from multiple shards and merging it in the application layer, or maintaining a denormalized copy of the needed data on each shard. Neither approach is free; both require engineering decisions about consistency, performance, and complexity.

Approaches used by platforms like ankertoto illustrate a common compromise: shard the primary player data by player ID for write scalability, and maintain a separate, eventually consistent projection for cross-shard queries like social graphs and leaderboards. The projection is updated by the event stream and serves read queries without touching the sharded primary data. The eventual consistency is acceptable for social queries but not for financial queries, which go directly to the sharded primary data.

Connection management at scale

Managing persistent connections for thousands of concurrent players requires infrastructure designed specifically for that load. HTTP servers designed for request-response workloads do not handle ten thousand long-lived WebSocket connections well. Dedicated connection managers - processes whose only job is to maintain connections and route messages - separate the concerns of connection management and game logic processing.

A connection manager holds open connections from players and translates between the player-facing protocol and the internal messaging system. When a player sends an action, the connection manager validates the message format, attaches the player's identity, and publishes it to the internal event bus. When the backend generates a response for a player, the connection manager looks up the player's connection and writes the response. This separation means game logic services do not need to be aware of connections at all.

Actor systems provide a natural model for connection management. Each connected player is an actor with a mailbox. Messages to the player are delivered to the actor, which writes them to the connection. The actor handles backpressure - if the player's connection is slow to drain, the actor's mailbox fills, signaling to the upstream that the player cannot accept more messages. Akka in the JVM world and the actors in Erlang/Elixir were designed precisely for this pattern and have production deployments handling very large connection counts.

Caching strategies in game systems

Caching in game backends requires careful thought about what to cache, how long to cache it, and how to invalidate it. Player profile data that changes rarely - username, avatar, total experience - is a good cache candidate. Current session state that changes on every action is not, because the cache would be invalidated on nearly every write, eliminating the benefit.

Read-through caches, where the cache automatically fetches from the database on a miss and stores the result, simplify application code but require careful TTL (time-to-live) management. A profile cached for ten minutes means that profile changes are invisible to other players for up to ten minutes. For most profile changes this is acceptable. For changes that affect gameplay - a ban status, for example - a shorter TTL or explicit cache invalidation is required.

Consistent hashing in cache clusters allows the cluster to grow without invalidating large portions of the cache. When a new cache node is added, only the keys that map to that node need to be re-fetched from the database. Without consistent hashing, a change in cluster size could invalidate most of the cache simultaneously, creating a cache stampede where the database receives a sudden spike of requests that the cache was previously absorbing.

Scalability is an ongoing engineering discipline, not a feature that gets shipped and completed. The patterns described here provide a foundation, but every game has unique load characteristics that require specific tuning. The teams that scale well are the ones that measure before optimizing, understand why the patterns they adopt work, and are prepared to replace them when the game's growth exceeds what the current approach can handle.

FB
Finnian Burke

Finnian Burke spent eight years building distributed systems before leaving to write about the programming patterns that actually hold up under pressure. He has worked with Haskell, Rust, and Go in production, and writes mostly about the parts of functional programming that turn out to be useful regardless of the language you use day to day.

More posts by Finnian

More in Games