WebTech1223411 logo WebTech1223411Web tech, read closely
Back-End

How Online Game Platforms Scale Their Servers on Busy Nights

Friday evening, 8 p.m. local time. A new season of an online game goes live, a popular streamer starts playing, and within ten minutes the number of people trying to log in is five times higher than on a normal Tuesday.

Abstract illustration for How Online Game Platforms Scale Their Servers on Busy Nights

Friday evening, 8 p.m. local time. A new season of an online game goes live, a popular streamer starts playing, and within ten minutes the number of people trying to log in is five times higher than on a normal Tuesday. Some platforms stay smooth. Others show error pages, endless queues and timeouts for an hour, and the next day the reviews are full of complaints.

I have been on call for several nights like that. The difference between platforms that cope and those that do not is rarely raw hardware. It is preparation and architecture, decided weeks before the busy night arrives.

Traffic in online games is spiky by nature

Most web services have fairly gentle daily curves. Online game platforms do not. Traffic jumps at the start of events, after patches, when streamers go live and at the same evening hours in each time zone. Peaks of three to ten times the average load are normal.

That changes the design goal. A platform built to handle the average comfortably will fall over at the peak. Buying enough fixed servers for the peak means most of them sit idle the rest of the week. The answer is to scale up and down with demand, and to design every part of the system so it can be scaled.

Autoscaling the stateless parts

The easiest components to scale are stateless services: web front ends, login APIs, store pages and matchmaking request handlers. Because any instance can serve any request, you can add more behind a load balancer and they start sharing the work immediately.

Cloud platforms and Kubernetes can do this automatically based on CPU use, request rates or queue length. The catch is speed. A new virtual machine can take a minute or two to start, and a container image that is several gigabytes takes time to download. Experienced teams scale ahead of known events, keep warm spare capacity, and use small images that start in seconds. Reactive autoscaling handles surprises. Planned scaling handles the predictable peaks.

Game servers are harder

Match servers are a different story. Each one holds the live state of a match, so you cannot simply move players between instances mid-game. Platforms use fleet managers such as Agones on Kubernetes, or managed services from the large cloud providers, which keep a pool of ready servers and allocate one when matchmaking forms a match.

The size of that ready pool is the key setting. Too small and players wait in queue while new servers boot. Too large and you pay for idle machines. Good fleet managers adjust the buffer based on the recent rate of match creation and spread servers across regions so players connect to one nearby. How those servers then keep matches consistent is explained in how online game servers keep many players in sync.

Protecting the database

The database is where scaling usually breaks. You can add web servers in a minute, but a primary database has hard limits on writes and connections. On a busy night, thousands of new instances can each try to open connections, and the database collapses under the connection load before it runs out of anything else.

Several techniques keep it healthy. Connection poolers like PgBouncer share a small number of real database connections among many app instances. Read replicas take read-heavy traffic such as profiles and leaderboards. Caches like Redis hold data that changes rarely, such as item definitions or event configuration, so most requests never reach the database. Writes that do not need to be instant, such as analytics events or match history, go into a queue and are processed steadily.

Queues as a safety valve

Queues are the most underrated tool for busy nights. A login queue in front of the platform lets in players at a rate the back end can handle and shows everyone else their position and an estimated wait. That is frustrating, but far less so than random errors, and it prevents the whole system from collapsing.

Internally, message queues such as Kafka or SQS separate fast front-end actions from slower processing. A player finishing a match gets an immediate response, while rewards and statistics are calculated a few seconds later. When I compared how several platforms behaved during peak hours, the ones that stayed responsive had clearly designed around this, including an online game platform such as ankertoto that showed a short, clearly explained queue rather than failing outright during a busy evening.

Why "the cloud scales automatically" is a misleading promise

Cloud marketing often suggests that moving to the cloud means scaling takes care of itself. I think that belief causes more outages than almost any technical bug.

The cloud gives you the ability to add capacity quickly. It does not make your application able to use that capacity. A service that keeps session data in local memory, a database that cannot accept more connections, a third-party payment API with its own rate limit, or a single global lock in the code will all cap your throughput no matter how many servers you add. Cloud accounts also have quotas on how many instances you can start, which teams often discover during the very event they were preparing for. Scaling is a property of the architecture, not of the hosting provider.

Testing before the big night

The only way to know how a platform behaves at five times normal load is to try it. Load testing tools such as k6, Locust or Gatling simulate thousands of players logging in, browsing and joining matches. Run these tests against a production-like environment, increase load until something breaks, fix it, and repeat.

Teams that do this well also run game day exercises, deliberately shutting down a database replica or a whole region to check that failover works. They keep runbooks for common failures, and they monitor what players actually feel, such as login success rate and time to match, not just server CPU.

The busy night, handled

When a platform survives its biggest night smoothly, players rarely notice anything at all. That invisible success comes from stateless services that scale, fleet managers with sensible buffers, a protected database, queues as safety valves, and testing beyond expected limits. For more on designing services that hold up, see how to design a REST API other developers enjoy using, or read more in our Back-End section.

KO
Kofi Oosterhuis

Kofi spent years keeping APIs and game servers alive through traffic spikes and the occasional bad deploy. He covers back-end design, scaling and networking, with a preference for boring systems that do not page anyone at night.

More posts by Kofi

More from the blog