The four in the morning restart is beautiful. Nobody is online, the server comes back in forty seconds, the logs are clean, everything is fine. The six in the evening restart is a different animal: forty people hit reconnect at the same moment, the loading screen sticks at the same percentage for everyone, a third of them time out, and Discord fills up with the word "broken" before the server has finished starting.
Nothing is broken. Your server just tried to do forty players' worth of database work in fifteen seconds, and databases have opinions about that. Slow player loading on ESX is almost always one of five specific things, and four of them are measurable in about ten minutes.
What actually happens when somebody connects
The connect path is longer than it looks. Roughly, in order: the client says hello, your deferral handlers run and can hold the player in the queue, es_extended looks their identifier up in users, character selection happens if you run multicharacter, the player spawns, and then every resource that cares about players fires its own query for that player.
That last step is the one that bites. One join is not one query. It is one query for the user row, then one for vehicles, one for property, one for the phone, one for appearance, one for whatever your housing script keeps, and so on across every resource on your server. On a forty resource city, a single join can be twenty round trips to MySQL.
Multiply by forty simultaneous joins and you have eight hundred queries arriving in a clump, competing for the same connection pool, while the server is also still starting resources.
The five causes, in the order worth checking
A missing index on identifier. The first thing to check and the most common single fix. Every lookup keyed on identifier does a full table scan if that column is not indexed, and a full scan of a users table with thirty thousand rows, forty times at once, is exactly the pause your players are complaining about.
SHOW INDEX FROM users;
SHOW INDEX FROM owned_vehicles;
You want identifier indexed on users, and owner indexed on owned_vehicles. Same for any of your own tables keyed on identifier. Adding an index takes seconds on a table this size and is the cheapest performance work available to you.
Connection pool exhaustion. oxmysql runs a pool with a configurable size. When every connection is busy, the next query waits, and the wait shows up as a player stuck on the loading screen doing absolutely nothing. If queries are queueing rather than running slowly, raising the pool helps. If the queries themselves are slow, raising the pool just means more slow queries at once, which is why you check the index first.
A resource doing something silly per player. Somewhere in your server is a script that runs SELECT * FROM something enormous with no WHERE clause every time a player loads. It has always done this. It was fine at eight players. Turn on the slow query log and find it:
SET GLOBAL slow_query_log = 'ON';
SET GLOBAL long_query_time = 0.5;
-- then read the file, and turn it off afterwards
Half a second is a good threshold for this. Anything on a join path taking longer than that is worth a look.
Deferrals that talk to the internet. Whitelist checks, Discord role lookups and external API calls inside your deferral handlers put a network round trip in front of every single join. Individually invisible. During a restart storm, forty of them at once against an API with rate limits, and your queue stops moving entirely. Cache the answers, set a timeout, and decide what happens when the third party is slow, because it will be.
Everything trying to start at once. After a restart the server is starting resources while players are connecting. Resources that build caches at start compete with the join traffic for the same database. Ten seconds of separation fixes a surprising amount of this.
Measuring it instead of guessing
Time the join path. Stamp when a player is seen at deferral, and again when they are fully loaded, and print the gap:
local joinAt = {}
AddEventHandler('playerConnecting', function()
joinAt[source] = os.clock()
end)
AddEventHandler('esx:playerLoaded', function(playerId)
local started = joinAt[playerId]
if started then
print(('[join] %d took %.2fs'):format(playerId, os.clock() - started))
joinAt[playerId] = nil
end
end)
Watch that number during a quiet period and again during the evening restart. A two second join that becomes a forty second join under load is a queueing problem. A twelve second join at three in the morning with nobody else online is a slow query problem, and no amount of staggering will help.
Both numbers matter, and they tell you different things, which is why the guessing stage usually goes wrong.
Fixes that actually move the number
- Index first. Free, instant, and frequently the whole fix.
- Stagger the returning crowd. A join queue that admits a few players a second turns a spike into a slope. Players wait slightly longer on average and nobody times out, which is the trade you want.
- Move work off the join path. Not everything has to be loaded before the player spawns. Property, phone contacts and garage contents can load a moment later without anyone noticing.
- Cache what does not change. Item definitions, job lists, shop prices and vehicle prices get loaded once at start, not per player.
- Announce restarts properly. Ten minutes, five minutes, one minute, and a message on the way back up. The crowd spreads itself out when it knows what is happening.
The restart schedule is part of the problem
Most of this only hurts because everybody comes back at the same moment, and that moment is one you chose.
Three restarts a day at fixed times is the common pattern, and the usual mistake is putting one of them right in the middle of your peak because it seemed tidy. Look at your concurrency graph and move the evening restart to the shoulder rather than the summit. An hour either side of the peak can halve the size of the returning crowd.
Warnings matter more than timing. Ten minutes, five, one, and a message when the server is back. Given warning, players finish what they are doing and drift back over several minutes. Given no warning, every single one of them hits reconnect within the same fifteen seconds, which is the stampede this whole article is about.
A few other things worth building into the schedule:
- Let the server finish starting before the queue opens. A short hold at the start of the boot means resources are running by the time the first player asks for their data.
- Do not restart to fix things. A restart that happens because something went wrong is a restart with no warning, which is the worst kind. Fix the resource, restart on schedule.
- Watch the first two minutes after a restart every time. That window contains most of your database errors for the day, and nobody is ever looking at it.
- Count the timeouts. If players are dropping out of the queue after a restart, that number is your actual score, and it should be zero.
A restart is an event your city experiences three times a day. It deserves the same design attention as any other event you run.
When it is not your server at all
Two things look exactly like slow loading and are not. A player on a bad connection sits at the same percentage as a player waiting on your database, and only one of those is your problem. And an asset heavy server makes new players download hundreds of megabytes before they see anything, which looks like a load time and is really a download.
Tell them apart by checking whether the same player loads fast on a quiet server and slow on a busy one. If it is slow on both and fast for everyone else, it is them.
Practical takeaway
Index the identifier columns, find the one resource running a silly query on join, keep external calls out of your deferrals, stagger the crowd after a restart, and measure the join time so you know whether you fixed anything.
If you are standing a fresh server up and want a clean baseline to measure against before forty resources pile on top, a stock framework install joins in a couple of seconds on a modest box. Anything much worse than that on your own server is work you have added, and work you can take back out.
The evening restart will still be busier than the four in the morning one. It just does not have to be an event.