Login Failed: Server Is in Script Upgrade Mode (Error 18401), and How Long to Wait Before You Worry

🚨Part of the SQL Server Errors series, the exact messages and what actually causes them.In: Server & Configuration

Msg 18401  ·  Severity 14  ·  State 1
Login failed for user 'sa'. Reason: Server is in script upgrade mode. Only administrator can connect at this time. (Microsoft SQL Server, Error: 18401)
⚡The patch is still installing. Open the ERRORLOG file on disk and watch it, do not restart the service. While it is healthy the log adds a line every second or so and ends with Recovery is complete. If it stops with Error: 912 and Error: 3417, the upgrade script failed and the fix is further down.

This is the error every patch window can raise and almost nobody plans for. The Cumulative Update installed, the service came back, and now every application login is refused with 18401. Nothing is broken yet. SQL Server is running the second half of the patch, and it will not let anyone in until that half is done.

The trap is the word “administrator”. You will try sa, then your own sysadmin login, then maybe the DAC, and get refused every time. In the two reproductions behind this post, all three were refused for the whole window. The way to see what is happening is the error log on disk, not a connection.


Read the Error Log File, Not the Login

You cannot run xp_readerrorlog because you cannot log in, so open the file itself. On Windows it is ERRORLOG with no extension in the instance’s MSSQL\Log folder; on Linux it is /var/opt/mssql/log/errorlog. The path pattern below is the default for a SQL Server 2025 default instance, so adjust the version folder and instance name to yours:

# Follow the log live. Ctrl+C to stop.
Get-Content -Path 'C:\Program Files\Microsoft SQL Server\MSSQL17.MSSQLSERVER\MSSQL\Log\ERRORLOG' -Tail 40 -Wait

What a healthy run looks like. These lines are from a SQL Server 2022 instance being taken from CU26 to CU27, with every login attempt during the window landing between the upgrade lines:

21:21:22.34 spid27s  Database 'master' is upgrading script 'u_tables.sql' from level 268439721 to level 268439751.
21:21:22.90 Logon    Error: 18401, Severity: 14, State: 1.
21:21:22.90 Logon    Login failed for user 'sa'. Reason: Server is in script upgrade mode. Only administrator can connect at this time. [CLIENT: 172.17.0.1]
21:21:23.21 spid27s  Database 'master' is upgrading script 'repl_upgrade.sql' from level 268439721 to level 268439751.
21:21:23.27 spid27s  Executing sp_vupgrade_replication.
21:21:26.22 spid27s  Upgrading publication settings and system objects in database [UpgradeLabNormal].
21:21:34.79 spid27s  Database 'master' is upgrading script 'msdb110_upgrade.sql' from level 268439721 to level 268439751.
21:21:34.80 spid27s  Starting execution of PRE_MSDB.SQL
21:21:35.00 spid27s  Setting database option COMPATIBILITY_LEVEL to 100 for database 'msdb'.
21:21:42.68 spid27s  Starting execution of MSDB.SQL
21:22:05.24 spid27s  Starting execution of EXTENSIBILITY.SQL
21:22:29.46 spid27s  SQL Server is now ready for client connections. This is an informational message; no user action is required.
21:22:29.47 spid27s  Recovery is complete. This is an informational message only. No user action is required.

Three things to read off that. The “is upgrading script” lines are the milestones, and msdb110_upgrade.sql is the big one. The Logon lines are your applications and your own attempts, and they tell you the server is up and refusing on purpose, not down. And “Recovery is complete” is the finish line. When it appears, logins work again, and on both runs it came a few hundredths of a second after the last script.

Once you are back in, the same lines are there for the record:

-- The milestones, the refusals and the finish line, from the current log
EXEC xp_readerrorlog 0, 1, N'upgrading script';
EXEC xp_readerrorlog 0, 1, N'18401';
EXEC xp_readerrorlog 0, 1, N'Recovery is complete';

-- And the build you are now on
SELECT SERVERPROPERTY('ProductVersion') AS build, SERVERPROPERTY('ProductUpdateLevel') AS cu;

What Script Upgrade Mode Actually Is

A Cumulative Update installs in two halves. Setup replaces the binaries, the DLLs and the executable, and restarts the service. The first time the new binaries start, they find the system databases still at the old script level and run the T-SQL upgrade scripts that shipped with the CU, from the instance’s MSSQL\Install folder, to bring master and msdb up to the new level. Until those scripts finish, the instance is in script upgrade mode and ordinary logins get 18401.

It is not only system databases. One of the steps, sp_vupgrade_replication, visits every database on the instance to upgrade replication settings and system objects, which is where the count of databases starts to matter. On the lab run above it logged one “Upgrading publication settings” and one “Upgrading subscription settings” line for each user database it could open, three of each.

Logins are refused because the scripts are rewriting the objects those logins would use. The same mechanism fires on any first start after the binaries change, which includes bumping a SQL Server container image over a persisted data volume. Both reproductions for this post were exactly that: a 2022-latest image pulled over a volume last started by the previous CU, and the log looks identical to a Windows patch.


How Long It Should Take, and the Signs It Is Stuck

On the lab, a single CU step with six small databases took 75 seconds on one run and 66 on the other, measured from the first 18401 to “Recovery is complete”. A real instance with hundreds of databases, an SSISDB, or a patch that jumps several CUs will take longer, and I am not going to put a number on it that I have not measured. What I can give you is how to tell slow from stuck, because they look the same from the application’s side and completely different in the log.

  • Slow: the log is still moving. New “upgrading script”, “Executing” or “Setting database option” lines are still arriving. Leave it alone. The msdb110_upgrade.sql step alone wrote about 2,400 lines on the lab.
  • Failed: the log has an Error: line, then 912, then 3417, then “SQL Server shutdown has been initiated”. The service is not in upgrade mode any more, it is stopped, and restarting it will run the same script into the same error. Go to errors 912 and 3417.
  • Looping: the log shows the same startup banner every few minutes. Something, a cluster resource, a monitoring agent or a colleague, is restarting the service each time it fails. Stop the restarts first, then read the error the first attempt hit.
  • Genuinely hung: no new line for many minutes, no error, and the file’s modified time is not moving. This is the rare one. Before anything else, check the drive the log and tempdb live on, because a full disk stops the log being written at the same moment it stops the script.

What Makes It Slow or Stuck

Slow is usually the number of databases, because of the per-database replication step, plus the size of the jump: several CUs at once means every intermediate script runs. Stuck, in the sense of failed, is one of a short list of documented causes, and the line immediately before Error: 912 in the log names which one you have:

  • Disk full for tempdb or the log. Microsoft’s own example for 912 is Error: 1101, could not allocate a new page for tempdb because of insufficient disk space, straight before the 912. The scripts create temporary objects, and a system drive that was fine for the installer is not always fine for those.
  • Edition. Error: 1712, online index operations are Enterprise only, is the classic. An upgrade script tried an ONLINE = ON operation on a Standard edition instance.
  • Objects the script cannot handle. Errors 2714 (object already exists), 15151 and 15173 (principals it cannot drop or alter, often in SSISDB), and objects owned by certificate-based principals are all on Microsoft’s list.
  • SSISDB in an availability group. Error 945, “Database ‘SSISDB’ cannot be opened”. The upgrade runs in single-user mode and an availability database has to be multi-user, so availability databases are taken offline for the run and skipped. Microsoft’s fix is to remove SSISDB from the AG, patch each node, then add it back.
  • TLS 1.0 disabled on a build that still needed it (error 17182), and error 5133 when the temporary database the script builds cannot be created.

One thing I expected to be on that list was not. A database the upgrade cannot open does not stop it. I set one lab database OFFLINE and another READ_ONLY before the image bump, and the run finished normally. The offline one was skipped with a line saying so, and the read-only one got a note at startup rather than an error:

Could not open database [UpgradeLabOffline]. Replication settings and system objects could not be upgraded. If the database is used for replication, run sp_vupgrade_replication in the [master] database when the database is available.
System objects could not be updated in database 'UpgradeLabReadOnly' because it is read-only.

So if you have databases offline for the patch window, they will not hold it up, but they will not be upgraded either, and the log tells you which ones to revisit.


When a Script Fails: Errors 912 and 3417

This is the bad outcome, and it is loud. The log carries the real error, then 912 naming the script that failed, then 3417 because master could not be recovered, then the service stops. Microsoft’s article for error 1712 shows the whole sequence, and the shape is the same whatever the underlying error:

Error: 1712, Severity: 16, State: 1.
Online index operations can only be performed in Enterprise edition of SQL Server.
Error: 912, Severity: 21, State: 2.
Script level upgrade for database 'master' failed because upgrade step 'ISServer_upgrade.sql' encountered error 917, state 1, severity 15. This is a serious error condition which might interfere with regular operation and the database will be taken offline. If the error happened during upgrade of the 'master' database, it will prevent the entire SQL Server instance from starting. ...
Error: 3417, Severity: 21, State: 3.
Cannot recover the master database. SQL Server is unable to run. Restore master from a full backup, repair it, or rebuild it. ...
SQL Server shutdown has been initiated

The 3417 wording frightens people into restoring master. Do not start there. During a CU that message almost always means the script that upgrades master hit an error and the engine refused to run half-upgraded, not that the database is damaged. The documented path is:

  1. Read the error above the 912. That is the cause. The 912 and 3417 are consequences and tell you nothing on their own.
  2. Start the service with trace flag 902. It skips the upgrade scripts so the instance comes up on the new binaries and you can work. Add -T902 on the Startup Parameters tab in SQL Server Configuration Manager and start the service, or from an elevated prompt: NET START MSSQLSERVER /T902 (NET START MSSQL$INSTANCENAME /T902 for a named instance).
  3. Fix the cause. Free the disk, drop or re-own the object, take SSISDB out of the AG for the duration, or for 1712 install a build where the script was fixed. Microsoft has a short article per error, all reachable from the troubleshooting page linked below.
  4. Remove -T902 and restart. The scripts run again from the start. If the log now reaches “Recovery is complete”, the patch is finished. If it fails again, there is a new error above a new 912, and you go round once more.

Trace flag 902 is documented as troubleshooting only, and running with it permanently is unsupported: the CU is not complete until the scripts have run. It is a door, not a home. Restoring or rebuilding master is the answer only when the instance will not start even with -T902, which means the problem is not the upgrade script, and a rebuild means restoring or recreating every login, job and linked server afterwards.


What Not to Do

Do not restart the service to “give it a kick”. The upgrade scripts run at startup and they run from the top. Stopping the service in the middle throws that run away, and the next start begins the same sequence again, so every restart adds the full duration to your outage and gets you no closer. If it was going to finish in four minutes, you have just made it eight. Read the log instead; the log tells you within seconds whether it is moving.

Do not put the service in a restart loop and walk away. A cluster resource with automatic restart, or a monitoring tool with “restart on failure”, will hammer a failing script every few minutes and fill the log with identical failures. If the first run hit 912, every run will, until the cause is fixed.

Do not add -T902 and forget it. The instance will come up, the applications will connect, the ticket will close, and the CU will sit half-installed until the next patch fails in a way that is far harder to diagnose. Take the flag out the same day.

Microsoft’s reference covers upgrade script failures when applying an update, error 912 and the trace flag 902 steps, and trace flag 902 itself in full.


Common Questions

It says only an administrator can connect. I am sysadmin, so why am I refused?
Because the word does not mean sysadmin. In both reproductions for this post, sa, a second login in the sysadmin role and the dedicated administrator connection were all refused with 18401 for the entire window, 24 refusals on the DAC alone. The DAC connected in the seconds before the first upgrade script started and again the moment the last one finished, which is no use mid-window. Treat the instance as closed until the log says “Recovery is complete”.
How do I know it has actually finished?
Two lines within a hundredth of a second of each other: “SQL Server is now ready for client connections” and “Recovery is complete”. Then SELECT SERVERPROPERTY('ProductVersion') shows the new build. If the ready line appears without “Recovery is complete”, keep watching.
I already restarted it twice. Have I made it worse?
You have made it longer, not worse, as long as the log shows the upgrade lines starting again after each restart. Stop restarting, watch the log, and let this run finish. If instead the log shows a 912 on each attempt, the restarts were never going to help and the failed-script section is where you are.
The application saw 18401 for two minutes and now it is fine. Do I need to do anything?
Confirm “Recovery is complete” is in the log and the build number is the one you installed, then no. That was the patch finishing. What is worth doing before the next window is telling the application owners the window includes this phase, and checking the log for “Could not open database” lines, because any database that was offline during the run was skipped and needs sp_vupgrade_replication run once it is back.
Can I connect while it runs to see progress?
Not in my testing, not even over the DAC. You do not need to: the log file on disk is the progress bar, and Get-Content -Tail 40 -Wait against it is the only monitor that worked throughout the window.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *