powershelldba.de · Uwe Janke

SQL Server Failover Cluster or Always On Availability Groups? Which High-Availability Solution Should You Choose?

Both run on a Windows Server Failover Cluster, both fail over automatically, and both are called "Always On" in Microsoft's marketing. That is where the similarity ends. A failover cluster instance protects a server. An availability group protects databases. Which one you need depends on what is actually likely to fail in your environment, which edition you license, and how much of your application lives outside the user databases.

The short answer: choose an availability group if you need a second copy of the data, a second site, readable secondaries or failover in seconds. Choose a failover cluster instance if you have reliable shared storage, many databases or many instance-level dependencies (jobs, logins, linked servers, certificates), and an RTO measured in minutes is acceptable. If storage is your biggest risk, the FCI is the wrong answer no matter how attractive the rest looks.

Two Different Units of Protection

The most useful way to compare the two is to ask what exactly moves during a failover.

Almost every practical difference follows from this: one copy versus several copies, and instance scope versus database scope.

What Survives Which Failure

High availability is only meaningful against a concrete failure. This table is the one I would put in front of anyone who has to sign off on the design.

Failure Failover cluster instance Availability group
Node crash, OS hang, hardware fault Covered. Service restarts on another node, crash recovery runs. Covered. Synchronous secondary takes over.
Patching OS or SQL Server Rolling, one failover. Downtime = service start + recovery. Rolling, one failover. Downtime usually a few seconds.
Storage failure or LUN loss Not covered. One copy of the data; all nodes lose it. Covered. Each replica has its own storage.
Corrupted page on disk Not covered. Restore from backup. Often repaired automatically from a replica (automatic page repair).
Loss of a data center Only with storage replication (SAN or Storage Replica) and a multi-subnet cluster. Covered with an asynchronous replica in the second site.
DELETE without WHERE, dropped table Not covered. Not covered. The mistake is replicated in milliseconds.
Agent job, login, linked server deleted on primary Not applicable, there is only one instance. Only the primary is affected, but the replicas may already have been out of sync before.

Two rows deserve emphasis. The storage row is the classic blind spot of FCIs: three nodes look redundant on a diagram, but they share one array. And the DELETE row applies to both: neither technology replaces backups, and an AG replicates human error faster than any batch job could. If you need protection against logical errors, look at log shipping with a restore delay in addition to HA.

The Questions That Decide It

1. Which edition are you licensing?

This narrows the choice before any technical argument starts.

Standard Edition Enterprise Edition
Failover cluster instance Yes, 2 nodes Yes, as many nodes as the cluster supports
Availability group Basic AG only: 2 replicas, 1 database per AG, no readable secondary, no backups or integrity checks on the secondary Full AG: up to 9 replicas (5 synchronous since SQL Server 2019), readable secondaries, backups on secondaries, distributed AGs, contained AGs

On Standard Edition the FCI is often the more practical option: one 2-node FCI protects all databases and all instance objects, while Basic AGs mean one AG and one listener per database. With 40 databases that is 40 AGs to create, monitor and fail over. Basic AGs make sense for a handful of important databases, not for a consolidated instance.

Licensing of the second node is the same for both: with Software Assurance, one passive replica or passive node used only for failover does not need its own core licenses. Since the 2019 licensing terms, a passive replica may also run log and full backups and DBCC CHECKDB. Once a secondary serves read queries, it must be fully licensed. Check the current Product Terms for your agreement before you plan around this.

2. How much of your application lives outside the user databases?

This is the question that is skipped most often, and it causes the most trouble after go-live. An AG replicates the contents of its databases and nothing else. Everything the application needs at instance level has to be kept in sync on every replica:

None of this is hard to solve, but it is ongoing work and it fails quietly. The typical symptom is a failover that works perfectly at database level and an application that still cannot log in. I have described a scripted approach in Synchronizing AlwaysOn Logins with PowerShell. SQL Server 2022 added contained availability groups, which carry their own master and msdb with users, logins, permissions and Agent jobs. That closes much of the gap, with its own limitations; see standard vs. contained availability groups.

An FCI has none of these problems, because there is only one instance. For applications with many jobs, cross-database queries, linked servers or SSIS packages in the file system, this alone can be the deciding argument.

3. What are your real RTO and RPO?

RPO: both deliver zero data loss for committed transactions in the local case. The FCI does it because there is only one copy of the data. The AG does it with synchronous commit, which means every commit waits for the secondary to harden the log. That wait shows up as HADR_SYNC_COMMIT and is a real cost on write-heavy systems; see diagnosing HADR_SYNC_COMMIT. For a remote site, AGs usually run asynchronous, and the RPO is then whatever was in the send queue at the moment of the disaster.

RTO: here the AG has a clear advantage. An FCI failover is a service restart: detection, disk arbitration, service start, then crash recovery of every database. Recovery time depends on how much log has to be redone and undone. Indirect checkpoints keep redo short (checkpoint types and configuration), Accelerated Database Recovery keeps undo short, but a failover under load of 30 seconds to a few minutes is normal, and the buffer pool starts cold. An AG secondary is already online and redoing continuously, so a failover usually takes a few seconds. The exception is a secondary with a large redo queue, which must catch up before the database is available.

If the business requirement is "under a minute", the FCI will meet it on a good day and miss it on a bad one. If "a few minutes" is acceptable, both are fine.

4. Do you need a second site or readable copies?

Multi-site FCIs exist, but they need storage replication underneath, either SAN replication or Windows Storage Replica, and the storage then fails over together with the cluster. That is a lot of moving parts for a team that does not run it every day. An asynchronous AG replica in the second data center is the simpler and more common design, and distributed availability groups extend this across clusters, which is also the cleanest path for migrations with minimal downtime.

Readable secondaries for reporting, backups and DBCC CHECKDB exist only with an Enterprise Edition AG. An FCI has exactly one active instance; the passive nodes do no work at all. Read offload is not free, though: reporting on readable secondaries covers the redo blocking and version store effects.

5. How many databases, and how much change?

An FCI does not care whether it hosts 5 or 500 databases. An AG does: every database uses worker threads for log transport and redo, every new database has to be seeded and joined, and the log of the primary cannot be truncated while a secondary is behind (log_reuse_wait_desc = AVAILABILITY_REPLICA). A disconnected secondary over a long weekend can fill the log drive of the primary. Environments where databases are created and dropped daily by an application need automation for joining them to the AG; without it, new databases are simply unprotected and nobody notices.

6. Who runs it at 3 a.m.?

Both rest on a Windows Server Failover Cluster, so quorum, witness and cluster networking have to be understood either way (quorum and witness explained). On top of that, the FCI adds shared storage and its dependencies, and the AG adds replica states, synchronization health, send and redo queues, listener configuration and the instance-level synchronization described above. In my experience the AG needs more routine attention, while the FCI fails less often but harder when the storage side is involved. Choose the one your team can diagnose under pressure.

The Client Side

Both present a single network name: the virtual network name of the FCI, or the AG listener. Applications should connect only to that name, never to node names. For multi-subnet designs, use MultiSubnetFailover=True in the connection string, otherwise clients may try the offline IP address first and wait for a timeout. With an AG, read-only routing also requires ApplicationIntent=ReadOnly. See SQL Server connection strings for the details, and DNS aliases if you need to keep old server names working for legacy clients.

Combining Both

An FCI can be a replica in an availability group: for example, a 2-node FCI on shared storage in the primary data center and a standalone instance as asynchronous AG replica in the second site. This gives fast local failover for node failures and a separate copy of the data for site loss. Two restrictions apply. Automatic failover between AG replicas is not supported when a replica is an FCI, so the site switch is always manual. And you now run both technologies, with the operational load of both. It is a valid design for large environments, not a default.

Decision Guide

If this is true for you Lean towards
Standard Edition and more than a few databases FCI
Many Agent jobs, linked servers, cross-database queries, SSIS in the file system FCI, or contained AG on SQL Server 2022 and later
No reliable shared storage, or storage is the component you trust least AG
RTO under one minute AG
Second data center or cloud replica AG (asynchronous replica or distributed AG)
Reporting, backups or CHECKDB offloaded from the primary AG, Enterprise Edition
Hundreds of databases, created and dropped by the application FCI, unless joining to the AG is automated
Write-heavy workload with tight latency requirements FCI, or AG only after measuring synchronous commit overhead
Whatever you choose: test the failover under load before go-live, measure the time from failure to the first successful application transaction, and repeat the test after every major change. A high-availability design that has never failed over is a hypothesis, not a solution.

The Bottom Line

A failover cluster instance protects the server and everything on it, but it has one copy of the data and a restart in every failover. An availability group protects the data with several independent copies and fails over in seconds, but it protects only the databases and leaves the rest of the instance to you. In most new Enterprise Edition environments the AG is the better default, because storage failures and site loss are the risks that actually hurt. On Standard Edition, on consolidated instances with hundreds of databases, and for applications deeply tied to the instance, the FCI is still the simpler and often the more robust choice. Decide by the failures you need to survive, not by which technology is newer.

Related reading: Failover Clustering vs. SQL AlwaysOn, AlwaysOn internals: how the log flows and AlwaysOn setup: traditional vs. automated.

← Back to Blog