Two Different Units of Protection
The most useful way to compare the two is to ask what exactly moves during a failover.
- Failover cluster instance (FCI): the whole SQL Server instance moves. There is one installation, one set of system databases, and one copy of the data on shared storage (SAN, Storage Spaces Direct, an SMB 3 file share or shared virtual disks). When the active node fails, the cluster starts the SQL Server service on another node, which mounts the same disks. The instance keeps its network name, so
master,msdb, logins, Agent jobs, linked servers, credentials and certificates simply come along. - Availability group (AG): a group of user databases moves. Every replica is a separate SQL Server instance with its own storage and its own copy of the data. The primary ships log records to the secondaries, which redo them continuously. On failover a secondary becomes the primary, and clients follow the listener name to it. Everything outside the user databases stays where it was: each replica has its own
masterandmsdb.
Almost every practical difference follows from this: one copy versus several copies, and instance scope versus database scope.
What Survives Which Failure
High availability is only meaningful against a concrete failure. This table is the one I would put in front of anyone who has to sign off on the design.
| Failure | Failover cluster instance | Availability group |
|---|---|---|
| Node crash, OS hang, hardware fault | Covered. Service restarts on another node, crash recovery runs. | Covered. Synchronous secondary takes over. |
| Patching OS or SQL Server | Rolling, one failover. Downtime = service start + recovery. | Rolling, one failover. Downtime usually a few seconds. |
| Storage failure or LUN loss | Not covered. One copy of the data; all nodes lose it. | Covered. Each replica has its own storage. |
| Corrupted page on disk | Not covered. Restore from backup. | Often repaired automatically from a replica (automatic page repair). |
| Loss of a data center | Only with storage replication (SAN or Storage Replica) and a multi-subnet cluster. | Covered with an asynchronous replica in the second site. |
DELETE without WHERE, dropped table |
Not covered. | Not covered. The mistake is replicated in milliseconds. |
| Agent job, login, linked server deleted on primary | Not applicable, there is only one instance. | Only the primary is affected, but the replicas may already have been out of sync before. |
Two rows deserve emphasis. The storage row is the classic blind spot of FCIs: three nodes look redundant on a diagram, but they share one array. And the DELETE row applies to both: neither technology replaces backups, and an AG replicates human error faster than any batch job could. If you need protection against logical errors, look at log shipping with a restore delay in addition to HA.
The Questions That Decide It
1. Which edition are you licensing?
This narrows the choice before any technical argument starts.
| Standard Edition | Enterprise Edition | |
|---|---|---|
| Failover cluster instance | Yes, 2 nodes | Yes, as many nodes as the cluster supports |
| Availability group | Basic AG only: 2 replicas, 1 database per AG, no readable secondary, no backups or integrity checks on the secondary | Full AG: up to 9 replicas (5 synchronous since SQL Server 2019), readable secondaries, backups on secondaries, distributed AGs, contained AGs |
On Standard Edition the FCI is often the more practical option: one 2-node FCI protects all databases and all instance objects, while Basic AGs mean one AG and one listener per database. With 40 databases that is 40 AGs to create, monitor and fail over. Basic AGs make sense for a handful of important databases, not for a consolidated instance.
Licensing of the second node is the same for both: with Software Assurance, one passive replica or passive node used only for failover does not need its own core licenses. Since the 2019 licensing terms, a passive replica may also run log and full backups and DBCC CHECKDB. Once a secondary serves read queries, it must be fully licensed. Check the current Product Terms for your agreement before you plan around this.
2. How much of your application lives outside the user databases?
This is the question that is skipped most often, and it causes the most trouble after go-live. An AG replicates the contents of its databases and nothing else. Everything the application needs at instance level has to be kept in sync on every replica:
- SQL logins, with identical SIDs, otherwise database users become orphaned after failover
- SQL Server Agent jobs, schedules, operators, proxies and credentials, and every job needs logic for "only run on the primary"
- Linked servers, server-level permissions and server roles
- Server certificates, including the TDE certificate, without which an encrypted database cannot be opened on the new primary
- Server configuration (
sp_configure, trace flags, Database Mail), server audits, endpoints
None of this is hard to solve, but it is ongoing work and it fails quietly. The typical symptom is a failover that works perfectly at database level and an application that still cannot log in. I have described a scripted approach in Synchronizing AlwaysOn Logins with PowerShell. SQL Server 2022 added contained availability groups, which carry their own master and msdb with users, logins, permissions and Agent jobs. That closes much of the gap, with its own limitations; see standard vs. contained availability groups.
An FCI has none of these problems, because there is only one instance. For applications with many jobs, cross-database queries, linked servers or SSIS packages in the file system, this alone can be the deciding argument.
3. What are your real RTO and RPO?
RPO: both deliver zero data loss for committed transactions in the local case. The FCI does it because there is only one copy of the data. The AG does it with synchronous commit, which means every commit waits for the secondary to harden the log. That wait shows up as HADR_SYNC_COMMIT and is a real cost on write-heavy systems; see diagnosing HADR_SYNC_COMMIT. For a remote site, AGs usually run asynchronous, and the RPO is then whatever was in the send queue at the moment of the disaster.
RTO: here the AG has a clear advantage. An FCI failover is a service restart: detection, disk arbitration, service start, then crash recovery of every database. Recovery time depends on how much log has to be redone and undone. Indirect checkpoints keep redo short (checkpoint types and configuration), Accelerated Database Recovery keeps undo short, but a failover under load of 30 seconds to a few minutes is normal, and the buffer pool starts cold. An AG secondary is already online and redoing continuously, so a failover usually takes a few seconds. The exception is a secondary with a large redo queue, which must catch up before the database is available.
If the business requirement is "under a minute", the FCI will meet it on a good day and miss it on a bad one. If "a few minutes" is acceptable, both are fine.
4. Do you need a second site or readable copies?
Multi-site FCIs exist, but they need storage replication underneath, either SAN replication or Windows Storage Replica, and the storage then fails over together with the cluster. That is a lot of moving parts for a team that does not run it every day. An asynchronous AG replica in the second data center is the simpler and more common design, and distributed availability groups extend this across clusters, which is also the cleanest path for migrations with minimal downtime.
Readable secondaries for reporting, backups and DBCC CHECKDB exist only with an Enterprise Edition AG. An FCI has exactly one active instance; the passive nodes do no work at all. Read offload is not free, though: reporting on readable secondaries covers the redo blocking and version store effects.
5. How many databases, and how much change?
An FCI does not care whether it hosts 5 or 500 databases. An AG does: every database uses worker threads for log transport and redo, every new database has to be seeded and joined, and the log of the primary cannot be truncated while a secondary is behind (log_reuse_wait_desc = AVAILABILITY_REPLICA). A disconnected secondary over a long weekend can fill the log drive of the primary. Environments where databases are created and dropped daily by an application need automation for joining them to the AG; without it, new databases are simply unprotected and nobody notices.
6. Who runs it at 3 a.m.?
Both rest on a Windows Server Failover Cluster, so quorum, witness and cluster networking have to be understood either way (quorum and witness explained). On top of that, the FCI adds shared storage and its dependencies, and the AG adds replica states, synchronization health, send and redo queues, listener configuration and the instance-level synchronization described above. In my experience the AG needs more routine attention, while the FCI fails less often but harder when the storage side is involved. Choose the one your team can diagnose under pressure.
The Client Side
Both present a single network name: the virtual network name of the FCI, or the AG listener. Applications should connect only to that name, never to node names. For multi-subnet designs, use MultiSubnetFailover=True in the connection string, otherwise clients may try the offline IP address first and wait for a timeout. With an AG, read-only routing also requires ApplicationIntent=ReadOnly. See SQL Server connection strings for the details, and DNS aliases if you need to keep old server names working for legacy clients.
Combining Both
An FCI can be a replica in an availability group: for example, a 2-node FCI on shared storage in the primary data center and a standalone instance as asynchronous AG replica in the second site. This gives fast local failover for node failures and a separate copy of the data for site loss. Two restrictions apply. Automatic failover between AG replicas is not supported when a replica is an FCI, so the site switch is always manual. And you now run both technologies, with the operational load of both. It is a valid design for large environments, not a default.
Decision Guide
| If this is true for you | Lean towards |
|---|---|
| Standard Edition and more than a few databases | FCI |
| Many Agent jobs, linked servers, cross-database queries, SSIS in the file system | FCI, or contained AG on SQL Server 2022 and later |
| No reliable shared storage, or storage is the component you trust least | AG |
| RTO under one minute | AG |
| Second data center or cloud replica | AG (asynchronous replica or distributed AG) |
| Reporting, backups or CHECKDB offloaded from the primary | AG, Enterprise Edition |
| Hundreds of databases, created and dropped by the application | FCI, unless joining to the AG is automated |
| Write-heavy workload with tight latency requirements | FCI, or AG only after measuring synchronous commit overhead |
The Bottom Line
A failover cluster instance protects the server and everything on it, but it has one copy of the data and a restart in every failover. An availability group protects the data with several independent copies and fails over in seconds, but it protects only the databases and leaves the rest of the instance to you. In most new Enterprise Edition environments the AG is the better default, because storage failures and site loss are the risks that actually hurt. On Standard Edition, on consolidated instances with hundreds of databases, and for applications deeply tied to the instance, the FCI is still the simpler and often the more robust choice. Decide by the failures you need to survive, not by which technology is newer.
Related reading: Failover Clustering vs. SQL AlwaysOn, AlwaysOn internals: how the log flows and AlwaysOn setup: traditional vs. automated.