VMware Cloud Foundation 9.1
Part 3Design — Architectural Options & All Component Models
📐 4 Fleet Deploy Models
📏 3 Sizing Profiles
🔀 5 Consumption Models
💾 8 Storage Models
🌐 6 NSX Models
🔒 4 Recovery Options
📋 28 Total Model Categories
đŸ—ī¸ VCF Fleet Deployment Models — 4 Designs
▸ VCF Fleet Deployment Model Selection Spectrum — from Single AZ to Multi-Region
← Simpler More Complex → Basic / Single Site Design Single AZ or region 1 VCF Instance · vSphere HA only N+1 host redundancy No cross-zone fault tolerance Blueprints: ▸ Single Site Minimal Footprint ▸ Single Site (Standard) Site HA — Across Zones Enhanced Design Multiple AZs within single region Fault domains across physical HW groups Protects from: AZ power/network/cooling/HW Unified operations maintained Blueprints: ▸ Multiple Sites in a Single Region ▸ Stretched clusters across AZs Multi-Site / Multi-Region Distributed Design Multiple VCF Instances across regions Single VCF Fleet: 1 VCF Ops + 1 VCF Auto Centralized mgmt + geo-distributed infra Multiple fleets for AMER/EMEA/APAC Blueprints: ▸ Multi-Sites Across Multiple Regions ▸ Multi-Sites Single Region + Add'l Regions VCF Edge (Remote) Edge / Remote Design Min 10 sites · Max 256 cores/site 8+ CPU cores/host · VCF Operations required Edge vCenter: DC or Co-location 10 design patterns incl. 2-node, Argo CD Use Cases: ▸ Retail · Banking · Energy · Healthcare ▸ Manufacturing · Government/Defense
âš™ī¸ Common Elements Across All Fleet Designs
  • Single VCF Operations instance manages all VCF Instances in the fleet
  • Single VCF Automation instance provides self-service across the fleet
  • Centralized management reduces operational complexity + ensures consistency
  • Latency: Must verify all latency requirements between VCF Operations and ALL fleet components before finalizing design
  • Placement: VCF fleet in low-latency, high-bandwidth network segment (ideally)
  • Scale: Multiple fleets may be required (e.g., 1 fleet per region: AMER, EMEA, APAC)
  • Foundation: Basic Design is architectural foundation for all fleet deployments; supports incremental expansion
  • Architecture Flexibility: Single or multi-domain per instance; mixed cluster types within domains; flexible workload segmentation (VDI, K8s, GP)
📊 Fleet Deployment Model Comparison
AttributeBasicSite HAMulti-Site/RegionVCF Edge
AZ Count12+Multiple regions10+ sites
Cross-AZ HA✗✓✓Partial
Stretched Clusters✗✓OptionalOptional
Multiple Instances✗Optional✓✓
ComplexityLowMediumHighMedium-High
Log/MetricsOptionalOptionalRecommendedOptional
📏 VCF Fleet Sizing Models — Simple · HA-Medium · HA-Large
â„šī¸ Planning Workbook
Download the Planning and Preparation Workbook → populate Management Domain Sizing tab to determine exact resource requirements based on chosen design.
Component Simple High Availability — Medium High Availability — Large
VCF Management Services1× Control Plane Node
3× Large Worker Node
3× Control Plane Node
3× X-Large Worker Node
3× Control Plane Node
4× X-Large Worker Node
SDDC Manager1× Default1× Default1× Default
vCenter1× Small1× Medium1× Large
NSX1× Medium3× Medium3× Large
VCF Operations1× Small3× Medium3× Large
License Server1× Default1× Default1× Default
Cloud Proxy1× Small1× Standard1× Standard
VCF Automation1× Medium3× Medium3× Large
Component Simple High Availability — Medium High Availability — Large
VCF Management Services1× Control Plane Node
2× Large Worker Node
3× Control Plane Node
2× X-Large Worker Node
3× Control Plane Node
3× X-Large Worker Node
SDDC Manager1× Default1× Default1× Default
vCenter1× Small1× Medium1× Large
NSX1× Medium3× Medium3× Large
Cloud Proxy1× Small1× Standard1× Standard
VCF OperationsFleet-level (1st Instance) — not redeployed in additional instances
VCF AutomationFleet-level (1st Instance) — not redeployed in additional instances
License ServerFleet-level (1st Instance) — not redeployed in additional instances
â„šī¸ Additional Instance Note
VCF Operations, VCF Automation, and License Server are fleet-level components. Additional VCF Instances leverage the existing fleet components — no redeployment required.
â˜ī¸ Network Consumption Models — 5 Models
ModelInterfaceKey CapabilitiesSupports All Apps OrgSupports VKSNSX FederationImplications
VCF Automation
All Apps Orgs
VCF Automation portal/API VPC-based networking (prescriptive) · Centralized or Distributed VLAN connectivity · Supervisor services · Avi LB ✓✓✗ Admins should NOT interact directly with NSX UI/API (except troubleshooting) · Best for new workloads
Virtual Private Cloud
(VPC)
vCenter UI / NSX UI / API Self-service N/S services · NAT (Centralized+Dist-VLAN) · VPN · VKS (Centralized+Dist-VLAN) · IPAM embedded · Projects + multi-tenancy ✗✓✗ (not via Global Mgr) Without All Apps Orgs: IaaS tenancy at network/security level only · Software-defined allows scaling N/S stack
NSX Segment NSX UI / API Tier-0/Tier-1 gateways + segments · NSX Federation full support · NSX Projects + multi-tenancy · VPN · Advanced L7 security · VM Mobility across L3 ✗✓✓ Full Requires NSX-specific knowledge · No embedded IPAM · NSX segments as DVPG in vCenter (view only – no CRUD)
VLAN Networking vCenter + Switch Admin VMs on DVPG or NSX VLAN segments · Physical fabric handles routing · Familiar traditional model · VKS supported · IPv6 supported ✗✓— Static config · Network admin required for every new VLAN · Max 4094 VLANs in L2 domain · No IaC approach possible
VCF Automation
VM Apps Orgs
VCF Automation portal/API VM-based self-service · Avi LB self-service · Multiple external connections · VLAN extension subnets · External IP Blocks + Infoblox IPAM VM Apps onlyLimited— Traditional VM workloads · Provider delegates LB, FW, org resources · Quotas enforced
â˜¸ī¸ vSphere Supervisor Models — Zones, Control Plane, Storage Topologies
vSphere Zone Types
Zone TypeMappingBenefitsImplications
Single vSphere Cluster ZoneMaps to a single vSphere clusterSimplest deployment modelScale by adding new vSphere cluster as additional workload zone ¡ Removing only cluster requires draining ALL workloads from zone
Multi vSphere Cluster ZoneMaps to multiple vSphere clustersScale compute by aggregating clusters ¡ Easier capacity expansion ¡ Seamless cluster replacements/upgrades/retirements with no downtimeRequires shared storage for cluster decommissioning ¡ Storage policies must be identical between clusters ¡ Network segments uniformly available between clusters ¡ Similar performance characteristics required across clusters ¡ Namespace can only allocate from a single vSphere Cluster from the zone
Control Plane Availability Models
ModelNodesDefault?BenefitsImplications
Simple1 Control Plane Node✓ Default at WLD creationSimplified activation · Minimal resourcesSingle point of failure · Downtime during upgrade · NOT supported for Three Zone Management deployments · Can scale to HA post-activation (Single Management Zone only)
High Availability3 Control Plane Nodes (1 per zone in 3-zone model)✗Protects against single CP node failure · Required for Three Management Zone modelOnly option for Three Management Zone deployment · Additional resources required
vSphere Supervisor Zone Deployment Models
ModelManagement ZonesWorkload ZonesSupports vSphere PodsAll Apps OrgNotes
Simplified Supervisor1 (combined)Same as mgmt✗✗Single CP VM · Single vNIC · No LB · Minimal resources · VM Service only
Single Mgmt Zone — Combined Workload1Same zone✓✓Simple or HA CP · No extra clusters for workloads · Cluster availability impacts both mgmt + workload
Single Mgmt Zone — Isolated Workload1 dedicatedSeparate zones✓✓Simple or HA CP · Mgmt and workload availability independent · Requires extra clusters for workloads
Three Mgmt Zones — Combined Workload3Same as mgmt✓✓HA CP only · Activation via API only · Protected from single cluster outage · No extra clusters for workloads
Three Mgmt Zones — Isolated Workload3 dedicatedSeparate zones✓✓HA CP only · API activation only · Maximum isolation · Additional clusters required
Supervisor Storage Topologies for vSphere Zones
TopologyScopeSupported vSAN StorageZone Failure Behavior
Zonal DatastoreLocal to single vSphere Zone ¡ Attached to hosts in that cluster ¡ Only accessible by workloads in that zonevSAN Cluster ¡ vSAN Cluster + File Svcs ¡ Stretched vSAN ¡ Stretched vSAN + File Svcs ¡ vSAN Storage Cluster ¡ vSAN Storage Cluster + File Svcs ¡ Stretched vSAN Storage Cluster ¡ Stretched vSAN Storage Cluster + File SvcsFailure of zone impacts workloads in that zone; Stretched vSAN options auto-restart workloads in surviving zone
Multi-Zone DatastoreMultiple zones each with own dedicated datastore ¡ No storage-level replication or failovervSAN Cluster ¡ vSAN Cluster + File Svcs ¡ vSAN Storage Cluster ¡ vSAN Storage Cluster + File SvcsApplication must provide own HA/replication ¡ StatefulSet replicas distributed across zones ¡ RWX file volumes local to zone (unavailable on zone failure)
Cross-Zone DatastoreSingle logical datastore spanning multiple zones ¡ Infrastructure-level HA and fault tolerancevSAN Storage Cluster ¡ vSAN Storage Cluster + File SvcsInfrastructure protects from single zone failure ¡ RWX volumes accessible during zone failure (File Svcs model) ¡ App replication still recommended
VCF Automation Models
ModelNodesUse CasesScalingImplications
Simple1 node (small appliance)PoC/Evaluation · Small-scale non-critical · Dev/TestScale-out to HA by resizing to Medium/Large → auto scales to 3 nodesNo application-level HA
High Availability3-node cluster (medium or large)Production · Mission-critical · Enterprise · Future growthScale-up/down (medium ↔ large) · Supports external LBApplication-level availability · Supports external load balancer
🔀 Workload Connectivity Models — 5 N-S Data Path Architectures
ModelN-S PathNATVPNLB SupportAll Apps/VKSMulti-Rack L3NSX FederationvCenter UI
Centralized (NSX Edge) NSX Edge nodes → physical uplinks ✓ (Auto-SNAT, NAT, GW FW) ✓ Avi LB + VNA LB ✓ Both ✓ ✓ ✓
Distributed VLAN (VNA) Direct from ESX host pNIC to fabric ✓ (Default Out NAT, NAT only) ✗ VNA-based LB ✓ Both ✗ — ✓
Distributed VXLAN (EVPN) Host pNIC → EVPN-VXLAN fabric External IP (1:1 stateless NAT) ✗ None native ✗ All Apps/VKS ✓ — ✗
NSX Segment (Tier-0/1) NSX Edge (active/active or active/standby T0) ✓ (T1 stateful) ✓ Advanced T1 LB ✗ All Apps Org ✓ ✓ Full ✗
VLAN Networking Physical fabric (traditional routing) — — External only ✗ All Apps Org Limited (4094 VLAN) — ✓
â„šī¸ Distributed VXLAN Requirements
Distributed VXLAN requires EVPN Route Controller node(s) to exchange MP-BGP EVPN NLRI and a standard EVPN-VXLAN network fabric. Near host pNIC line rate for traffic; no NSX Edge nodes required.
🔧 VCF Management Services Models + Management Network Models
đŸ“Ļ Component Models — What Runs Where
Component1st VCF InstanceAdditional Instances
Fleet Lifecycle✓ Fleet-level—
Salt RaaS✓ Fleet-level—
Log Management✓ Fleet-level (Day-N)—
SDDC Lifecycle✓ Instance-level✓ Instance-level
Telemetry✓ Instance-level✓ Instance-level
Salt Master✓ Instance-level✓ Instance-level
Identity Broker✓ Instance-level✓ Instance-level (Day-N)
Real-Time Metrics✓ Instance-level (Day-N)✓ (Day-N)
Software Depot✓ Instance-level✓ (Day-N)
âš ī¸ Day-N Deployment Note
Log Management, Real-Time Metrics, Identity Broker (in additional instances), and Software Depot (in additional instances) require adding more worker nodes as a Day-N operation.
🚀 Deployment Models — Simple vs HA
ModelControl PlaneWorkersFailure Impact
Simple1 CP node3 workersCP failure → redeploy cluster + restore from backup
High Availability3 CP nodesSpec + count depends on deployment sizeHA for CP + workers; no impact during upgrade
🌐 Management Network Models
  • Standard: VCF components on vSphere Distributed Port Group (DVPG)
  • NSX Stretched Overlay: VCF components on NSX stretched overlay segment + DVPG. Supported ONLY for: VCF Operations, VCF Automation, License Server, VCF Ops for Networks
  • IP Mobility: Stretched overlay enables IP mobility for VCF Ops + Ops for Networks (DR scenarios)
  • VCF Automation DR: Same IPs reusable at recovery site via NSX stretched overlay
  • NSX Federation: Global Managers must be deployed Active/Standby on this model
  • VCF Mgmt Services: Deployed Day 0 on DVPG — NOT on NSX overlay segments
  • NSX Edge Cluster: Additional steps required; BGP routing required for T0 gateway
📊 Operations Models — VCF Operations · Log Management · Ops for Networks · Recovery
ModelNodesAdditional AppliancesScalingImplications
Simple1 nodeLicense Server ¡ Cloud Proxy ¡ SDDC ManagerScale-up + scale-out to HASlower recovery from failure ¡ vSphere HA restarts on host failure ¡ Service interruption possible for monitoring/alerting + fleet mgmt
High Availability3-node cluster: Primary + Replica + Data nodeLicense Server · Cloud Proxy · SDDC Manager · Optional external LBScale-up all nodes · Scale-out with additional data nodes · Cloud Proxy → collector group (manual)Rapid single-node failure recovery · External LB: extra IP+FQDN + SAN cert entries for all nodes + LB FQDN
Continuous AvailabilityNode pair across 2 AZs: Primary ↔ Primary Replica · Data ↔ Data ReplicaLicense Server · Cloud ProxyCross-AZ scaleHighest availability model; requires 2 AZs; data replication between zones
ModelNodesBenefitsKey Notes
Simple1 replicaSmallest footprint; can extend to HADay-2 deploy via VCF Operations ¡ VIP IP + FQDN required at deploy (can add later) ¡ No app HA ¡ vSphere HA restarts on failure
High Availability3+ replica cluster (3–19 replicas)Distributed cluster + integrated LB · Full to partial service maintained on node failureDay-2 deploy via VCF Operations · VIP IP + FQDN required at deploy
Multi-InstanceMultiple independently managed instances (each Simple or HA)Large-scale environments ¡ Full isolation between instances ¡ Remote sites (low BW) ¡ Organizationally separate environments1 instance integrated with VCF Operations; others run standalone Aria Ops for Logs 8.18 ¡ Day-2 deploy (primary via VCF Ops; secondary manually)
ModelPlatform NodesCollector NodesScaleImplications
Simple1 Platform node (XL size required)1 Collector nodeScale-up + scale-out to HA cluster ¡ Supports 10,000 VMs/objectsSlower recovery ¡ vSphere HA restarts ¡ No app HA ¡ 100% CPU+RAM reservation required ¡ Deploy via VCF Operations (Day-2)
High Availability3-node cluster1 Collector nodeGreater scale + availabilityDay-2 deploy via VCF Operations
Recovery OptionScopeWhen to UseKey Requirements
Component Backup & RestoreIndividual component levelIndividual VCF component fails; restore without full Instance recoverySFTP server ¡ Third-party image backup tool
VCF Instance Backup & RestoreEntire VCF Instance (same location)Entire VCF Instance failed; recover in same locationSFTP server ¡ Accounts for sequencing + dependencies ¡ PowerShell automation ¡ Recovery replicates exact identity (VLANs, IPs, hostnames) ¡ Test requires isolated lab (IP conflict prevention)
VCF Fleet DRFleet-level components to another siteEntire VCF Instance failed; must recover in different location2+ VCF Instances + physical sites ¡ IP mobility between sites ¡ Recovers: VCF Ops + Fleet Lifecycle + Salt RaaS + Log Mgmt + VCF Automation
Cyber RecoveryVCF Instance affected by ransomwareRansomware attack; isolated clean room recovery needed2+ VCF Instances ¡ Isolated clean room ¡ Push-button network isolation ¡ EDR integration ¡ On-premises recovery
â„šī¸ DORA Compliance
VCF Instance Backup & Restore satisfies regulatory requirements including the Digital Operational Resilience Act (DORA).
🔐 Security Models — Identity Broker · VCF SSO · Lateral Security (vDefend)
Identity Broker Architecture — 3-Component Auth Design
  • External IdP: Centralized source of digital identities (corporate directory) — authenticates users
  • Identity Broker: Connects corporate IdP to VCF fleet components; enables VCF SSO; streamlines auth across components (except SDDC Manager)
  • VCF Single Sign-On: Centralized authentication, identity propagation, and session management across VCF fleet
Supported External IdPs
Okta (SAML) Okta (OIDC) Ping Identity MS Entra ID MS ADFS Symantec Generic SAML 2.0 Generic OIDC AD/LDAP OpenLDAP
Identity Broker Models
ModelDeploymentNotes
VCF Instance IdP ModelIdentity Broker on VCF Management Services instance in mgmt domainDefault fleet + instance level auth; VCF SSO + SCIM/JIT/LDAP provisioning methods
Multi-Instance IdP ModelEach VCF Instance has own Identity BrokerUsed for additional VCF Instances; fleet-level SSO via 1st Instance broker remains
âš ī¸ Identity Broker Upgrade Requirement
Must redeploy VCF Identity Broker 9.0.x to supported network/datastore before upgrading to 9.1. This is a critical pre-requisite step.
VCF Single Sign-On Models
ModelDescriptionIdP ConfigUse Case
VCF Fleet-Wide SSOSingle Identity Broker per fleet; unified SSO across all VCF Instances; user accesses all fleet resources with single loginOne IdP → 1 Identity Broker → all fleet componentsStandard enterprise deployment; centralized identity
Cross-Instance SSOMultiple Identity Brokers (per instance); SSO works across instances via federationMultiple IdPs possible; Identity Brokers federatedMulti-instance with separate IdP configs per instance; acquired organizations; regulatory isolation
Single VCF SSOSSO scoped to a single VCF Instance only; no cross-instance auth propagationOne IdP → 1 Identity Broker → 1 instanceSmall deployments; isolated VCF Instances; airgapped environments
ModelScopeKey CapabilitiesImplications
License HubCentral vDefend license managementConnected + disconnected ops ¡ Automated renewal reactivation ¡ License endpoints: NSX Mgr + vDefend SSP + Avi Controller ¡ Up to 120 endpoints per License Hub ¡ Usage updated every 180 daysSimplifies license ops ¡ Single pane across multiple vDefend + Avi installations
Security Services Platform (SSP)Appliance-based advanced securitySingle SSP Installer deploys + manages single SSP instance ¡ 1 SSP instance per NSX Manager cluster ¡ Security Intelligence, NDR, Malware Prevention ¡ Security Segmentation Score + Report ¡ Firewall Rule Analysis ¡ DFW for Bare Metal workloads ¡ Security MetricsMgmt domain or workload domain deployment ¡ May need to expand domain resources for SSP
Management Domain SecurityVCF management componentsvDefend DFW prescriptive methodology · Staged approach: Infrastructure → Environments → Application · Terraform automation script for implementationProtects VCF mgmt components; simplifies design + planning
Workload Domain SecurityCustomer workload VMsDFW 1-2-3-4: Four-Stage Prescriptive Segmentation Journey · Stage 1: Segmentation Planning (CSV metadata import) · Stage 2–4: Progressive policy enforcement · RBAC-labeled dynamic security groupsSystematic segmentation implementation; reduces change risk
đŸ›ī¸ VCF Domain Models ¡ vSphere Cluster Models ¡ Distributed Switch Models
Domain Type Comparison
AttributeManagement DomainWorkload Domain
Created ByVCF Installer (auto)VCF Operations (manual/API)
vCenter1 (shared mgmt components)1 dedicated per domain
vCenter SSO DomainShared (vsphere.local)Dedicated per domain; supports dedicated IdP
Max Workload Domains—40 per VCF Instance
Can Run Workloads?✓ (also hosts fleet mgmt)✓ (primary purpose)
NSX Edge Nodes✓✓
Lifecycle ManagementIndependent from WLDIndependent per domain
NSX Instance Sharing — Dedicated vs Shared
AttributeDedicated NSX per WLDShared NSX across WLDs
FootprintHigher (each WLD has own NSX Mgr cluster)Reduced (1 NSX Mgr cluster shared)
AvailabilityIndependent mgmt + control plane per WLD → higher availabilitySingle mgmt + control plane shared → potential single point
Scale LimitPer-WLD limits apply independentlyShared limits; NSX capacity shared across WLDs
LifecycleIndependent upgrade per WLD NSXCoordinated upgrade (impacts all shared WLDs)
Edge SharingEdge clusters per WLD onlyEdge clusters can be shared when sharing NSX
ModelFault DomainsStorageHA TypeRack Failure?AZ Failure?Notes
Single-Rack vSphere Cluster1 (single rack/L2 domain)vSAN HCI / ExternalvSphere HA (host level)✗✗Default for mgmt domain · Simplest model · All hosts L2 adjacent
Multi-Rack Layer 3Multiple storage fault domainsvSAN HCI* or Compute (see storage models)vSphere HA · Rack-level fault domain isolation✓✗L3 between ESX hosts · Isolated L2 per rack · Extra subnets/VLANs per rack · NOT supported for mgmt domain default cluster · More complex networking
Stretched vSphere Cluster2 AZ storage fault domainsStretched vSAN (see notes)vSphere HA across AZs · Witness in 3rd AZ✓✓Can stretch single-rack clusters only* · L3 multi-rack clusters CANNOT be stretched · Storage Client traffic separation = cannot stretch Storage Clusters · DR via vSphere Metro Storage Cluster (vMSC) for FC/NFS
â„šī¸ vSAN Node Minimums
Management Domain Default Cluster: 3 nodes (Simple, tolerates 1 host failure, no auto-rebuild) or 4 nodes (HA, tolerates 1 failure with auto-rebuild). All Workload Domain Clusters: Minimum depends on storage policy; typically 4+ nodes for RAID-6 with auto-rebuild.
  • Role: vSphere Distributed Switches provide centralized provisioning, administration, and monitoring across all associated ESX hosts in a vSphere cluster
  • Network services provided: (1) VMs ↔ physical network + other VMs ¡ (2) VMkernel services (mgmt, vMotion, vSAN, NSX host TEP) ↔ physical network
  • Deployment: At least 1 VDS per vSphere cluster for system traffic; additional VDS for NSX overlay segments on workloads
  • Design qualities driving VDS selection: Availability ¡ Manageability ¡ Performance ¡ Recoverability ¡ Security
  • LACP: UI-native configuration for LACP on VDS supported in VCF 9.1 (previously API-only); includes physical fabric validation prechecks
  • NIOC (Network I/O Control): Used to prioritize traffic types (mgmt, vMotion, vSAN, NFS, iSCSI, NSX TEP) on shared uplinks
  • MTU: 9000 bytes recommended for vSAN and NSX overlay traffic (jumbo frames)
  • Out-of-band vCenter operations: In 9.1, VDS changes (add/remove uplinks, teaming, MTU, PNICs) can be performed directly in vCenter without impacting SDDC Manager
💾 Storage Models — Principal Storage Comparison + vSAN Topologies
Principal Storage Feature Comparison
CapabilityvSANFibre ChannelNFSiSCSINVMe
Full SDDC Solution✓✗✗✗✗
Storage Policy Based Management✓✗✗✗✗
Automated Deployment & Scale✓✗✗✗✗
Automated LCM/Patching/Upgrades✓✗✗✗✗
Stretched Clusters✓✓ (vMSC)✓ (vMSC)✗✗
Remote Clusters✓✓✓✓✓
Compute Only Clusters✓✗✗✗✗
VCF Automated Config✓✗ (manual/import)✗ (manual/import)✗✗
vSAN Storage Models
ModelArchitectureDrive TypeStretched?Key UseImplications
Single-Rack vSAN ESAHCI · Single-tier storageNVMe only✓ (2 AZs)Performance-optimized HCI workloads; default in VCF 9.1Extra network config · CPU/RAM overhead · Extra raw capacity for failure scenarios · Auto RAID-6 default
Single-Rack vSAN OSAHCI · Two-tier (cache + capacity)SATA/SAS/NVMe✓ (2 AZs)Broader hardware compatibility; legacy workloadsSame overhead as ESA; two-tier architecture adds cache tier
vSAN Storage ClusterDisaggregated storage (ESA required)NVMe only✓ (Stretched or Non-Stretched)Storage-dense nodes; compute scales independently of storageDo NOT run VMs directly on storage cluster · Requires vSAN ESA
vSAN Compute ClusterCompute-only; mounts remote vSAN datastoreNone local✓ (2 AZs)Independent compute + storage scaling; Site CouplingRequires vSAN Storage Cluster · No local vSAN storage
vSAN Stretched ClusterSingle logical cluster across 2 AZs + witnessESA or OSA✓ (native)Site-level HA; tolerates full AZ failureWitness host in 3rd site/AZ · vSphere HA auto-restarts VMs in surviving AZ
vSAN Multi-RackvSAN cluster across multiple physical racksESA or OSA✗Rack-level fault tolerance + scalabilityvSAN Fault Domains configured per rack; tolerates rack-level failure
â„šī¸ vSAN HCI Mesh (Datastore Sharing)
Permitted across all vSAN storage models within a VCF domain. Multiple independent vSAN HCI Clusters can consume storage from adjacent vSAN storage resources — flexible cross-cluster capacity sharing.
🌐 NSX Manager, Edge Cluster, Load Balancer & VNA Cluster Models
ModelNodesVIPAPI ScalingUpgrade ImpactNotes
Simple NSX Manager1 nodeVIP (for future expansion)LimitedManagement + control plane impactedvSphere HA protects with high restart priority ¡ Smallest footprint ¡ Slower recovery
HA NSX Manager3-node clusterVIP (UI + API)All API via VIP → single nodeNo impact during upgradeAnti-affinity rule (nodes on different ESX hosts) · Rapid single-node failure recovery · vSphere HA high restart priority
HA NSX Manager + External LB3-node clusterVIP (UI) + External LB VIP (API)API distributed across all 3 nodesNo impact during upgradeBest for high-churn API environments ¡ Requires external LB
NSX GLOBAL MANAGER MODELS (Federation)
HA Global Manager3-node clusterVIP (UI + API)VIP targets single nodeNo impact during upgradeAnti-affinity rules ¡ vSphere HA high restart priority ¡ Centralized management for NSX Federation
HA Global Manager + Standby3-node active + 3-node standbyVIP per clusterActive cluster handles APINo impact during upgradeActive/Standby cross-site model for Global Manager DR ¡ Standby promoted on primary failure
ModelvSphere Cluster BasisForm FactorRack Failure?AZ Failure?Key Notes
Single-Rack Edge ClusterSingle-Rack vSphere ClusterVM or Bare Metal✗✗Standard deployment · Shared storage between racks required for multi-rack variant · North-south throughput limited by edge node count/type
Rack Fault Tolerant (Multiple Single-Rack Clusters)Multiple Single-Rack ClustersVM or Bare Metal✓✗No dependency on physical fabric for rack failure · Optimal fault isolation between racks · Dependency on BGP config for symmetric routing
Rack Fault Tolerant (Multi-Rack L3 Cluster)Multi-Rack Layer 3 vSphere ClusterVM or Bare Metal✓✗Higher N-S throughput across racks · Fewer resources vs multiple single-rack clusters · L2 adjacency for VLANs across racks (VXLAN) required
Dual AZ — NSX Edge Node HAMultiple Single-Rack OR Stretched vSphere ClusterVM or Bare Metal✓✓ Fast failoverFast network failover across AZs · Multiple single-rack: no dependency on physical fabric · Stretched cluster: common datastore required across AZs · BGP config dependency for symmetric routing
Dual AZ — vSphere HA RecoveryStretched vSphere Cluster onlyVM only✓✓ Slow failoverSymmetric routing maintained · Slow failover (vSphere HA restart time) · Requires physical fabric to extend NSX Edge uplink VLANs + TEP VLAN + mgmt VLANs · Common datastore across AZs
Virtual Network Appliance (VNA) Cluster Models
  • Role: VNA node = appliance providing stateful network services (NAT, L4 LB) that cannot be distributed to hypervisors running Distributed Transit Gateway
  • N-S traffic: Traffic requiring stateful services goes through VNA cluster; bypasses NSX Edge nodes entirely
  • Supported services: Default Outbound NAT (Auto-SNAT), NAT, Gateway Firewall, L4 Load Balancing
  • NOT supported: VPN, centralized routing — use NSX Edge Cluster for those
  • New in 9.1: L4 LB for VPCs via VNA; supports VCF Automation All Apps Org + VKS
  • Placement: Deployed within workload domain; size + node count based on stateful service requirements
  • vCenter UI: VNA Cluster managed via vCenter UI (Transit Gateway / VPC context)
  • Near line-rate: Traffic NOT redirected to VNA achieves near host pNIC line rate
  • Distributed VLAN Connectivity: Requires VNA nodes for stateful services; no NSX Edge required
LB ModelTypeIntegrationUse CasesNotes
NSX Native LBSoftware LB on Tier-1 GatewayNSX onlyL4/L7 LB for NSX Segment workloads ¡ VPN termination LBManaged via NSX Policy API; no extra appliance
Avi Load BalancerSoftware-defined L4/L7 (add-on, v32.1.1)NSX + VPC + Transit GWProduction L4/L7 for VMs + Supervisor + VKS ¡ Full self-service via VCF Automation ¡ Quotas/SE limits/app limitsSeparate versioning (32.1.1); NOT in VCF or vSF base SKU; add-on only ¡ External LB for VCF Ops HA (NSX T1 one-arm or Avi)
VNA L4 LBL4 LB via Virtual Network ApplianceVPC / Distributed VLANL4 LB for VPC workloads + VKS; stateful services in distributed topologyRequires VNA cluster deployment; no NSX Edge required for Distributed VLAN
NSX T1 One-Arm LBNSX Tier-1 one-arm LBNSX + management networkVCF Operations HA cluster load balancing (management plane)Used specifically for VCF Ops HA cluster in NSX stretched overlay network model
🔌 Network Fabric Models · VCF Edge Models · Private AI Models
  • Spine-Leaf (IP Fabric): Recommended for VCF deployments; ECMP L3 routing; no STP; BGP/OSPF underlay
  • EVPN-VXLAN Fabric: Required for Distributed VXLAN Connectivity Model; MP-BGP EVPN NLRI exchange; near host pNIC line rate
  • L2 (ToR) Fabric: Supports Single-Rack vSphere Clusters; L2 VLANs extended to hosts; simpler but less scalable
  • L3 Fabric: Required for Multi-Rack Layer 3 vSphere Cluster Model; isolated L2 per rack; extra subnets and VLANs per rack required
  • Key requirement: Verify ALL latency requirements (Site to Site) before finalizing fabric design
  • Uplink speeds: 25 GbE (standard) or 100 GbE (high-performance / AI/ML workloads)
  • LACP: Supported and configurable via VCF Installer/SDDC Manager UI in 9.1 (with physical fabric prechecks)
  • MTU: 9000 bytes (jumbo frames) for vSAN + NSX overlay traffic on physical fabric
âš ī¸ VCF Edge Requirements
RequirementValueJustification
Min Sites10 sitesGeographically diverse edge locations
Min CPU per Host8 CPU coresMinimum resources per edge host
Max CPU per Site256 CPU coresMax scale per individual edge location
Host PlacementPhysically distinct from DCSeparate rack/switch from DC workload hosts
Edge vCenterDC or Co-location facilityDedicated or existing vCenter for Edge workload
VCF OperationsMANDATORYRequired for licensing of Edge environment
đŸ—ēī¸ VCF Edge Use Cases & Patterns
  • Retail: POS, inventory, customer analytics at edge
  • Banking: Branch computing, ATM infrastructure, local compliance
  • Energy: Substation automation, SCADA, operational data processing
  • Healthcare: Medical imaging, patient monitoring, remote diagnostics
  • Manufacturing: OT/IT convergence, factory floor automation
  • Government/Defense: Classified/hardened deployments
  • 10 Design Patterns including: Single Node + Argo CD, 2-Node vSAN edge cluster, Full VCF stack at edge, Compute-only clusters
  • If deploying without full VCF stack: manually deploy VCF Operations + vCenter
🤖 Private AI Foundation Platform Models
ModelAbstraction LevelKey CapabilitiesBest For
vSphere Supervisor (VKS)Kubernetes-nativeGPU partitioning (vGPU/DirectPath) ¡ VKS clusters with GPU node pools ¡ OCI registries (Harbor) ¡ Deep Learning VMsDevOps/Data Scientists using Kubernetes; model training + inference; CI/CD for AI apps
VCF Automation All AppsSelf-service platformAI workload catalog items ¡ Pre-configured GPU templates ¡ Avi LB for AI inference endpoints ¡ VLAN extension for AI data pipelinesPlatform teams delivering AI IaaS; multi-tenant AI consumption
Bare Metal AI ClusterNear-bare-metalDirectPath I/O for GPUs (AMD MI350, NVIDIA) ¡ vMotion + Storage vMotion + Live Patch ¡ RDMA passthrough (ConnectX-7, BlueField-3)Maximum GPU performance; HPC/training workloads; UALink fabric for GPU interconnect
🎮 Private AI Compute Models — GPU Support in VCF 9.1
HardwareTypeModeNew in 9.1
NVIDIA ConnectX-7NIC + ComputeEnhanced DirectPath + RDMA✓ vMotion+LivePatch
NVIDIA BlueField-3DPUEnhanced DirectPath✓ vMotion+LivePatch
AMD MI350GPUEnhanced DirectPath✓ GA in 9.1
AMD IOMMU Virt.IOMMUDirectPath virtualization✓ New in 9.1
Intel E825 NICLOM NICXeon Gen 6 platform✓ New in 9.1
â„šī¸ AI Partnerships
Hugging Face, PyTorch, OPEA, and UALink partnerships in VCF 9.1 for optimized AI workload performance on AMD MI350 + GPU clusters.
Detail
VCF 9.1 — Part 3: Design — Architectural Options & All Component Models | Principal VCF Technical Reference Doc Lines 8,357–16,000 | GA: 12 MAY 2026 · Broadcom / VMware Cloud Foundation 9.1.0.0