# Platform9 Private Cloud Director Documentation

Complete technical documentation for Platform9 Private Cloud Director - the enterprise VMware alternative. Deploy and manage your private cloud with comprehensive guides, API references.

<h2 align="center"><mark style="color:$primary;">Platform9 Private Cloud Director</mark></h2>

<p align="center">Learn about Platform9 products. Turn your existing infrastructure into a full-featured private cloud with Platform9. Manage VMs and containers at scale with a familiar user experience and automated APIs in a private, secure environment.</p>

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th><th data-hidden data-card-cover data-type="image">Cover image</th></tr></thead><tbody><tr><td><mark style="color:$primary;"><strong>Getting Started</strong></mark></td><td><mark style="color:$info;">Transition from VMware in minutes. Deploy VMs, configure networking, and manage storage with built-in automation, high availability, and unified control.</mark></td><td><a href="https://docs.platform9.com/private-cloud-director/getting-started/transition-from-vmware">https://docs.platform9.com/private-cloud-director/getting-started/transition-from-vmware</a></td><td><a href="/files/UsVZ7oqmLkID517oKSV6">/files/UsVZ7oqmLkID517oKSV6</a></td></tr><tr><td><mark style="color:$primary;"><strong>Self Hosted</strong></mark></td><td><mark style="color:$info;">Full control, on your infrastructure. Deploy a highly available management plane across multiple servers with enterprise-grade automation and support.</mark></td><td><a href="https://docs.platform9.com/private-cloud-director/getting-started/self-hosted">https://docs.platform9.com/private-cloud-director/getting-started/self-hosted</a></td><td><a href="/files/3ZgPj5djBcJ791iOvfxw">/files/3ZgPj5djBcJ791iOvfxw</a></td></tr><tr><td><mark style="color:$primary;"><strong>Community Edition</strong></mark></td><td><mark style="color:$info;">Enterprise private cloud, simplified. Deploy on bare metal or a VM with a single command, with full functionality, and experience a production-grade private cloud with zero licensing costs.</mark></td><td><a href="https://docs.platform9.com/private-cloud-director/getting-started/getting-started-with-community-edition">https://docs.platform9.com/private-cloud-director/getting-started/getting-started-with-community-edition</a></td><td><a href="/files/ChLeibeRH1syQS0b22BE">/files/ChLeibeRH1syQS0b22BE</a></td></tr></tbody></table>

{% columns %}
{% column %}

<figure><img src="/files/KD4tc6h6ALaahmkSEmv3" alt=""><figcaption></figcaption></figure>
{% endcolumn %}

{% column %}

## <mark style="color:$primary;">Release Notes</mark>

<mark style="color:$info;">Continuous innovation with regular releases. New features, enhanced security, and performance improvements. Stay current with the latest capabilities.</mark>

<a href="https://docs.platform9.com/release-notes" class="button secondary">Learn More</a>

{% endcolumn %}
{% endcolumns %}

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><mark style="color:$primary;"><strong>API Documentation</strong></mark></td><td><mark style="color:$info;">Automate everything with REST APIs and the unified pcdctl CLI. Manage compute, storage, networking, and Kubernetes clusters programmatically at scale.</mark></td><td><a href="/files/1DlT0om8foBvzZcMzw34">/files/1DlT0om8foBvzZcMzw34</a></td><td><a href="https://docs.platform9.com/api-docs">https://docs.platform9.com/api-docs</a></td></tr><tr><td><mark style="color:$primary;"><strong>Virtualized Clusters</strong></mark></td><td><mark style="color:$info;">Group hypervisors, enable VM high availability, and auto-balance workloads with Dynamic Resource Rebalancing. Production-ready from day one.</mark></td><td><a href="/files/D0SK9ZJKrxaVhh5Xoo7m">/files/D0SK9ZJKrxaVhh5Xoo7m</a></td><td><a href="https://docs.platform9.com/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint">https://docs.platform9.com/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint</a></td></tr><tr><td><mark style="color:$primary;"><strong>Kubernetes Clusters</strong></mark></td><td><mark style="color:$info;">Managed Kubernetes with hosted control planes, Cluster API automation, and built-in multi-tenancy. Create, scale, and upgrade clusters effortlessly.</mark></td><td><a href="/files/Y9b9g5hLG0QO6FTXmYUb">/files/Y9b9g5hLG0QO6FTXmYUb</a></td><td><a href="https://docs.platform9.com/private-cloud-director/kubernetes-clusters/k8s-overview">https://docs.platform9.com/private-cloud-director/kubernetes-clusters/k8s-overview</a></td></tr></tbody></table>

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><mark style="color:$primary;"><strong>Troubleshooting</strong></mark></td><td><mark style="color:$info;">Diagnose and resolve issues fast. Service-specific debugging guides, log file analysis, and direct paths to resolution for every component.</mark></td><td><a href="/files/92nmQbbL4hALoCQIZ5kS">/files/92nmQbbL4hALoCQIZ5kS</a></td><td><a href="https://platform9.com/kb/pcd-ts">https://platform9.com/kb/pcd-ts</a></td></tr><tr><td><mark style="color:$primary;"><strong>Knowledge Base</strong></mark></td><td><mark style="color:$info;">Searchable solutions from real-world deployments. Common issues, proven fixes, and best practices, continuously updated by Platform9 engineers.</mark></td><td><a href="/files/Ph0EgnzLVZeUek6sngMX">/files/Ph0EgnzLVZeUek6sngMX</a></td><td><a href="https://platform9.com/kb/pcd">https://platform9.com/kb/pcd</a></td></tr></tbody></table>

### Other docs

<table data-view="cards"><thead><tr><th></th><th data-hidden data-card-cover data-type="image">Cover image</th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td>Managed OpenStack</td><td><a href="/files/Y9b9g5hLG0QO6FTXmYUb">/files/Y9b9g5hLG0QO6FTXmYUb</a></td><td><a href="/spaces/zFKnLSEoXJl99ahgFggb">/spaces/zFKnLSEoXJl99ahgFggb</a></td></tr><tr><td>Managed Kubernetes</td><td><a href="/files/1DlT0om8foBvzZcMzw34">/files/1DlT0om8foBvzZcMzw34</a></td><td><a href="/spaces/tZryvDiHvZIhOZofYOWU">/spaces/tZryvDiHvZIhOZofYOWU</a></td></tr><tr><td>Kubernetes Self Managed Cloud Platform</td><td><a href="/files/3ZgPj5djBcJ791iOvfxw">/files/3ZgPj5djBcJ791iOvfxw</a></td><td><a href="/spaces/gOAZI28gS3Bp6LuO0Qjs">/spaces/gOAZI28gS3Bp6LuO0Qjs</a></td></tr></tbody></table>


# Overview

Platform9's <code class="expression">space.vars.product\_name</code> is an enterprise private cloud platform that enables IT & DevOps administrator to manage their virtualized and Kubernetes infrastructure at scale while leveraging their existing investments in storage and server hardware.

<code class="expression">space.vars.product\_name</code> is built using an **open source KVM hypervisor** behind the scenes. It uses a combination of best-of-breed open source technologies for virtualization and container management, along with our unique features to provide a familiar experience to virtualization and container administrators, along with all enterprise capabilities required to build, maintain, and scale private clouds.

### Who is this documentation for? <a href="#who-is-this-documentation-for" id="who-is-this-documentation-for"></a>

The <code class="expression">space.vars.product\_name</code> documentation is for **administrators, operators** **and users** of <code class="expression">space.vars.product\_name</code> who are responsible for creating and maintaining virtualized clusters and applications provisioned on them, on behalf of their organization.

This documentation assumes a level of understanding of fundamental infrastructure management technologies such as virtualization, container orchestration, storage, networking etc.

### Benefit of <code class="expression">space.vars.product\_name</code>

<code class="expression">space.vars.product\_name</code> is a private cloud platform that enables you to manage your VMs and containers side by side. <code class="expression">space.vars.product\_acronym</code> offers following benefits to system administrators:

* **Familiar, virtualization first management interface** - <code class="expression">space.vars.product\_name</code> provides a familiar interface for virtualization administrators, where they can find all traditional virtualization features they've relied on.
* **All enterprise virtualization features built-in** - <code class="expression">space.vars.product\_name</code> comes built in with all key virtualization features such as VM HA, Clusters enabled with Distributed Resource Rebalancing & Scheduling, VM cloning and snapshotting, VM affinity, anti-affinity, and much more.
* **Integrates with all enterprise storage** - <code class="expression">space.vars.product\_name</code> integrates with a wide array of enterprise storage arrays and systems, enabling you to keep your existing storage investments. <code class="expression">space.vars.product\_acronym</code> also integrates with all your existing server hardware, without requiring a change or refresh.
* **In-place vSphere cluster conversion with vJailbreak** - In addition to <code class="expression">space.vars.product\_name</code>, Platform9 also offers a free migration tool for VMware users, called [vJailbreak](https://platform9.com/vjailbreak/). With vJailbreak, you can perform in-place conversion of your vSphere clusters in a rolling manner to convert them to <code class="expression">space.vars.product\_name</code> cluster, using automation.

### Design Principles

Read here to understand the [Design Principles](/private-cloud-director/introduction/design-principles) that <code class="expression">space.vars.product\_name</code> has been built with.

<br>


# Design Principles

We've built <code class="expression">space.vars.product\_name</code> with these design principles.

### Familiar, Virtualization First Experience

We are a product built by VMware engineers, for virtualization administrators. We believe that virtualization must be the bedrock of any modern private cloud platform. That virtualization must be implemented as a first class citizen, and presented without complex abstractions that create complexity. We've utilized our decades of experience working at VMware to build a Private Cloud Platform with a familiar user experience compared to what you are accustomed to with your traditional virtualization interface.

### Integration and Interoperability with Existing Hardware

We aim to be broadly compatible with all major server, storage, and network investments in use by enterprises.

We believe that one of the core value propositions of private clouds is that for most businesses, there are existing hardware investments of diverse makes and generations. Often, this hardware is partially depreciated, reducing its operational costs.

We strive to allow our customers to *sweat* their existing hardware assets, and to do so, we prioritize interoperability and a wide HCL. We strive for our HCL to include all Enterprise Linux compatible hardware, and will degrade features gracefully when niche hardware capabilities are unavailable.

### Opinionated Design and UX

We are inspired by beautifully designed products that simplify end-user operations. While built on powerful open-source technologies, we aim to make these technologies simpler to learn, deploy, and operate at scale.

<code class="expression">space.vars.product\_name</code> will strive to simplify cloud operations and management for IT and platform engineering teams, while ensuring critical control elements are available.

For instance, to provide the best possible experience, we may abstract away lower-level APIs and controls in the underlying technology, offering a higher-level abstraction. At other times, we may make simplifying assumptions that reduce semantic complexity.

This opinionated design will be done in a manner that retains upstream API compatibility, but we may introduce additional APIs - via composition - to support our design philosophy.

We design for IT operations teams and platform engineering teams as the Administrator user, who will then optionally invite other administrator users, as well as DevOps teams and developer teams as self-service users.

### Curated Technologies

<code class="expression">space.vars.product\_name</code> is built using best of breed open source components under the hood, that are then coupled with our IP for enterprise readiness. Open-source projects vary widely in their maturity and readiness for enterprise usage. We select projects and technologies with a view to ensuring the following, in priority order:

1. Projects that can help us deliver a deterministic, highly available experience that enterprises can depend upon (when used with our proactive operations model)
2. Projects that have broader customer interest, vendor participation, and that are not single-vendor technologies
3. Projects that have either achieved maturity or are likely to achieve maturity due to a growing contributor base

### Cloud Completeness via Extensions

A long-standing challenge with private clouds has been their lack of the breadth of features available from public cloud providers. We are clear-eyed about the amount of effort required to match the features and services of hyperscale public clouds; yet, we are ambitious about making private clouds competitive and more attractive to developers.

Our approach to achieving Cloud Completeness will be to build foundational platform features, such as critical Infrastructure-as-a-Service (IaaS) capabilities, and to leverage the ecosystem of open-source technologies, including Kubernetes operators, to enable a broader range of higher-level application services.


# Transition from VMware

If you’re considering transitioning from VMware to an alternative virtualization stack, the <code class="expression">space.vars.product\_name</code> is an excellent choice. Platform9's <code class="expression">space.vars.product\_name</code> offers features similar to those you’ve relied on in VMware for years.

To help you get started, here’s a simple glossary mapping familiar VMware features to their counterparts in the <code class="expression">space.vars.product\_name</code> environment.

| **Your favorite VMware feature**                          | **Equivalent in PCD**                                                                                                                                                                                                                                                                                                                                                                                                                                                                                |
| --------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| VMware vSphere Clusters                                   | In <code class="expression">space.vars.product\_acronym</code>, [Virtualized Clusters](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint) and host aggregates serve a similar purpose to vSphere clusters.                                                                                                                                                                                                                                                                  |
| VMware vCenter Server                                     | The <code class="expression">space.vars.product\_acronym</code> management plane manages infrastructure resources, similar to how vCenter orchestrates VMware environments.                                                                                                                                                                                                                                                                                                                          |
| VMware ESXi Host                                          | <code class="expression">space.vars.product\_acronym</code> uses KVM (Kernel-based Virtual Machine) as the underlying hypervisor.                                                                                                                                                                                                                                                                                                                                                                    |
| VMware Distributed Resource Scheduler (DRS)               | <code class="expression">space.vars.product\_acronym</code> Dynamic Resource Rebalancing (DRR) manages the distribution of workloads & resource consumption across hosts.                                                                                                                                                                                                                                                                                                                            |
| VMware CPU and Memory over-commitment ratios              | <code class="expression">space.vars.product\_name</code> supports over-commitment ratio specification for CPU and Memory at a host or a host aggregate level.                                                                                                                                                                                                                                                                                                                                        |
| VMware High Availability (HA)                             | <code class="expression">space.vars.product\_acronym</code> VM HA (2/4 nodes cluster) feature supports continuous availability of services.                                                                                                                                                                                                                                                                                                                                                          |
| VMware vMotion                                            | <code class="expression">space.vars.product\_acronym</code> supports Live Migration to move instances across hosts without downtime.                                                                                                                                                                                                                                                                                                                                                                 |
| VMware Storage vMotion                                    | <code class="expression">space.vars.product\_name</code> [Storage Live Migration / Volume Retyping](/private-cloud-director/storage/volume#storage-live-migration--volume-retyping)                                                                                                                                                                                                                                                                                                                  |
| VMware Virtual or Distributed Switch (vSwitch / DVSwitch) | <code class="expression">space.vars.product\_acronym</code> uses Open vSwitch to provide Distributed Virtual Switch equivalent capabilities.                                                                                                                                                                                                                                                                                                                                                         |
| VMware Template                                           | In <code class="expression">space.vars.product\_acronym</code>, templates are referred to as [Images](/private-cloud-director/images-and-image-library/image-library---images), which are used to create new virtual machines.                                                                                                                                                                                                                                                                       |
| VMware Snapshot                                           | <code class="expression">space.vars.product\_acronym</code> support VM Snapshots for capturing the current state of a VM.                                                                                                                                                                                                                                                                                                                                                                            |
| VMware NSX (Networking and Security)                      | <code class="expression">space.vars.product\_acronym</code> [Networking Service](/private-cloud-director/virtualized-networking/networking-overview) offers full Software Defined Networking capabilities, including managing physical or virtual networks, routers, IP management, micro-segmentation, security policies, LBaaS, DNSaaS and more.                                                                                                                                                   |
| VMware Datacenter                                         | <code class="expression">space.vars.product\_acronym</code> Regions, Domains, Tenants, and much more                                                                                                                                                                                                                                                                                                                                                                                                 |
| VMware VMFS Datastore                                     | <code class="expression">space.vars.product\_acronym</code> direct mount storage w/ storage LUNs (no shared file system) or with NFS backend                                                                                                                                                                                                                                                                                                                                                         |
| VMware vVol                                               | <code class="expression">space.vars.product\_acronym</code> Volumes, Volume Types, Volume Snapshots                                                                                                                                                                                                                                                                                                                                                                                                  |
| VMDK (Virtual Machine Disk) Format                        | <code class="expression">space.vars.product\_acronym</code> uses the QCOW2 file format for virtual machine disk images.                                                                                                                                                                                                                                                                                                                                                                              |
| vCenter Storage Policies                                  | <code class="expression">space.vars.product\_acronym</code> Storage Types and Storage Groups                                                                                                                                                                                                                                                                                                                                                                                                         |
| vApps and vRealize Orchestrator (vRO)                     | <code class="expression">space.vars.product\_acronym</code> uses Terraform to automate composite cloud application deployment.                                                                                                                                                                                                                                                                                                                                                                       |
| VMware Site Recovery Manager (SRM)                        | <code class="expression">space.vars.product\_acronym</code> supports disaster recovery through integration with most popular third-party backup and disaster recovery solutions. See [Veeam Integration with PCD](/private-cloud-director/integrations/veeam-integration-with-pcd), [Rubrik Integration with PCD](/private-cloud-director/integrations/rubrik-integration-with-pcd) Cohesity, [Commvault Integration with PCD](/private-cloud-director/integrations/commvault-integration-with-pcd). |
| VMware Resource Allocation (CPU/Memory)                   | <code class="expression">space.vars.product\_acronym</code> uses flavors to define VM configuration, including CPU and memory allocations.                                                                                                                                                                                                                                                                                                                                                           |
| VMware Content Library                                    | <code class="expression">space.vars.product\_acronym</code> uses Image Library Service for VM templates and snapshots.                                                                                                                                                                                                                                                                                                                                                                               |


# Pre-requisites

This document outlines the infrastructure prerequisites for setting up your private cloud. If you are looking to deploy the Self-Hosted version, please follow the [Pre-requisites](/private-cloud-director/getting-started/self-hosted/self-hosted-pre-requisites) first.

## Hypervisor Host Prerequisites

### General Hypervisor Configuration Pre-requisites

Each physical server or host that you will use as a hypervisor with <code class="expression">space.vars.product\_name</code> must meet the following requirements:

* x86 server - <code class="expression">space.vars.product\_name</code> only supports x86 server hardware today.
* Running Ubuntu 22.04 LTS (Jammy Jellyfish) server version or Ubuntu 24.04 LTS (Noble Numbat) server version for operating system.
  * Download Ubuntu 22.04 - <https://mirrors.usinternet.com/ubuntu/releases/22.04.5/ubuntu-22.04.5-live-server-amd64.iso>
  * Download Ubuntu 24.04 - <https://mirrors.usinternet.com/ubuntu/releases/24.04.3/ubuntu-24.04.3-live-server-amd64.iso>
* Each server must have hardware virtualization enabled
  * an Intel processor with the Intel VT-x and Intel 64 virtualization extensions enabled or
  * an AMD processor with the AMD-V and the AMD64 virtualization extensions enabled.
  * Read [Hardware Virtualization Extension](/private-cloud-director/getting-started/pre-requisites/hardware-virtualization-extension) for steps to check if your server hardware has virtualization extensions enabled.
* The server must meet the [CPU Model Pre-requisites for Hypervisor Hosts](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model#cpu-model-pre-requisites-for-hypervisor-hosts).
* Your linux installation may come pre-installed with LVM package and your hypervisor operating system maybe be installed using an LVM based partition. If this is the case, you must:
  * Follow [Hypervisor LVM Configuration](/private-cloud-director/getting-started/pre-requisites/hypervisor-lvm-configuration) guide to configure the required LVM filters on each hypervisor host.
  * Not doing so may result in undesirable side effects listed under [Hypervisor LVM Configuration](/private-cloud-director/getting-started/pre-requisites/hypervisor-lvm-configuration#symptoms-when-not-configured).
* Each server should have the following minimum amount of resources:
  1. 8 vCPUs
  2. 16GB RAM
  3. 250 GB storage (to include sufficient disk space for the operating system, Platform9 installer packages, logs, temporary files, virtual machine config files, virtual machine disk files when using ephemeral disks).

{% hint style="info" %}
**Note**

If you plan to only use non-ephemeral (block) storage for your virtual machines, having local storage of 100 GB at per hypervisor server level should be sufficient.
{% endhint %}

* `sudo` access enabled for the Administrator to log into the server and install the Platform9 agent
* Server `hostname` should contain at least one non-numeric character
* Make sure that the content under `/opt/pf9` is not shared across hosts. Either make this a local directory, or if using shared storage, ensure that this path mounts to a unique shared storage file share or volume that is not shared across any other hosts in your setup.
* When using the SaaS-hosted deployment model, outbound connectivity (port 443) must be enabled on each server so that the Platform9 agent can connect to the <code class="expression">space.vars.product\_name</code> SaaS management plane.
* In the case of a multi-domain environment, host onboarding should be done by the Administrator user in the `default` domain and not the secondary domains.
* Each server must run an NTP service (`chronyd`, `ntpd`, or `systemd-timesyncd`) **and** have its system clock actively synchronized. `pcdctl prep-node` mandatorily checks both — a host whose clock is not synchronized cannot be onboarded. Verify with `timedatectl status` and look for `System clock synchronized: yes`.
* If planning to use [VM Live Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration) feature, follow the [Live Migration Prerequisites](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration-prerequisites)
* If planning to use the [Virtual Machine High Availability Vm Ha ](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha)feature, follow the [VM HA Prerequisites](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha#pre-requisites).
* If you plan to use the Dynamic Resource Rebalancing (DRR) feature, follow the [DRR Pre-requisites](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr#drr-pre-requisites)

### Operating System Version Compatibility

<code class="expression">space.vars.product\_name</code> supports both Ubuntu 22.04 LTS and Ubuntu 24.04 LTS. However, when running a mixed environment with both versions, be aware of the following limitations:

**VM Migration Between Ubuntu Versions**

Live migration of virtual machines between Ubuntu 22.04 and Ubuntu 24.04 hosts has known limitations:

* Initial VM migration from Ubuntu 22.04 to Ubuntu 24.04 typically succeeds.
* Subsequent migrations back to Ubuntu 22.04 may fail.
* Multiple round-trip migrations between versions are not supported.

**Rolling Cluster Upgrades**

When performing a rolling upgrade from Ubuntu 22.04 to Ubuntu 24.04:

* Disable VM High Availability (VM HA) before starting the upgrade.
* Disable Dynamic Resource Rebalancing (DRR) before starting the upgrade.
* Drain each host before upgrading it to Ubuntu 24.04.
* Complete all host upgrades before re-enabling VM HA and DRR This ensures workload stability during the upgrade process and prevents migration failures.

### Hypervisor Local Storage Pre-requisites

If you plan to use [Ephemeral Local Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-local-storage) for VM root disk for critical VMs in your environment, **we recommend using NFS shared storage** for your [Cluster Blueprint](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint#virtual-machine-storage-path) location. Having NFS shared storage for your VM storage path will make it easy to ensure that you have sufficient storage space for your VM root disks, and that you can expand this space quickly in the future, if required.

Regardless of whether you use NFS shared storage or local storage, you must ensure that there's sufficient local disk space at per hypervisor host level to store the virtual machine root disk files for the maximum number of VMs that will run on a hypervisor host at any given point.

The recommended minimum storage per hypervisor host:

* **250GB** of minimum disk size
* This storage will be used by operating systems, <code class="expression">space.vars.product\_name</code> service components, log files and **virtual machine root disk files** for VMs using ephemeral storage.

If you plan to use [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) or [Block Storage Volumes](/private-cloud-director/storage/volume) for VM root disk, then per hypervisor host local disk requirements will be lower:

* **100GB** of local disk space
* This storage will be used by operating system, <code class="expression">space.vars.product\_name</code> service components, and log files.

### Swap Configuration

We recommend that you configure swap memory on all hosts that you add to your <code class="expression">space.vars.product\_name</code> setup, to enhance system performance and to effectively manage memory-intensive VM workloads. Follow the steps in [Configure Swap](/private-cloud-director/getting-started/pre-requisites/configure-swap-ubuntu) to configure your swap memory and the swappiness value for your host.

### Networking Prerequisites

Read [Overview & Architecture](/private-cloud-director/virtualized-networking/networking-overview) for a detailed understanding of networking in <code class="expression">space.vars.product\_name</code>.

All hypervisor hosts should have a minimum of one network interface, and ideally four network interfaces to enable redundancy across network interface failure.

Use of bonded network interfaces is recommended to ensure high availability in case of a physical network interface failure.

A typical configuration would look like:

* bond0 mapped to two network adapters: eth0 and eth1
* bond1 mapped to two network adapters: eth2 and eth3

<code class="expression">space.vars.product\_name</code> allows you to designate separate network interfaces for the following types of traffic:

* Management network
* Image library I/O network
* VM console network
* Virtual network tunnels

#### Outbound Connectivity Requirements

You would need to configure outbound access on port 443 from your hosts for the below domain names to ensure they can be onboarded to the <code class="expression">space.vars.product\_name</code> management plane.

1. <code class="expression">space.vars.product\_name</code> management plane url is accessed over port 443.
2. For `pcdctl` CLI download on hosts, <https://pcdctl.s3.us-west-2.amazonaws.com/pcdctl-setup>
3. APT sources list for installing packages on the Ubuntu host using `pcdctl prep-node` :
   1. <http://security.ubuntu.com/ubuntu>
   2. <http://us.archive.ubuntu.com/ubuntu>
   3. <http://ubuntu-cloud.archive.canonical.com/ubuntu>
   4. <http://nova.clouds.archive.ubuntu.com/ubuntu>
   5. <https://wiki.ubuntu.com/OpenStack/CloudArchive>

#### Connectivity via http(s) proxy server

When outbound connectivity needs to be routed via a proxy server, ensure that the `/etc/environment` file has the http\_proxy, https\_proxy and no\_proxy variables defined as follows:

{% tabs %}
{% tab title="Bash" %}

```bash
https_proxy=http://pf9:squid@squid.pf9.io:3128
http_proxy=http://pf9:squid@squid.pf9.io:3128
no_proxy=127.0.0.1,localhost,..<IP address of other hosts>
```

{% endtab %}
{% endtabs %}

The format of a typical proxy server URL is `[<protocol>][<username>:<password>@]<host>:<port>` .

Also, the `no_proxy` variable should include IP addresses for which the traffic should not be routed via the proxy, e.g. other servers that would be set up as Hypervisor, Image Library or Storage roles.

Additionally, ensure that apt uses the proxy server to fetch packages. This can be configured by updating the `/etc/apt/apt.conf.d/proxy.conf` file with the following entries

{% tabs %}
{% tab title="Bash" %}

```bash
Acquire::http::Proxy "http://pf9:squid@squid.pf9.io:3128";
Acquire::https::Proxy "http://pf9:squid@squid.pf9.io:3128";
```

{% endtab %}
{% endtabs %}

### VNC Console Prerequisites

The VNC console service is added to all hypervisor hosts as part of configuration of hypervisor role. This allows you to access the VMs on the host from your web browser. The following prerequisites must be met for this feature to work

* Ensure that the port `6080` is open on each host.
* Ensure that you are on the same network as the host.
* To access the console using a public or floating IP address, or a domain name, configure the corresponding option in the cluster blueprint. Ensure that proper routing and DNS resolution are configured for the specified IP or domain. This setting applies to all hypervisor hosts in the region.

{% hint style="info" %}
**Console security model**

The underlying QEMU VNC sockets on each hypervisor are bound to the hypervisor's IP address and are reachable from the network if routing exists. VNC sessions are encrypted with TLS and authenticated with an automatically generated per-session password — no user configuration required. All authenticated browser console access flows through the noVNC service on port `6080`. TLS material for the libvirt VNC server is managed automatically under `/etc/pki/libvirt-vnc/`.
{% endhint %}

## Persistent / Block Storage Prerequisites

<code class="expression">space.vars.product\_name</code> supports a wide variety of enterprise storage solutions. Verify you have access to the administrative console of your storage solution and can look up the required configuration information from your admin console.

* Read more in the [Storage Overview](/private-cloud-director/storage/storage-overview) article about types of storage supported by <code class="expression">space.vars.product\_name</code>.
* For [block storage](/private-cloud-director/storage/volume), see the list of [supported block storage drivers](/private-cloud-director/storage/block-storage/supported-storage-drivers)
* <code class="expression">space.vars.product\_name</code> expects each hypervisor that connects to iSCSI storage must present **one unique iSCSI Qualified Name (IQN)**. Duplicate IQN can exist across hypervisor hosts when multiple hosts boot with an identical IQN, often because their OS image was cloned. Please refer to the [knowledge base article](https://platform9.com/docs/private-cloud-director/kb/how-to-rename-iscsi-initiator-names-to-the-existing-hypervisors-) to address duplicate IQNs.

### Temporary Volume Storage Requirements

Anytime you create a volume using an image, the directory `/opt/pf9/pf9-cindervolume-base/state/conversion` is used as a temporary staging area to store the image file while the conversion is happening. Therefore, for any hosts with persistent storage role assigned, this directory must have sufficient storage space to **support the cumulative size based on the maximum 'uncompressed' image size in your environment and the maximum number of concurrent new volume creations using images** you expect to happen in your setup. You can estimate the uncompressed size of a qcow2 image by doubling it's current disk size.

Also see here [latest compatibility matrix of Cinder storage drivers and devices](https://docs.openstack.org/cinder/latest/reference/support-matrix.html#driver-support-matrix).

## Image Library Prerequisites

The Image Library service manages virtual machine images in the <code class="expression">space.vars.product\_name</code> environment.

To enable its proper operation, the following prerequisites must be met on **all hosts with Image Library role assigned**:

* The Image Library service runs on port `9494`. **Ensure this port is open** to allow image operations, specifically image uploads, edit or delete operations from web browsers and CLI clients.
* Using shared storage for the image library service is required to create a highly available image library setup.
* Image uploads and deletes using the CLI require the use of the [Image Library Admin Endpoint](/private-cloud-director/images-and-image-library/image-library---images#image-library-admin-endpoint), and the client machine being used to upload the image must have network reachability to the IP address of the image library host that is serving as the admin endpoint.

### Image Library Temporary Storage Requirements

* Anytime a new image is being uploaded to your image library, the directory **`/var/lib/nginx`** is used as a temporary staging area to store the image file while the upload is happening. This means at any given point, **sufficient storage space** should be available under this directory to **support the cumulative size based on the maximum image size in your environment and the maximum number of concurrent image upload** operations you expect to happen to the image library.
  * We therefore recommend using shared storage (eg NFS) for this directory if you plan to use images with disk size greater than approximately 50GB.

### Browser Connectivity

The host that you've assigned the image library role (the image library host) must be accessible via a web browser from the machines your users will use to access <code class="expression">space.vars.product\_name</code> UI . This requirement is necessary for:

* Uploading images through the <code class="expression">space.vars.product\_name</code> UI.
* Verifying and accepting image library self-signed certificates.

### Self-Signed Certificates

The <code class="expression">space.vars.product\_name</code> image library service uses self-signed certificates today to secure the communication for image uploads. Since browsers and CLI tools only trust publicly verified certificates, users must manually accept the self-signed certificate before they can upload any images to the image library service.

To accept the self-signed certificate:

* Navigate to the image library endpoint in a browser.
  * Click Access & Security Menu -> API Access -> and look for glance-cluster.
* Accept the insecure certificate when prompted.

## Load Balancer As a Service (LBaaS) Prerequisites

Follow the [Load balancer as a service Prerequisites](/private-cloud-director/virtualized-networking/load-balancer#prerequisites) if you plan to use the <code class="expression">space.vars.product\_name</code> built-in LBaaS component for your applications.

## Kubernetes Prerequisites

Read [Kubernetes Pre-requisites](/private-cloud-director/kubernetes-clusters/k8s-pre-requisites) for setting up a Kubernetes cluster in <code class="expression">space.vars.product\_name</code>


# Configure Swap

This document describes the steps to configure swap memory and swapiness on a Ubuntu server.

## Configure Swap Memory

We recommend configuring your swap memory size to be **1.5x** **of your host physical RAM** for best performance.

There are two ways to configure swap memory for your host. You can either create a swap file in an existing partition, or allocate a dedicated disk partition for swap.

### Configure Swap Memory using Swapfile

Create a new Swapfile with a specific size using `fallocate`.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo fallocate -l 200G /swapfile.img
```

{% endtab %}
{% endtabs %}

Modify the Swapfile permissions to allow only the root user to read and write changes on the file.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo chmod 0600 /swapfile.img
```

{% endtab %}
{% endtabs %}

Format the file as swap using `mkswap`.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo mkswap /swapfile.img
```

{% endtab %}
{% endtabs %}

### Configure Swap Memory using Disk Partition

Convert your designated block storage partition to swap. In the example below, `/dev/vdb1` will be dedicated as your swap partition.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo mkswap /dev/vdb1
```

{% endtab %}
{% endtabs %}

### Enable Swap Memory

Once you have configured your swap memory, the next step is to enable swap using the `swapon` command.

If using swapfile:

{% tabs %}
{% tab title="Bash" %}

```bash
sudo swapon /swapfile.img
```

{% endtab %}
{% endtabs %}

If using dedicated partition:

{% tabs %}
{% tab title="Bash" %}

```bash
sudo swapon /dev/vdb1
```

{% endtab %}
{% endtabs %}

Finally, verify that the swap partition is active.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo swapon -s
```

{% endtab %}
{% endtabs %}

### [Make the Swap File Permanent](https://www.digitalocean.com/community/tutorials/how-to-add-swap-space-on-ubuntu-22-04#step-5-making-the-swap-file-permanent)

If using swap file, you will need to add the swap file to your `/etc/fstab` file, so that the swap settings are retained across server reboot.

Backup your fstab file first.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo cp /etc/fstab /etc/fstab.bak
```

{% endtab %}
{% endtabs %}

Now add the swap file to fstab.

{% tabs %}
{% tab title="Bash" %}

```bash
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
```

{% endtab %}
{% endtabs %}

## Configure Swappiness

### What is Swappiness

Swappiness is a linux kernel parameter that determines the system's tendency to move data from physical RAM to the swap space on a disk. It's a value between 0 and 100.

Default swappiness value for Ubuntu is 60.

### Recommended Swappiness Value

For <code class="expression">space.vars.product\_name</code> hypervisors, **we recommend setting the swappiness value to 10**, to minimize swapping and favor keeping processes in RAM to avoid performance degradation.

A swappiness value of 10 means swapping is only used as a last resort when RAM is nearly exhausted, prioritizing the performance and responsiveness of VMs

Setting swappiness higher increases the likelihood of swap usage, which is not ideal for hypervisors due to the significant latency swap imposes on VM workloads.

### Configure Swappiness

Use the following command to configure swappiness on your Ubuntu server.

Step 1 - Edit `/etc/sysctl.conf`

Step 2 - Add `vm.swappiness = 10`

Step 3 - Apply the changes by running `sudo sysctl -p`

Step 4 - Confirm that the changes are applied by running `cat /proc/sys/vm/swappiness`


# Hardware Virtualization Extension

This document describes the steps to validate that your server hardware that you plan to use with <code class="expression">space.vars.product\_name</code> has hardware and optionally I/O virtualization extensions enabled.

## Hardware Virtualization

<code class="expression">space.vars.product\_name</code> uses KVM hypervisor underneath which relies on hardware virtualization extensions (Intel VT-x or AMD-V) for full virtualization.

Specifically, <code class="expression">space.vars.product\_name</code> requires that your physical servers run with:

* an Intel processor with Intel VT-x and Intel 64 virtualization extensions enabled, or
* an AMD processor with AMD-V and AMD64 virtualization extensions enabled

Follow the steps below to check if your server hardware comes with hardware virtualization extensions and if they are enabled.

Run the following command to validate that CPU virtualization extensions are available on your server:

{% tabs %}
{% tab title="Bash" %}

```bash
$ grep -E 'svm|vmx' /proc/cpuinfo
```

{% endtab %}
{% endtabs %}

If the above command returns any output, that means your server hardware contains the hardware virtualization extensions.

In the output, look for a `vmx` entry, indicating an Intel processor with the Intel VT-x extension enabled, or an `svm` entry, indicating an AMD processor with the AMD-V extensions enabled.

In some cases, manufacturers may have disabled the virtualization extensions in the server BIOS settings. If the output of the above command is empty, or if full virtualization does not work, find your server and processor specific instructions on enabling the extensions in your server BIOS configuration.

## I/O Virtualization

Both Intel and AMD have hardware virtualization extensions available to allow virtual machines to have direct access to hardware devices. This is specifically important to enable passthrough access to GPU devices, when using GPUs with <code class="expression">space.vars.product\_name</code>.

If you plan to consume GPUs via GPU passthrough technology, then the following virtualization extension must be enabled:

* VT-d extensions enabled for an Intel processor, or
* AMD IOMMMU extensions enabled for an AMD processor

To check if the I/O virtualization extensions are enabled for your physical server, run the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
$ ls /sys/class/iommu/
```

{% endtab %}
{% endtabs %}

Examine the output.

* If the directory contains entries like `dmar0`, `dmar1` (for Intel systems), it indicates that IOMMU (and thus VT-d) is enabled
* If the directory contains entries like `amd-iommu-0` or similar, it suggests that the AMD IOMMU is recognized and potentially enabled by the kernel.

If the directory is empty, IOMMU might be disabled or not recognized.


# Hypervisor LVM Configuration

### Overview

PCD compute nodes (hypervisors) have LVM configured and running on the hypervisor node, must follow the guidelines below to properly configure LVM filters to prevent system hangs and performance issues.

This document outlines the required LVM filter configuration.

### Check if LVM Package is Installed

Run the following command to check if LVM package is installed on your system.

```bash
dpkg -l | grep lvm
ii  libllvm15:amd64                        1:15.0.7-0ubuntu0.22.04.3               amd64        Modular compiler and toolchain technologies, runtime library
ii  liblvm2cmd2.03:amd64                   2.03.11-2.1ubuntu4                      amd64        LVM2 command library
ii  lvm2                                   2.03.11-2.1ubuntu5                      amd64        Linux Logical Volume Manager
```

If LVM is installed, you will see output similar to above. In that case, follow the steps below to set LVM filters.

### Identify Physical Disk Devices On Your System

Run the command below to identify all the block devices mounted on this host. NOTE however that this command will also list any VM disks as well.

```bash
lsblk -o NAME,SIZE,TYPE,MOUNTPOINT,VENDOR,MODEL | grep disk
sda                        1.7T disk                    HPE      LOGICAL VOLUME
sdb                        1.0T disk                    HPE      LOGICAL VOLUME
sdg                         49G disk                    3PARdata VV
sdk                         49G disk                    3PARdata VV
sdl                         49G disk                    3PARdata VV
sdm                         49G disk                    3PARdata VV
```

Identify from the output above the **subset of entries that are the physical disks for your hypervisor host**.

In this example, you would only pick `sda` and `sdb`. The other four devices are volumes attached to VMs running on this host.

### Add Or Enable LVM Filters For All Disks

Edit `/etc/lvm/lvm.conf` and locate the devices section.

Then edit the `filter` and `global_filter` values and add the **entries for the physical disks for your hypervisor host** that you identified above.

Note that this is a regular expression, where 'a' stands for accepting the path and 'r' stands for rejecting the path in the syntax below.

This example

```bash
devices {
    filter = [ "a|^/dev/sda|", "a|^/dev/sdb|", "r|.*|" ]
    global_filter = [ "a|^/dev/sda|", "a|^/dev/sdb|", "r|.*|" ]
}
```

### Validate that the LVM Filters are Active

```bash
sudo lvm dumpconfig devices | grep -E "filter|global_filter"
```

### Symptoms When Not Configured

If LVM is installed on your hypervisor but the LVM filters are not configured properly, here are some of the side effects and symptoms that you may observe on your <code class="expression">space.vars.product\_name</code> hosts:

* QEMU processes stuck in uninterruptible sleep (D state)
* VM operations hang or timeout
* Slow system responsiveness on the host
* High I/O wait times during LVM operations
* VMs fail to start or become unresponsive


# Getting Started

This document provides steps to get started with your <code class="expression">space.vars.product\_name</code> setup.

If you haven't yet, make sure to follow the [Pre Requisites](/private-cloud-director/getting-started/pre-requisites) to get ready for install.

## Step 1 - Setup your Private Cloud Director Instance

The commercial version of <code class="expression">space.vars.product\_name</code> offers two deployment options: **SaaS or Self Hosted**. Or you can use the free community edition of <code class="expression">space.vars.product\_name</code> for your home lab / test setup.

### 1. SaaS Hosted

With SaaS Hosted model, we setup and host <code class="expression">space.vars.product\_name</code> in the cloud for you and you consume it as a SaaS service. **Your infrastructure still remains in your data center.** The SaaS hosted version of <code class="expression">space.vars.product\_name</code> communicates remotely with your infrastructure to get your private cloud setup and running. This is generally the best option for organizations that prefer the ease of use of SaaS and don't want to manage the technical complexity of hosting the product themselves.

### 2. Self Hosted

With Self Hosted model, you can host <code class="expression">space.vars.product\_name</code> on your on hardware in your data center or a co-location hosted data center. This is generally the best option for organizations with extra compliance and security requirements where SaaS hosting is not a viable option.

Follow [Self Hosted Install](/private-cloud-director/getting-started/self-hosted/self-hosted-install) for steps to install the self hosted management plane.

### 3. Community Edition

<code class="expression">space.vars.product\_name</code> also offers a community edition that is completely free to download and use.

Follow [Community Edition](/private-cloud-director/getting-started/getting-started-with-community-edition) guide for steps to install <code class="expression">space.vars.product\_name</code> Community Edition.

### 4. Platform9 OS (Beta)

Platform9 OS bundles Rocky Linux by CIQ (RLC) and <code class="expression">space.vars.product\_name</code> into a single ISO-based installation. It is designed for organizations that want a fully integrated, appliance-style deployment where the operating system and <code class="expression">space.vars.product\_acronym</code> are installed together in one workflow — no separate OS provisioning required.

{% hint style="warning" %}
Platform9 OS is currently in closed beta. Access is by invitation only. Contact your Platform9 account manager to request access.
{% endhint %}

Follow [Platform9 OS Beta Installation](/private-cloud-director/getting-started/platform9-os-beta-install) for steps to install using the ISO-based installer.

## Step 2 - Create Cluster Blueprint & Virtualized Cluster

Once you have a <code class="expression">space.vars.product\_name</code> instance setup, next step is to start onboarding your hypervisors. Follow the steps below to create a [Cluster Blueprint](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint) and start onboarding hypervisors.

1. Create a virtualized cluster blueprint. Follow [Virtualized Cluster Blueprint](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint) article for more information on what cluster blueprints are and how to create one.
2. Once the cluster blueprint is created, you are now ready to create one or more virtualized clusters.
3. Follow [Create a Virtualized Cluster](/private-cloud-director/virtualized-clusters/virtualized-cluster#create-a-virtualized-cluster) for steps to create a cluster.
4. Your first virtualized cluster is now created.

## Step 3 - Add Hosts to Cluster

1. The next step is to start onboarding your physical servers or hosts and adding them to the cluster.
   1. For SaaS hosted management plane - follow [Add a Host (SaaS Deployment)](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#add-a-host-saas-deployment)
   2. For Self Hosted or Community Edition - follow [Add a Host (Self-Hosted & Community Edition Deployment)](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#community-edition-deployment)
2. Now you will have a number of hosts showing up as unauthorized in your <code class="expression">space.vars.product\_name</code> UI.
3. Follow [Authorize Host And Assign Roles](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#authorize-host-and-assign-roles) to authorize one or more roles for the hosts and assign them to the cluster.

## Step 4 - Create Images and Networks

1. Follow [Image Library Images](/private-cloud-director/images-and-image-library/image-library---images) to create your first image in your image library.
2. Create at least one [Virtual Network](/private-cloud-director/virtualized-networking/networks-and-ports) or [Physical Network](/private-cloud-director/virtualized-networking/physical-network).

## Step 5 - Create Virtual Machines

1. Read [Virtual Machines](/private-cloud-director/virtualized-clusters/virtualmachine) for steps to provision your first virtual machine on your cluster.

## Step 6 - Create Kubernetes Cluster

1. Read [Getting Started with Kubernetes](/private-cloud-director/kubernetes-clusters/getting-started-with-kubernetes-in-pcd) to create your first Kubernetes cluster using <code class="expression">space.vars.product\_name</code>.


# Community Edition

Understand what Community Edition is, how it’s different, and where to start.

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

{% hint style="info" %}
**New release!**

For details on the latest release of <code class="expression">space.vars.product\_name</code> Community Edition, see the [Release Notes](/release-notes/april-2026-release).
{% endhint %}

## Community Edition At a Glance

Community Edition (CE) is the community-supported, free, full-featured version of <code class="expression">space.vars.product\_name</code>.

You get the same core virtualization experience as the commercial product.

CE is different in how it’s deployed and supported:

* **Single-host management plane.** CE runs the management plane on one machine. Note that this is only the limitation on how the management plane is provisioned. There is no limit on the number of hypervisors you can add to the CE management plane.
* **You bring the hypervisors.** Workloads always run on separate hypervisor hosts.
* **Community support.** Use <https://www.reddit.com/r/platform9/> and these docs.

{% hint style="info" %}
If you want a Platform9-managed control plane, use SaaS Hosted. If you need a customer-managed highly available control plane, use [Self Hosted](/private-cloud-director/getting-started/self-hosted).

See [Getting Started](/private-cloud-director/getting-started/getting-started) for a quick comparison.
{% endhint %}

***

## How Community Edition Works

CE is two system types working together:

* **CE host (control plane)**
* **Hypervisor hosts (run VMs)**

### Community Edition Host (Control Plane)

The **CE host** is where you manage the platform.

It:

* Serves the web UI
* Exposes REST APIs
* Monitors hosts and VMs
* Orchestrates lifecycle operations (create, resize, migrate, HA, and more)
* Hosts Grafana dashboards

It does **not** run your VMs.

### Hypervisor Hosts (Compute)

**Hypervisor hosts** run the virtual machines.

Each hypervisor runs a **Platform9 host agent**. The agent:

* Connects the host to the CE control plane
* Executes lifecycle actions
* Sends health and metrics back to CE

### Shared Storage (Optional, But Recommended)

Shared storage (NFS, iSCSI, and other supported backends) enables:

* **Persistent volumes** for VM disks
* **Live migration** between hypervisors
* **VM High Availability (VM HA)** that can restart VMs on another host

If you use **only local storage** on hypervisors:

* Live migration requires VM shutdown.
* VM disks stay tied to the original host.
* VM HA can’t protect against host loss for those VMs.

***

## Get Started

1. Validate requirements in [Prerequisites](/private-cloud-director/getting-started/getting-started-with-community-edition/prerequisites).
2. Run the standard install in [Install](/private-cloud-director/getting-started/getting-started-with-community-edition/install).
3. If you need custom CIDRs, domains, or other changes, read [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation) **before** installing.
4. If anything fails, use [Common Issues](/private-cloud-director/getting-started/getting-started-with-community-edition/common-issues).

{% hint style="info" %}
Want a walkthrough instead of reference docs? Use the tutorial: [Beginner's Guide to Deploying PCD Community Edition](/private-cloud-director/tutorials/beginners-guide-to-deploying-pcd-community-edition).
{% endhint %}

## Operations

* Backup strategy and procedures: [Backup and Restore CE](/private-cloud-director/getting-started/getting-started-with-community-edition/backup-and-restore-ce)


# Prerequisites

Host, network, and storage requirements for installing Community Edition.

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

## Quick Start Checklist

Use this if you want the minimum viable lab setup.

* [ ] Community Edition (CE) Host: [Ubuntu 24.04](https://cloud-images.ubuntu.com/releases/noble/release/) VM with 8 CPUs, 32GB RAM, 100+ GB free space
* [ ] Hypervisor host: [Ubuntu 24.04](https://cloud-images.ubuntu.com/releases/noble/release/) VM with 8+ CPUs, 16+ GB RAM, 100+ GB free space
* [ ] Both the CE host and hypervisor host(s) can reach the internet
* [ ] You can access both CE and the hypervisor(s) via SSH (port 22)
* [ ] You can access the CE host via HTTPS (port 443)
* [ ] Install Community Edition on the CE host (see [Install](/private-cloud-director/getting-started/getting-started-with-community-edition/install))

```bash
curl -sfL https://go.pcd.run | bash
```

## Detailed Prerequisites

Community Edition (CE) uses two system types:

* **CE host**: runs the control plane and web UI.
* **Hypervisor hosts**: run your VMs.

### Community Edition Host (CE Host)

This host runs the <code class="expression">space.vars.product\_name</code> control plane and web interface.

#### Operating system

* **Supported:** [Ubuntu 24.04 AMD64 cloud image](https://cloud-images.ubuntu.com/releases/noble/release/)
* **Not supported:** Minimal cloud image (missing required packages, such as `crontab`)
* **Other distributions:** Not supported by the installer

#### Sizing

* **CPU:** 8 logical CPUs, 12+ for Kubernetes workloads
* **RAM:** 28GB minimum, 32GB recommended
* **Disk:** 100GB minimum free space, SSD recommended
* **CPU architecture:** Intel Nehalem / AMD Bulldozer or newer (x86-64-v2)

#### Kubernetes management plane (optional)

If you enable the Kubernetes management plane during installation, the CE host requires:

| Resource | Minimum         | Recommended     |
| -------- | --------------- | --------------- |
| **CPU**  | 12 logical CPUs | 16 logical CPUs |
| **RAM**  | 32 GB           | 48 GB           |

The installer verifies the minimum requirements before proceeding. Disable the Kubernetes management plane to use the base CE host requirements.

#### Network

* **Internet access**
* **IPv6:** Must be enabled, but doesn't need an address
* **Firewalld:** Must be stopped and disabled (if installed)

```bash
# Stop the firewalld service
sudo systemctl stop firewalld
# Prevent it from starting on boot
sudo systemctl disable firewalld
```

<details open>

<summary>Ports and outbound access</summary>

**Ports open locally (CE host):**

* 443, 2379, 2380, 3306, 4194, 5395, 5672, 5673, 6264, 8023, 8158, 8285, 8558, 9080, 10250, 10255

**Outbound access:**

* TCP 53, 443
* UDP 53, 123

</details>

### Hypervisor Hosts

These hosts run your actual VMs.

#### Operating system

* **Supported:** [Ubuntu 24.04 AMD64](https://cloud-images.ubuntu.com/releases/noble/release/) cloud image
* **Not supported:** Minimal cloud image (missing required packages, such as `crontab`)
* **Other distributions:** Not supported by the installer

#### Sizing

* **CPU:** 8 logical minimum, 16+ recommended
* **RAM:** 16GB minimum, 32GB+ recommended
* **Disk:** 100GB minimum free space, 200GB+ recommended (shared storage is better)

#### Nested virtualization (if hypervisors are VMs)

If your hypervisor hosts run as VMs, nested virtualization must be enabled.

**VMware vSphere**

* For hypervisor hosts running as VMs on vSphere:
  * Enable hardware virtualization on the vCPUs before installing Ubuntu
  * Allocate at least 115GB thin-provisioned disk (to allow for 100+ GB free space after OS)
  * Accept promiscuous mode, MAC address changes, and forged transmits on the vSwitch (not at the port group!) to allow nested VMs running on the hypervisor VM to access the external network

**Proxmox**

* Enable CPU type "host" or "kvm64"
* Enable nested virtualization

{% hint style="info" %}
If hypervisor hosts are running as VMs, nested virtualization must be enabled and visible to the hypervisor host operating system.
{% endhint %}

#### Nested virtualization check

```bash
# Run this on the hypervisor host
egrep "svm|vmx" /proc/cpuinfo

# Should show output with vmx (Intel) or svm (AMD) flags
# No output = nested virtualization not enabled
```

### Storage Backend

You need shared storage for persistent VM disks and live migration. Ephemeral disks work without shared storage and are stored locally on each hypervisor host.

Options:

1. **NFS (Easiest for Testing)**
   * Minimum: 100GB+ available on NFS mount point
   * Best for: Lab environments, proof-of-concept
   * Performance: Good for testing, adequate for light workloads
2. **External Storage Array**
   * Minimum: Depends on your needs
   * Best for: dev/test environments, extended proof-of-concept
   * Performance: Excellent
3. **Local Storage**
   * Uses hypervisor local disk
   * Best for: Ephemeral VMs, temporary workloads
   * Limitation: VM live migration and VM HA are not supported

See [Storage Overview](/private-cloud-director/storage/storage-overview) for detailed configuration.

### Network Requirements

#### Default Kubernetes CIDRs

* Services: `10.21.0.0/16`
* Pods: `10.20.0.0/16`

{% hint style="info" %}
**Check for conflicts:** If your network uses these ranges, see [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation) to change them BEFORE installing.
{% endhint %}

#### Required external URLs

The CE installer must reach these endpoints. If you’re behind a firewall, whitelist them.

* `go.pcd.run` - Installer script
* `quay.io`, `registry.k8s.io` - Container images
* `github.com` - Software downloads
* `gcr.io` - Required if the Kubernetes management plane is enabled
* [Complete list](/private-cloud-director/getting-started/getting-started-with-community-edition/common-issues#required-ports-and-external-urls) of required ports and URLs

{% hint style="warning" %}
**Proxy environments:** The Kubernetes management plane can be installed in proxy environments, but provisioning Kubernetes clusters through it is not supported.
{% endhint %}


# Install

Default installation steps for Community Edition, plus first-login and basic operations.

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

{% hint style="info" %}
For a guided walkthrough, use [Beginner’s Guide to Deploying PCD Community Edition](/private-cloud-director/tutorials/beginners-guide-to-deploying-pcd-community-edition).
{% endhint %}

## Install Community Edition

This installs the Community Edition (CE) control plane on a single **CE host**. You then onboard one or more **hypervisor hosts** to run VMs.

{% hint style="info" %}
Read [Prerequisites](/private-cloud-director/getting-started/getting-started-with-community-edition/prerequisites) first. See [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation) if you need custom CIDRs or domains.
{% endhint %}

### Outcome

After install, you can:

* Log in to the web UI
* Onboard hypervisor hosts
* Create networks, images, and VMs
* View metrics in the web UI or Grafana

### Timeline

Typical runtime is \~45 minutes. Slow systems or slow registries can push it to \~90 minutes.

* 5–10 min: system prep and prerequisite checks
* 10–15 min: K3s cluster creation
* 20–30 min: <code class="expression">space.vars.product\_acronym</code> services deployment to the K3s cluster
* \~5 min: verification and credential output

### Deploy the CE host

Run the installer on the CE host as a user with `sudo` access.

```bash
curl -sfL https://go.pcd.run | bash
```

{% hint style="info" %}
By default in Ubuntu, all users in the `sudo` group have the ability to sudo. Group membership can be validated with the command `groups <username>` and sudo permissions can be verified with `sudo -l -U <username>` .
{% endhint %}

The installer will:

* Ask you to accept the [Community Edition EULA](https://platform9.com/ce-eula)
* Ask whether to enable the Kubernetes management plane (optional - see [Kubernetes management plane prompt](#kubernetes-management-plane-prompt))
* Run prerequisite checks
* Deploy the CE management plane
* Print the user interface web address and provide admin credentials

If prerequisite checks fail, see [Common Issues](/private-cloud-director/getting-started/getting-started-with-community-edition/common-issues). If installation fails, the installer can optionally upload a support bundle.

#### Kubernetes management plane prompt

The installer asks whether to enable the Kubernetes management plane, which lets you provision managed Kubernetes clusters. This is opt-in — the default is No.

Enabling it:

* Adds approximately 10 minutes to install time
* Requires a larger CE host - 12 CPUs and 32 GB RAM minimum (see [Prerequisites](/private-cloud-director/getting-started/getting-started-with-community-edition/prerequisites#kubernetes-management-plane-optional))

To skip the prompt for automation, pre-export the variable before running the installer:

```bash
# Enable Kubernetes management plane
export ENABLE_K8S=true

# Disable Kubernetes management plane
export ENABLE_K8S=false
```

<details>

<summary>Installation example</summary>

```bash
ubuntu@ce-host:~$ curl -sfL https://go.pcd.run | bash
Private Cloud Director Community Edition Deployment Started...

By continuing with the installation, you agree to the terms and conditions of the
Private Cloud Director Community Edition EULA.

Please review the EULA at: https://platform9.com/ce-eula

Do you accept the terms of the EULA? [Y/N]: y

Private Cloud Director can install a Kubernetes management plane
that lets you provision managed Kubernetes clusters.
This is optional and adds ~10 minutes to install time and requires additional
CPU / memory (see [Prerequisites](prerequisites.md)).

Enable Kubernetes management plane? [y/N]: y

Finding latest version...  Done
Downloading artifacts...  Done
Configuring system settings...  Done
Installing artifacts and dependencies...  Done
Configuring Docker Mirrors...  Done
 SUCCESS  Configuration completed
 INFO  Verifying system requirements...
 ✓  Architecture
 ✓  Disk Space
 ✓  Memory
 ✓  CPU Count
 ✓  OS Version
 ✓  Swap Disabled
 ✓  IPv6 Support
 ✓  Kernel and VM Panic Settings
 ✓  Port Connectivity
 ✓  Firewalld Service
 ✓  Default Route Weights
 ✓  Basic system services
Completed Pre-Requisite Checks on local node
 SUCCESS  Cluster created successfully
 INFO  Starting PCD management plane
 SUCCESS  Certificates generated
 SUCCESS  Base infrastructure setup complete
 SUCCESS  Installed Kubernetes management plane
 SUCCESS  pcd-virt deployment now complete
 SUCCESS  Final touches...
Private Cloud Director (Community Edition) deployment complete!
How would you rate your Private Cloud Director (Community Edition) installation experience?
(1-5, press Enter to skip): 5

------------- deployment details ---------------
fqdn:                pcd.pf9.io
region:              pcd
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2026.4-116
-------- region service status ----------
desired services:     56
ready services:       56

Login Details:
URL: https://pcd.pf9.io
email:     admin@airctl.localnet
password:  wVXJbocBrZaqcjtS

See documentation at https://pcd.run for next steps.

Join the support community at https://reddit.com/r/platform9
ubuntu@ce-host:~$
```

</details>

#### Enable Kubernetes management plane on an existing CE installation

To enable the Kubernetes management plane on an existing CE deployment:

{% hint style="info" %}
The base CE installation must be fully deployed and healthy before running this command.
{% endhint %}

```bash
sudo /opt/pf9/airctl/airctl start --config /opt/pf9/airctl/conf/airctl-config.yaml --k8s-only
```

This takes approximately 10 minutes.

To verify the installation succeeded:

```bash
kubectl get pods -n kaapi
```

All pods should show `Running`.

{% hint style="warning" %}
The CE host must meet the minimum requirements for the Kubernetes management plane: 12 CPUs and 32 GB RAM. See [Prerequisites](/private-cloud-director/getting-started/getting-started-with-community-edition/prerequisites#kubernetes-management-plane-optional).
{% endhint %}

### Access the web UI

For default installations, the web user interface can be accessed at `https://pcd.pf9.io`. That name must resolve from your browser.

{% hint style="info" %}
If you've [customized the Community Edition FQDN](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation#deployment-fqdn-fully-qualified-domain-name), then replace `pcd.pf9.io` with that value.
{% endhint %}

#### Option 1: Local hosts file (single machine)

Use this for quick testing, as only your computer will be able to resolve the URL to the user interface.

**Linux/MacOS**

```bash
# Replace <CE_HOST_IP> with your CE host IP
echo "<CE_HOST_IP> pcd.pf9.io" | sudo tee -a /etc/hosts

# Example
echo "192.168.1.100 pcd.pf9.io" | sudo tee -a /etc/hosts
```

**Windows**

1. Open Notepad as Administrator. From the Windows Start menu, right-click the Notepad icon and choose "Run as Administrator".
2. Open `C:\Windows\System32\drivers\etc\hosts`.
3. Add the following, but be sure to replace the example IP. Save and close the file afterwards.

```
192.168.1.100 pcd.pf9.io
```

#### Option 2: Real DNS (team access)

Create A records in your DNS server:

```
pcd.pf9.io            -> <CE_HOST_IP>
```

### Log in

Open the user interface URL in a browser (the default is [pcd.pf9.io](https://pcd.pf9.io/)). Accept the self-signed certificate on first visit.

* Select **Use local credentials**.
* Use the `admin@airctl.localnet` credentials printed by the installer.
* Keep **Domain** as `default`.
* Do not check "I have an MFA token".

{% hint style="info" %}
SSO & MFA are unavailable until configured. Errors with an SSO or MFA login attempt will fail until they are configured.
{% endhint %}

<figure><img src="/files/sUGZrOckQIKwWkm8naA8" alt="Private Cloud Director Community Edition local credentials login screen"><figcaption><p>Private Cloud Director Community Edition local credentials login screen</p></figcaption></figure>

### Admin credentials and Grafana access

Do not rename the default admin user, `admin@airctl.localnet`. Do not change its password.

This account is used internally by <code class="expression">space.vars.product\_name</code> services.

Create separate users for administrator or self-service access.

{% hint style="warning" %}
The default admin account is also used for Grafana authentication. Changing the default admin email or password can break Grafana access.
{% endhint %}

### Next steps

1. See the Getting Started guide at [Getting Started](/private-cloud-director/getting-started/getting-started#step-2-create-cluster-blueprint-and-virtualized-cluster)
2. Follow the rest of the Getting Started guide to onboard hypervisor hosts, create networks, upload images, and create virtual machines.

***

## Operations and troubleshooting

<details>

<summary>Validate the installation</summary>

Your deployment is healthy if the region is marked as `Ready`. Ready services should match desired services.

```bash
/opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

Example output:

```bash
root@pcd-ce:~# /opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
------------- deployment details ---------------
fqdn:                pcd.pf9.io
region:              pcd
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2026.4-108
-------- region service status ----------
desired services:     54
ready services:       54
```

</details>

<details>

<summary>Retrieve admin credentials</summary>

```bash
/opt/pf9/airctl/airctl get-creds --config /opt/pf9/airctl/conf/airctl-config.yaml
```

</details>

<details>

<summary>Logs and service names</summary>

**Hypervisor host logs**

* `/var/log`
* `/var/log/pf9`

**CE host logs**

* `/var/log/pf9/fluentbit/ddu.log` (JSON; includes Kubernetes pod logs)

You can also use `kubectl logs` to fetch pod logs.

**Hypervisor host services**

Service names start with `pf9`. Examples:

* `pf9-hostagent`
* `pf9-imagelibrary`
* `pf9-ostackhost`

</details>

<details>

<summary>Grafana credentials</summary>

Grafana uses the same admin credentials printed at install time.

* Username: `admin@airctl.localnet`
* Password: retrieve from the CE host:

```bash
/opt/pf9/airctl/airctl get-creds --config /opt/pf9/airctl/conf/airctl-config.yaml
```

Default Grafana URL:

* `https://pcd.pf9.io/grafana/login`

</details>

<details>

<summary>Reset Grafana admin password (advanced)</summary>

{% hint style="danger" %}
Resetting Grafana independently is not recommended. Keep Grafana and UI admin passwords in sync.
{% endhint %}

The namespace is `pcd` by default. If you customized the FQDN, use the FQDN shortname for the namespace. Example: If the FQDN is *ce.acmeretail.com*, assign `ce` to the `NS` variable in the below example.

```bash
NS=pcd
kubectl exec -it deploy/grafana -n $NS -- grafana cli admin reset-admin-password <strong-password>
```

</details>

<details>

<summary>Uninstall Community Edition</summary>

```bash
/opt/pf9/airctl/airctl delete-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml
```

</details>

<details>

<summary>Expand an Ubuntu logical volume (LVM)</summary>

If your root filesystem is smaller than the physical volume, expand it.

1. Find the root device:

```bash
df -h /
```

2. Resize the LV (example name shown):

```bash
sudo lvresize -l +100%FREE /dev/mapper/ubuntu--vg-ubuntu--lv
```

3. Resize the filesystem:

```bash
sudo resize2fs /dev/mapper/ubuntu--vg-ubuntu--lv
```

</details>

### Troubleshooting

Start with [Common Issues](/private-cloud-director/getting-started/getting-started-with-community-edition/common-issues).

#### Optional learning resources

Platform9 YouTube:

* Channel: <https://www.youtube.com/@Platform9>
* Playlist: [Private Cloud Director playlist](https://www.youtube.com/playlist?list=PLUqDmxY3RncUmegG6dv8XxjSr0g1THpLq)


# Custom Installation

The purpose of this page is to document the ways that Community Edition can be customized.

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

There are a number of environment variables that can be defined prior to deploying the Community Edition host that will allow you to customize the default IP address used for the install, the URL for the Community Edition user interface, the use of install telemetry, and more.

{% hint style="info" %}
**Info**

All environment variables exported via the command line will not persist beyond the current terminal session, and only apply to the user that exports the variables. If a variable needs to be exported, run the install from the same terminal session. All environment variables will be exported using the following command structure: `export <variable_name>=<value>`
{% endhint %}

### HTTP & HTTPS proxy support

If you would like to install Community Edition behind a proxy, you can use the `HTTPS_PROXY` and `HTTP_PROXY` environment variables. Additionally, the `NO_PROXY` environment variable can be used to add addresses that should not be proxied to the default list of addresses.

```
export HTTPS_PROXY=https://proxy:1234
export HTTP_PROXY=http://proxy:5678
export NO_PROXY=<IP address or FQDN>
```

### Change IP address

The install script will choose the IP address of the interface set as the default route using the `ip route` command. Should you choose to use a different default IP address, you can export the `IP_ADDRESS` environment variable.

```bash
export IP_ADDRESS=192.168.1.150
```

### Deployment FQDN (fully qualified domain name)

The FQDN for the deployment (`DU_FQDN`) sets the URL for the Community Edition user interface. It may be useful to customize the fully qualified domain name based on an existing DNS domain, for example.

Defaults:

* `DU_FQDN="pcd.pf9.io"`
* Resulting fully qualified domain name: `pcd.pf9.io`

```bash
export DU_FQDN="pf9.lan"
# This will result in pf9-test.lan as the FQDN for the user interface
```

### Installing with a signed SSL certificate

CE can be installed with a signed SSL certificate instead of the default self-signed certificate. You may do this by setting both `USER_CERT_PATH` and `USER_KEY_PATH` environment variables with the fully qualified paths to the certificate and key files.

It is required to generate the certificates with the appropriate wildcard SANs and Key Usage:

* \*.pf9.localnet
* \*.domain.net

The first, \*.pf9.localnet is required for internal usage. The second depends on the shortname/FQDN used. For example if the `DU_FQDN` is "air99.platform9.net", then ensure the certificate has SANs for \*.platform9.net.

In addition, ensure the following Key Usage extensions are enabled:

```none
X509v3 extensions:
  X509v3 Key Usage: critical
  	Digital Signature, Key Encipherment  
  X509v3 Extended Key Usage:
  	TLS Web Server Authentication, TLS Web Client Authentication
```

Defaults:

* These variables are not set by default, as installing a self-signed certificate is the default behavior.

```bash
export USER_CERT_PATH=/path/to/cert
export USER_KEY_PATH=/path/to/key
```

### Community Edition deployment networking

Community Edition is installed as a Kubernetes deployment, and uses internal networking for communication between pods (`POD_CIDR` ) and for exposing applications as service IPs (`SERVICE_CIDR` ). If these IP address ranges conflict with your organization's networking ranges, they can be changed as needed. If there is no conflict, there's no need to change these default ranges.

Defaults:

* `SERVICE_CIDR="10.21.0.0/16"`
* `POD_CIDR="10.20.0.0/16"`

```bash
export SERVICE_CIDR="10.21.0.0/16"
export POD_CIDR="10.20.0.0/16"
```

### Installation telemetry

The Community Edition installation collects anonymous telemetry during the installation process allowing Platform9 to improve the installation process. Should you wish to opt out, the `TELEMETRY` environment variable can be set to `false`.

Default:

* `TELEMETRY=true`

```bash
export TELEMETRY=false
```

### Skip install pre-requisites check

The Community Edition installer runs a series of pre-requisite checks before installing, such as validating the amount of CPU & memory available. This behavior can be changed by setting the `SKIP_PRECHECKS` environment variable to `true` .

Default:

* `SKIP_PRECHECKS=false`

{% hint style="warning" %}
**Warning**

Forcing an install with less than the required amount of CPU & memory *will* cause it to fail!
{% endhint %}

```bash
export SKIP_PRECHECKS=true
```

### Skip post-install experience poll

As part of our ongoing effort to make the install experience as smooth & painless as possible, we've added a post-install poll to capture how you rated the installation experience. The `SKIP_RATING_PROMPT` variable can be set to `true` to disable this behavior.

Default:

* `SKIP_RATING_PROMPT=false`

```bash
export SKIP_RATING_PROMPT=true
```

### Automatically accept EULA

The `ACCEPT_EULA` environment variable can set to `true` before installing Community Edition, providing you have read and agreed to the EULA in advance. This environment variable can useful in the case of automating Community Edition installs.

Default:

* `ACCEPT_EULA=false`

```bash
export ACCEPT_EULA=true
```

### Automatically clean-up previous installation

The `SHOULD_CLEANUP` environment variable can be set to `true` to bypass the prompt to keep or delete a prior installation. The default is `false` .

Default:

* `SHOULD_CLEANUP=false`

```bash
export SHOULD_CLEANUP=true
```

### Extending default deployment timeout

If you are experiencing install failures due to timeouts, they may be resolved by setting the install script's default timeout period to longer than 600 seconds using the `SVC_DEPLOYMENT_TIMEOUT` environment variable. The normal installation process uses an internal variable, which can be changed with this environment variable.

Default:

* The default internal timeout period is `600` seconds.

```bash
export SVC_DEPLOYMENT_TIMEOUT=1200
```

### Kubernetes workload support

Kubernetes management plane support is optional in CE, enabled via the interactive install prompt or by pre-exporting `ENABLE_K8S` as true. The default is `false`.

```bash
export ENABLE_K8S=true
```

### Protected environment variables

The following environment variables are critical to the installation process and must not be changed, as this will cause installation failures: `S3_BUCKET` and `S3_USER_AGENT`.


# Common Issues

The purpose of this page is to provide solutions to common issues during Community Edition (CE) installation & operation.

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

## Resolving failed prerequisites checks

### Ports should not be in use

The CE installer expects to be installed in a dedicated system with the following ports available: 443, 2379, 2380, 3306, 4194, 5395, 5672, 5673, 6264, 8023, 8158, 8285, 8558, 9080, 10250, 10255. Typically, this error will be encountered after a previous installation failure. To completely clean up an installation failure, use `airctl delete-cluster` and then retry the installation.

```bash
# Deletes the CE installation, including k3s
sudo airctl delete-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml
# deletes logs & downloaded artifacts in the user's home directory
rm -r "$HOME/airctl-logs"
rm -r "$HOME/pcd-ce"
# runs the install script from the internet
curl -sfL https://go.pcd.run | bash
```

### Insufficient CPU or RAM

The CE install script checks the number of logical CPUs available at system start. If the install fails this check, run the installation again after provisioning more logical CPUs to the system (if running as a virtual machine) or on physical hardware with more logical CPUs available.

The install script calculates the amount of usable memory using the `MemTotal` line in `/proc/meminfo` , and checks for an amount less than 28GB. Additional memory must be made available to the CE system.

### Multiple default routes with the same metric

The CE install script checks for more than one default IPv4 route with the same metric. If more than one route with the same metric is found, this check will fail, as the system will not know which default gateway to use. To resolve this error, either change the metric on the additional default routes, or delete any additional default routes with the same metric if they aren't needed.

```bash
# list all IPv4 routes
ip -4 route show
# Change the metric on the additional default route(s)
sudo ip route change default via <gateway-ip> metric <new-metric>
# Delete the additional default route(s) if they are not needed
sudo ip route del default via <gateway-ip>
```

### Unsupported OS or version

The CE install script checks the operating system name and version. If this check fails, please try installing using a supported operating system and version as described in the CE [Pre-requisites](/private-cloud-director/getting-started/getting-started-with-community-edition/prerequisites#detailed-prerequisites).

### Kernel panic or vm panic on oom

The CE install script checks to make sure kernel & oom (out of memory) panics are configured correctly. `/proc/sys/kernel/panic` should be set to `10` and `/proc/sys/vm/panic_on_oom` should be set to `0` .

### Failed to check firewalld status or service not found

The CE install script checks `systemctl` to see if `firewalld.service` is inactive or not running, and will fail if the service is set to active. To resolve this, use the following commands.

```bash
# Stop the firewalld service
sudo systemctl stop firewalld
# Prevent it from starting on boot
sudo systemctl disable firewalld
```

### IPv6 is disabled

The CE install script checks `/proc/sys/net/ipv6/conf/all/disable_ipv6` to see if it is set to `0` . IPv6 should be enabled by default, but if it has been explicitly disabled, please enable IPv6 and reboot the CE system before trying the installation again.

***

## Resolving installation failures

### Calico CNI

CE installation can fail if it encounters issues with Calico CNI (container networking interface). Calico sets up networking within CE's Kubernetes pods and provides DNS resolution. Typically, these issues are either related to slowness when retrieving the container images from a remote image repository causing the install script to timeout while waiting, or are related to a previous installation not being fully cleaned up before an installation is attempted again. These issues can usually be resolved by fully cleaning up the installation attempt and trying again.

You may also consider setting the install's default timeout to an amount higher than 600 seconds. See the [Extending default deployment timeout](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation#extending-default-deployment-timeout) section of the [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation) page.

```bash
# Deletes the CE installation, including k3s
sudo airctl delete-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml
# deletes logs & downloaded artifacts in the user's home directory
rm -r "$HOME/airctl-logs"
rm -r "$HOME/pcd-ce"
# runs the install script from the internet
curl -sfL https://go.pcd.run | bash
```

### Failed to apply logrotation on node - exit status 127

The CE installation will fail if it is unable to create a cronjob to automatically rotate log files, and will return exit status 127 if `cron` is not installed on the system. This is a default package on the Ubuntu 22.04 AMD64 server cloud image distribution, but is not included with the Ubuntu 22.04 AMD64 minimal cloud image. Installing `cron` will resolve this issue, but using the minimal cloud image is not supported and doing so may result in additional install failures.

### Curl responds with an error, or does not work

This is typically caused by network restrictions such as firewalls or lack of internet access. Using `curl` with verbose mode enabled (`-v` ) is useful for debugging.

```bash
# Example
curl -v https://go.pcd.run | bash
```

### Installation fails due to curl timeouts

This issue typically occurs early in the installation process, and can happen when the CIDRs used for Kubernetes services & pods conflicts with external local area networks. The default for the services CIDR is `10.21.0.0/16`, and the default for the pods CIDR is `10.20.0.0/16`. See the [Community Edition deployment networking](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation#community-edition-deployment-networking) section on the [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation) page for information on how to change these defaults using environment variables. After the environment variables are set, use the same terminal session to follow the steps in the section titled "Recovering a failed installation" on this page.

### Failed to find running pod percona-db-pxc-db-pxc-0

This error can occur when installing Community Edition on CPUs older than Intel Nehalem or AMD Bulldozer generations, as Percona requires the `x86-64-v2` microarchitecture level. This can be confirmed by running the following command:

```bash
sudo /lib64/ld-linux-x86-64.so.2 --help
```

Look for `Subdirectories of glibc-hwcaps` directories in the output. The `x86-64-v2` line should not say `supported, searched` .

Please try installing Community Edition again on CPUs newer than Intel Nehalem or AMD Bulldozer generations.

### Required ports & external URLs

Required ports:

* TCP: 53, 443
* UDP: 53, 123

Should you need to allow access to external URLs in order to complete a Community Edition installation, here are all of the URLs that are currently accessed during install:

```none
api2.amplitude.com
auth.docker.io
cdn.dl.k8s.io
cdn01.quay.io
check.percona.com
checkpoint-api.hashicorp.com
cr.fluentbit.io
dl.k8s.io
dockermirror.platform9.io
docs.tigera.io
gcr.io
ghcr.io
github.com
go.pcd.run
grafana.com
opencloud-dev-charts.s3.us-east-2.amazonaws.com
pcd-community.s3-accelerate.amazonaws.com
pkg-containers.githubusercontent.com
prod-registry-k8s-io-us-east-1.s3.dualstack.us-east-1.amazonaws.com
production.cloudflare.docker.com
pypi.org
quay.io
registry-1.docker.io
registry.k8s.io
release-assets.githubusercontent.com
storage.googleapis.com
us-east4-docker.pkg.dev
usage.projectcalico.org
```

Here is the list of required external URLs in order to onboard a hypervisor host to Community Edition.

```none
esm.ubuntu.com
files.pythonhosted.org
github.com
motd.ubuntu.com
pcdctl.s3.us-west-2.amazonaws.com
pypi.org
www.python.org
```

***

## Recovering a failed installation

Situation: A Community Edition installation has failed, the issue has been rectified, and you'd like to restart the installation.

Resolution: You must first uninstall/delete the CE installation (and k3s) using the following command, and then run the install script again.

```bash
# Fully delete CE installation & k3s
/opt/pf9/airctl/airctl delete-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml
# Run install script again
curl -sfL https://go.pcd.run | bash
```

After these steps have completed successfully, you may complete the rest of the deployment process starting with the [Local DNS Entries](/private-cloud-director/getting-started/getting-started-with-community-edition#local-dns-entries) section in [Getting Started With Community Edition](/private-cloud-director/getting-started/getting-started-with-community-edition).

***

## General troubleshooting

For CE hosts, use the following commands to understand where the error is occurring:

* Check CE's status after installing to see which region (Kubernetes namespace) is experiencing issues. `/opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml`
* For troubleshooting a failed install, check the install logs at `airctl-logs/airctl.log`
* Check the size and amount of available filesystem space `df -h /`
* `kubectl describe node` will show information about CE's underlying Kubernetes infrastructure
  * if resources are constrained, the allocated resources block should show requests CPU and memory near or at 100%
* `kubectl get pods -A` Look for any pods that are not Running or Completed. Pods in CrashLoopBackOff or with high restart counts and a young age should have their logs investigated for more detail.
* `kubectl logs <pod> -n <namespace>` to view pod logs.
* The default Kubernetes namespace is `pcd` for the infrastructure region, unless the `DU_FQDN` environment variable was customized prior to installation as described in the Deployment FQDN section of [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation).


# Backup and Restore CE

This document describes the backup & restore process for Private Cloud Director Community Edition (CE).

{% hint style="danger" %}
**Community Edition is not for production workloads.** Use it for labs, evaluation, and learning only.
{% endhint %}

## Community Edition backup & restore process

### Scope

This procedure backs up and restores the **Community Edition management plane** on the CE host.

It does not back up VM disks, external storage, or anything outside the CE host.

### Before you begin

* Run as a user with `sudo` access on the CE host.
* Ensure `/tmp` has enough free space for the backup file.
* Plan for downtime during restore. `unconfigure-du` stops the CE management plane.

{% hint style="warning" %}
The backup can include secrets (tokens, passwords, certificates). Store it securely and restrict file permissions.
{% endhint %}

The backup process for Community Edition is:

* Create a backup using `airctl backup`. The resulting backup file will be saved in the `/tmp` directory on the CE host.
* Save backup copies of the `k3s.yaml` and `airctl-config.yaml` configuration files as well.
* Move the backup file, and copy the other files to another directory to prevent unintentional overwrite.

The restore process for Community Edition is:

* `airctl unconfigure-du` will stop the Community Edition management plane, including all of its Kubernetes pods, but will leave the Kubernetes install in place.
* `airctl restore` uses the backup file to rebuild & restore your Community Edition installation.

The full command list is outlined below:

{% stepper %}
{% step %}

#### Create a backup (saves to /tmp)

```bash
/opt/pf9/airctl/airctl backup --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endstep %}

{% step %}

#### Create a backup directory, and move the backup file

```bash
mkdir backups
mv /tmp/backup_* backups
```

{% endstep %}

{% step %}

#### Save a copy of k3s and airctl configuration files

```bash
cp /opt/pf9/airctl/conf/airctl-config.yaml backups/
cp /etc/rancher/k3s/k3s.yaml backups/
```

{% endstep %}

{% step %}

#### (Recommended) Lock down backup permissions

```bash
chmod 700 backups
chmod 600 backups/*
```

{% endstep %}

{% step %}

#### Verify the backup files exist

```bash
ls -lh backups/
```

{% endstep %}

{% step %}

#### Unconfigure the CE deployment unit

```bash
/opt/pf9/airctl/airctl unconfigure-du --config /opt/pf9/airctl/conf/airctl-config.yaml --force
```

{% endstep %}

{% step %}

#### Restore the deployment unit from the backup

```bash
/opt/pf9/airctl/airctl restore --config /opt/pf9/airctl/conf/airctl-config.yaml --backupdir backups
```

{% endstep %}

{% step %}

#### Verify the restore

Your deployment is healthy when regions show `Ready` and ready services match desired services.

```bash
/opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endstep %}
{% endstepper %}


# Self Hosted

<code class="expression">space.vars.product\_name</code> can be deployed in your own infrastructure, giving you complete control over your cloud management platform while maintaining data sovereignty and security.

### Overview

The <code class="expression">space.vars.self\_hosted\_product\_name</code> deployment model allows you to run <code class="expression">space.vars.product\_name</code> entirely within your own data center or private cloud environment. This option is ideal for organizations with strict compliance requirements, air-gapped environments, or specific infrastructure needs.

### Getting Started

#### Prerequisites

Review the system requirements, network configuration, and infrastructure prerequisites needed to deploy <code class="expression">space.vars.product\_acronym</code> in your environment.

[View Prerequisites →](/private-cloud-director/getting-started/self-hosted/self-hosted-pre-requisites)

#### Installation

Follow our step-by-step installation guide to deploy <code class="expression">space.vars.product\_acronym</code> in your infrastructure. The installation process covers initial setup, configuration, and verification.

[Installation Guide →](/private-cloud-director/getting-started/self-hosted/self-hosted-install)

#### Upgrade

Learn how to upgrade your <code class="expression">space.vars.self\_hosted\_product\_name</code> deployment to the latest version while maintaining service availability and data integrity.

[Upgrade Guide →](/private-cloud-director/getting-started/self-hosted/upgrade)

For operational depth on host upgrades — including pre-upgrade environment checks, host role sequencing, Ubuntu 22.04-to-24.04 OS upgrade caveats, and failure recovery — see the [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook).

After completing the upgrade, follow the [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification) checklist to confirm all services and roles are healthy before re-enabling VM HA and DRR.


# Pre-requisites

This document outlines the prerequisites for deploying the <code class="expression">space.vars.self\_hosted\_product\_name</code>.

## Management Cluster

As part of the installation process, the <code class="expression">space.vars.self\_hosted\_product\_name</code> creates a **K3s cluster** using the physical servers that you use to deploy it on. We refer to this cluster as the **management cluster**. The <code class="expression">space.vars.product\_name</code> management plane then runs as a set of Kubernetes pods and services on this management cluster.

Single-node deployments are currently not supported for <code class="expression">space.vars.self\_hosted\_product\_name</code>. The minimum supported configuration requires 3 servers to ensure high availability and proper operation of the K3s management cluster. The 3-node cluster also lets core control-plane components — including the RabbitMQ message broker — run as multi-replica HA clusters by default, so a single management-server failure does not interrupt the management plane. For development or testing purposes, contact your Platform9 support for alternative deployment options.

The following is the recommended capacity for the management cluster, based on the projected scale of your <code class="expression">space.vars.product\_name</code> deployment. These configurations assume production deployments with high availability requirements.

| Hypervisors You Plan to Use           | Minimum Capacity                                                     | Recommended Capacity                                                                                                                                                                       |
| ------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| <p>Small<br><br>(<20 hosts)</p>       | <p>3 servers, each with:<br><br>20 vCPUs, 40GB RAM and 200GB SSD</p> | <p>3 servers, each with:<br><br>24 vCPUs, 48GB RAM and 250GB SSD</p>                                                                                                                       |
| <p>Growth<br><br>(<100 hosts)</p>     | <p>3 servers, each with:<br><br>26 vCPUs, 48GB RAM and 250GB SSD</p> | <p><strong>Option A</strong> — 3 servers, each with:<br>32 vCPUs, 64GB RAM and 250GB SSD<br><br><strong>Option B</strong> — 5 servers, each with:<br>24 vCPUs, 48GB RAM and 250GB SSD</p>  |
| <p>Enterprise<br><br>(>100 hosts)</p> | <p>3 servers, each with:<br><br>32 vCPUs, 96GB RAM and 250GB SSD</p> | <p><strong>Option A</strong> — 3 servers, each with:<br>48 vCPUs, 128GB RAM and 250GB SSD<br><br><strong>Option B</strong> — 5 servers, each with:<br>32 vCPUs, 80GB RAM and 250GB SSD</p> |

The above recommendation is for a single region. ***For each additional region*** deployed on the same management cluster, use the same per-tier table to size the incremental capacity, based on the number of hosts that region will manage. It is recommended to have a separate management cluster in each geographical location, to avoid performance degradation and a single point of failure.

For Growth and Enterprise scale, two recommended options are provided. The 3-server option keeps the cluster footprint small and is suitable when minimizing the number of management servers is a priority. The 5-server option distributes the core services and their replicas across more nodes, providing additional headroom to absorb a node failure, while lowering the per-node resource requirement. The total cluster capacity is comparable in both options.

### Disk Partition Guidance

<code class="expression">space.vars.product\_name</code> normally runs on a single root filesystem. If your environment requires dedicated partitions for compliance or security, size the directories that <code class="expression">space.vars.product\_name</code> depends on: **/var**, **/opt**, and **/etc**. These directories hold logs, container data, PF9 components, and configuration files, so they must have enough space to support normal operations and upgrades.

Recommended Sizes for /var, /opt, and /etc

| Directory | Minimum | Recommended |
| --------- | ------- | ----------- |
| /var      | 130GB   | 150GB       |
| /opt      | 10GB    | 20GB        |
| /etc      | 1GB     | 2GB         |
| /tmp      | 15GB    | 20GB        |
| /home     | 15GB    | 20GB        |
| /root     | 15GB    | 25GB        |

### Server Configuration

Each physical server that you use to run as part of the management cluster should meet the following requirements:

**Operating System**: Ubuntu 22.04, Ubuntu 24.04

### **Swap config**

Make sure that each server has swap disabled. You can run the following command to do this.

{% tabs %}
{% tab title="Bash" %}

```bash
swapoff -a
```

{% endtab %}
{% endtabs %}

The above change will not survive a reboot; hence, it is recommended to update the `/etc/fstab` file and comment out the line that has the entry for the `swap` partition. e.g.

{% tabs %}
{% tab title="Bash" %}

```bash
UUID=aabbcc /               ext4    errors=remount-ro 0       1
UUID=xxyyzz /home           ext4    defaults        	0       2
UUID=mswmsw /media/windows  ntfs    defaults				  0       0

#/dev/sdb1 none swap sw 0 0   <--- comment out the line
```

{% endtab %}
{% endtabs %}

### **IPv6 support**

Ensure the below sysctl setting is set to 0, so that IPv6 support is enabled on the server.

{% tabs %}
{% tab title="Bash" %}

```bash
sysctl net.ipv6.conf.all.disable_ipv6
# If currently set to 1, change it to 0 as below:
echo net.ipv6.conf.all.disable_ipv6=0 >> /etc/sysctl.conf
sysctl -p
```

{% endtab %}
{% endtabs %}

### **Passwordless Sudo**

Many operations require sudo access (for example, installing Yum repositories, Docker, etc.). Please ensure that your server has passwordless sudo enabled.

**Kernel Panic Option**

Update the server configuration section to include a step for setting `kernel.panic=10`

{% tabs %}
{% tab title="YAML" %}

```yaml
echo "kernel.panic=10" >> /etc/sysctl.conf && sysctl -p
```

{% endtab %}
{% endtabs %}

### **SSH Keys**

* We rely on SSH to log in to the management cluster hosts and to install various components and manage them.
* Please generate ssh keys and sync them across all hosts of the management cluster. The airctl installation is performed using a ***non-root user,*** so generate and use the SSH keys with the same non-root user that you select to run airctl operations. This ensures SSH access between hosts is established under the same user that performs the airctl operations.
* We recommend generating the key pair on one host and then adding the public key to all other hosts in their `~/.ssh/authorized_keys` file. This will enable every host in the management cluster to ssh into every other host.

{% tabs %}
{% tab title="Bash" %}

```bash
cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
ssh-copy-id -i ~/.ssh/id_rsa.pub root@test-3
```

{% endtab %}
{% endtabs %}

### **Package Updates**

* Install `cgroup-tools` :

{% tabs %}
{% tab title="Bash" %}

```bash
apt-get update -y && apt-get install cgroup-tools -y && apt-get install iptables -y
```

{% endtab %}
{% endtabs %}

* Download and Update OpenSSL Version to 3.0.7 for Ubuntu 22.04:

{% tabs %}
{% tab title="Bash" %}

```bash
export AGENT_KEY=<YOUR_USER_AGENT_KEY>
# Download the OpenSSL package
curl --user-agent "${AGENT_KEY}" https://pf9-airctl.s3-accelerate.amazonaws.com/openssl-smcp-ubuntu/openssl_3.0.7-1_amd64.deb --output /tmp/openssl_3.0.7-1_amd64.deb
# Verify the MD5 checksum
md5sum /tmp/openssl_3.0.7-1_amd64.deb | grep 706caf || { echo "MD5 checksum does not match, exiting." && exit 1; }
# Install the OpenSSL package
sudo dpkg -i /tmp/openssl_3.0.7-1_amd64.deb || { echo "Failed to install OpenSSL, exiting." && exit 1; }
echo "/usr/local/ssl/lib64" | sudo tee /etc/ld.so.conf.d/openssl-3.0.7.conf
sudo ldconfig -v
# Create a symbolic link to the OpenSSL binary
sudo ln -sf /usr/local/ssl/bin/openssl /usr/bin/openssl
# Verify the OpenSSL version
openssl version | grep 3.0.7 || { echo "OpenSSL version does not match, exiting." && exit 1; }
```

{% endtab %}
{% endtabs %}

* User Agent Key For Installation

You will need a specific Platform9 user agent key for the installation of your self-hosted management plane. Your Platform9 sales engineer will share the key with you prior to the install.

## Networking

You will need 2 virtual IPs that are on the same L2 domain as the hosts in the management cluster.

* VIP #1: This is the IP where you can access the <code class="expression">space.vars.product\_name</code> management plane UI.
* VIP #2: This is used to serve the management K3s cluster's API server.

## Storage

For a production setup of <code class="expression">space.vars.self\_hosted\_product\_name</code> you will need a Kubernetes Container Storage Interface (CSI) compatible storage for persisting the state of the management cluster. To know more, see [CSI and Kubernetes Storage](https://kubernetes.io/docs/concepts/storage/volumes/).

The Terrakube component of PCD AppCatalog requires persistent storage with multiple access (RWX) for sharing among Terrakube pods. Ensure a compatible RWX storage solution (like NFS, CephFS, or any CSI-compliant RWX provider) is available and configured in the K3s cluster where Terrakube runs.

### Storage Class Customisation

Customers have the flexibility to utilize a custom storage provisioner, allowing them to modify the *storage class* and *disk size* for each component according to their specific deployment needs. This section is optional; feel free to skip it if the default configuration meets your requirements.

**Default Storage Classes**

| Component       | Default Storage Class            | Comment                  |
| --------------- | -------------------------------- | ------------------------ |
| RabbitMQ        | `pcd-sc`                         | NFS type storage         |
| MySQL           | Cluster default (`hostpath-csi`) |                          |
| OVN OVSDB NB    | `pcd-sc`                         | NFS type storage         |
| OVN OVSDB SB    | `pcd-sc`                         | NFS type storage         |
| Prometheus      | `pcd-sc`                         | NFS type storage         |
| Terrakube       | `pcd-sc`                         | Requires RWX access mode |
| Terrakube Redis | Cluster default (`hostpath-csi`) |                          |
| Grafana         | Cluster default (`hostpath-csi`) |                          |
| Audit           | `pcd-sc`                         |                          |

To use the custom storage class provisioner, please follow these steps:

1. Place your custom storage class YAML files in the directory located at `/opt/pf9/airctl/conf/ddu/storage/custom/` . It is essential to have a storage class named `pcd-sc` included along with any other YAML files in this custom path.
2. Utilize `-p` or `--storage custom` in the `airctl configure` command. This will automatically apply all the YAML files found in the custom path.

*Example storage class yaml files:*

{% tabs %}
{% tab title="NFS Storage Class Example" %}

```yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: <STORAGE_CLASS_NAME>
mountOptions:
- nfsvers=4.1
- nolock
parameters:
  server: <NFS_SERVER_IP>
  share: <NFS_SHARE_PATH>
provisioner: nfs.csi.k8s.io
reclaimPolicy: Delete
volumeBindingMode: Immediate
```

{% endtab %}
{% endtabs %}

{% tabs %}
{% tab title="Hostpath Provisioner Storage Class Example" %}

```yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: <STORAGE_CLASS_NAME>
parameters:
  storagePool: standard
provisioner: kubevirt.io.hostpath-provisioner
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer
```

{% endtab %}
{% endtabs %}

Override storage class and disk size for specific component:

Edit the `/opt/pf9/airctl/conf/options.json` file to configure the settings specifically related to storage parameters. This adjustment empowers you to customize the storage class to meet your precise requirements. Updating the disk size is optional, as default sizes are already set.

Ensure that the storage class overrides specified in `options.json` are available in the custom storage path. If the storage class is not present, the deployment will fail.

{% tabs %}
{% tab title="Options.json" %}

```json
# Values here are given for example and not actual values.
{
  "chart_url": "<chart_url>",
  "terrakube_sc": "efs-sc",
  "terrakuberedis_sc": "block-sc",
  "rabbitmq_sc": "block-sc",
  "mysql_sc": "block-sc",
  "ovn_ovsdb_nb_sc": "efs-sc",
  "ovn_ovsdb_sb_sc": "efs-sc",
  "grafana_sc": "block-sc",
  "prometheusopenstack_sc": "efs-sc",
  "audit_pvc_sc": "audit-sc",

  "terrakube_disk_size": "11Gi",
  "terrakuberedis_disk_size": "2048Mi",
  "rabbitmq_disk_size": "6144Mi",
  "mysql_disk_size": "10Gi",
  "ovn_ovsdb_nb_disk_size": "11Gi",
  "ovn_ovsdb_sb_disk_size": "12Gi",
  "grafana_disk_size": "9Gi",
  "prometheusopenstack_disk_size": "9Gi",
  "audit_pvc_size": "6Gi"
}
```

{% endtab %}
{% endtabs %}

**Note:**

* Disk size units can be specified as Mi, Gi, or Ti.
* Decreasing disk size is **not supported**.
* Changing storage class name is supported only during installation, **not during upgrade.**

### **Inotify and Kernel limit Tuning**

Increase inotify limits to prevent airctl install from failing due to inotify errors

```bash
echo 'fs.inotify.max_queued_events = 512000
fs.inotify.max_user_instances = 1024 
fs.inotify.max_user_watches = 524288
fs.aio-max-nr = 500000
vm.panic_on_oom = 0' | sudo tee /etc/sysctl.d/99-pf9-airctl.conf > /dev/null
sudo sysctl -p /etc/sysctl.d/99-pf9-airctl.conf
```


# Install

This guide outlines the steps for a self-hosted deployment of <code class="expression">space.vars.product\_name</code>. Before installing, refer to the [Pre-requisites](/private-cloud-director/getting-started/pre-requisites) section to ensure all required prerequisites are met.

## Concepts

**Management Cluster**

As part of the installation process, the Self-Hosted version of <code class="expression">space.vars.product\_name</code> creates a K3s cluster using the physical servers that you use to deploy it on. We refer to this cluster as the **management cluster**. The <code class="expression">space.vars.product\_name</code> **management plane** then runs as a set of Kubernetes pods and services on this management cluster. The RabbitMQ message broker that mediates inter-service communication is deployed as a 3-node cluster on the management cluster by default for high availability.

**Infra Region vs Workload Regions**

Read [Tenant](/private-cloud-director/identity-and-multi-tenancy/tenant) to understand the concepts of regions and `infra` region in <code class="expression">space.vars.product\_name</code>.

## Download Installer

`airctl` is the command-line installer for <code class="expression">space.vars.self\_hosted\_product\_name</code>. Run the following commands only on one of the management cluster hosts to download `airctl` along with the required installer artifacts.

All the following steps should be performed by a non-root user.

#### Step 1: Download the Installer Script

Run the following command to download the installer script and required artifacts into your home folder:

{% tabs %}
{% tab title="Bash" %}

```bash
curl --user-agent "<YOUR_USER_AGENT_KEY>" https://pf9-airctl.s3-accelerate.amazonaws.com/latest/index.txt | awk '{print "curl -sS --user-agent \"<YOUR_USER_AGENT_KEY>\" \"https://pf9-airctl.s3-accelerate.amazonaws.com/latest/" $NF "\" -o ${HOME}/" $NF}' | bash
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**NOTE**

Replace `YOUR_USER_AGENT_KEY` in the command with the user agent key you requested from Platform9. For more details see [Pre-requisites](/private-cloud-director/getting-started/pre-requisites)
{% endhint %}

#### Step 2: Make the Installer Executable

Set the execute permissions on the installation script.

{% tabs %}
{% tab title="Bash" %}

```bash
chmod +x ./install-pcd.sh
```

{% endtab %}
{% endtabs %}

#### Step 3: Run the Installation Script

Execute the installer with the specified version. This runs the installer using the version number found in `version.txt`.

{% tabs %}
{% tab title="Bash" %}

```bash
./install-pcd.sh `cat version.txt`
```

{% endtab %}
{% endtabs %}

#### Step 4: Add airctl to System Path

Add `airctl` to the system path to use it globally by creating a symlink in `/usr/bin` folder.

{% tabs %}
{% tab title="Bash" %}

```bash
sudo ln -s /opt/pf9/airctl/airctl /usr/bin/airctl
```

{% endtab %}
{% endtabs %}

### Configure airctl

Run the following command to generate a configuration file, which will be used to deploy the <code class="expression">space.vars.self\_hosted\_product\_name</code> management cluster.

{% tabs %}
{% tab title="Bash" %}

```bash
/opt/pf9/airctl/airctl configure --du-fqdn pcd.platform9.localnet --external-ip4 10.149.106.15 --ipv4-enabled --master-ips 10.149.106.11,10.149.106.12,10.149.106.13 --master-vip-interface ens3 --master-vip4 10.149.106.16 --storage-provider custom --regions Region1 --worker-ips 10.149.106.14 --node-az-labels /path/to/node-az-labels.yaml
```

{% endtab %}
{% endtabs %}

Following are the input parameters for the command:

1. **du-fqdn -** Specify the base FQDN you would like to use for product\_name. For eg `pcd.mycompany.com`
2. **external-ip4 -** Specify the VIP to be used for the management plane here.
3. **ipv4-enabled** - Enable the IPv4 networking for the cluster.
4. **master-ips** *-* Specify a comma separated list of IP addresses for the master nodes you'd like to use for the management cluster. We recommend minimum 3 master nodes for a production environment. These nodes act as worker nodes as well.
5. **worker-ips (optional)** - If you'd like to add worker nodes to the management cluster, then specify the comma separated list of IP addresses for worker nodes here.
6. **master-vip-interface** - Specify the name of the network interface to be used for the Virtual IP of master nodes. Note that each master node must have it's default network interface named with this name.
7. **master-vip4** - Specify the Virtual IP to be used for the management cluster. This will be used to serve the management Kubernetes cluster's API server.
8. **master-vip-vrouterid** - This is optional. If unspecified, one will randomly be generated and can be found in the updated cluster spec saved in the directory that contains the state of the management server (see 'Important Directories' section for directory location). It is recommended to specify one if you plan to deploy multiple Kubernetes clusters in the same VLAN to avoid collision.
9. **regions** - Specify one or more region names as a space-separated list for regions you would like to create in your <code class="expression">space.vars.product\_name</code> setup. When specifying more than one regions the list needs to be enclosed in "". The final FQDN for your <code class="expression">space.vars.product\_name</code> deployment will use a combination of your base FQDN and your region name. For example, if your base FQDN is `pcd.mycompany.com` and you specified a single region name here as region1, the final FQDN for your deployment will be `pcd-region1.mycompany.com`.
10. **storage-provider** - Specify the CSI storage provider that should be used to store any persistent state for the management cluster. If not specified, `hostpath-provisioner` storage provider option will be selected as default.
    1. For `custom` as the storage provider, copy provider specific CSI yaml files to `/opt/pf9/airctl/conf/ddu/storage/custom/` locally. `airctl` will configure the dependencies reading the yaml files at this path and use the provided storage class for provisioning volumes for <code class="expression">space.vars.product\_name</code> components. There must be a yaml containing a `StorageClass` named `pcd-sc` and marked as the default StorageClass via the annotation `storageclass.kubernetes.io/is-default-class: "true"`.
    2. Alternatively, you can create the storage provider resources and storage class out of band prior to running `airctl start` on the management cluster. **Ensure that the storage class is named `pcd-sc` and is marked as default** (via `storageclass.kubernetes.io/is-default-class: "true"`) in this case. We also recommend creating a test pod with a persistent volume to ensure correct connectivity and configuration is in place first.
    3. For non-production environments and where custom storage provider is not available/required, select `hostpath-provisioner`.
11. **nfs-ip,** **nfs-share -** If using hostpath-provisioner as storage provider, NFS server IP and mount location must be provided.
    1. You should have an NFS server pre-configured before selecting this option.
12. **node-az-lables (optional)** - Provide labels and distribute the workload equally ([detailed doc](https://docs.platform9.com/private-cloud-director/~/revisions/Vf3rc8gZIkCBQty9ryKP/getting-started/self-hosted/availability-zones))
13. **enable-k8s (optional)** - Install the Kubernetes management plane alongside the <code class="expression">space.vars.product\_name</code> management plane, which lets you provision managed Kubernetes clusters. This is opt-in; omit the flag to leave it disabled.

{% hint style="info" %}
**Kubernetes management plane**

The Kubernetes management plane is optional and uses additional resources on the management plane. If you skip it during installation, you can enable it later on the existing deployment. See [Enable Kubernetes Management Plane](/private-cloud-director/getting-started/self-hosted/airctl/enable-kubernetes-management-plane).
{% endhint %}

This command generates two configuration file templates:

* `/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml` – Contains the configuration required to bootstrap the management cluster.
* `/opt/pf9/airctl/conf/airctl-config.yaml` – Contains the configuration for the management plane.

To avoid passing the configuration file as a command-line option when running `airctl` commands, copy `/opt/pf9/airctl/conf/airctl-config.yaml` to your `$HOME` directory.

{% tabs %}
{% tab title="Bash" %}

```bash
ln -s /opt/pf9/airctl/conf/airctl-config.yaml $HOME/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

### Proxy Configuration (Optional)

If your environment uses a network proxy, set the required values in the `/opt/pf9/airctl/conf/helm_values/bork.template.yml` file as shown below:

{% tabs %}
{% tab title="Bash" %}

```bash
cat /opt/pf9/airctl/conf/helm_values/bork.template.yml | grep proxy
# Sample output:
https_proxy: "http://squid.platform9.horse:3128"
http_proxy: "http://squid.platform9.horse:3128"
no_proxy: "10.149.106.11,10.149.106.12,10.149.106.13,10.149.106.14,10.149.106.15,10.149.106.16,127.0.0.1,10.20.0.0/22,localhost,::1,.svc,.svc.cluster.local,10.21.0.0/16,10.20.0.0/16,.cluster.local,.platform9.localnet,.default.svc"
```

{% endtab %}
{% endtabs %}

The list of I.P. addresses in the no\_proxy list should include the master-ips, worker-ips, external-ip4, master-vip4 along with any other addresses for which the traffic should not be routed via proxy server.

Also, to ensure that containerd honors the proxy values and allows the creation of the <code class="expression">space.vars.product\_name</code> management cluster, update the proxy values on all management plane nodes as shown below:

{% tabs %}
{% tab title="Bash" %}

```bash
cat /etc/environment
# Sample output:
PATH="/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin:/usr/local/ssl/bin"
HTTP_PROXY="http://squid.platform9.horse:3128"
https_proxy="http://squid.platform9.horse:3128"
http_proxy="http://squid.platform9.horse:3128"
HTTPS_PROXY="http://squid.platform9.horse:3128"
NO_PROXY="10.149.106.11,10.149.106.12,10.149.106.13,10.149.106.14,10.149.106.15,10.149.106.16,127.0.0.1,10.20.0.0/22,localhost,::1,.svc,.svc.cluster.local,10.21.0.0/16,10.20.0.0/16,.cluster.local,.platform9.localnet,.default.svc"
no_proxy="10.149.106.11,10.149.106.12,10.149.106.13,10.149.106.14,10.149.106.15,10.149.106.16,127.0.0.1,10.20.0.0/22,localhost,::1,.svc,.svc.cluster.local,10.21.0.0/16,10.20.0.0/16,.cluster.local,.platform9.localnet,.default.svc"
```

{% endtab %}
{% endtabs %}

{% tabs %}
{% tab title="Bash" %}

```bash
cat /etc/systemd/system/containerd.service.d/http-proxy.conf
# Sample output:
[Service]
EnvironmentFile=/etc/environment
```

{% endtab %}
{% endtabs %}

## Deploy Management Cluster

Next, deploy the management cluster for your <code class="expression">space.vars.product\_name</code> environment.

#### Step 1: Run Pre-Checks

Before creating the management cluster, run the following command to perform pre-checks and resolve any issues identified:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl check
```

{% endtab %}
{% endtabs %}

Optionally, configure AWS Credentials for `airctl` backup

Before creating the cluster, ensure that the AWS credentials required for S3 backup are available by creating the following file.

{% tabs %}
{% tab title="Bash" %}

```bash
Path: /etc/default/airctl-backup
Contents:

AWS_ACCESS_KEY_ID=<YOUR_ACCESS_KEY>
AWS_SECRET_ACCESS_KEY=<YOUR_SECRET_KEY>
AWS_REGION=<YOUR_AWS_REGION>
AWS_S3_PATH=s3://<YOUR_BUCKET_NAME_/PATH>
```

{% endtab %}
{% endtabs %}

When you create this file before cluster deployment, the system automatically creates a Kubernetes secret named `aws-credentials` in the `pf9-utils` namespace. You need this secret to upload backups to your S3 bucket.

Without this file, you must manually create or patch the `aws-credentials` secret in the `pf9-utils` namespace after cluster creation.

#### Step 2: Deploy the K3s Cluster

Run the following command to deploy the K3s cluster:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl create-cluster --verbose
```

{% endtab %}
{% endtabs %}

#### Step 3: Validate Cluster Health

Once the cluster is created, verify that it is functioning properly and that all nodes are healthy:

{% tabs %}
{% tab title="Bash" %}

```bash
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
kubectl get nodes
# Sample output:
NAME           STATUS   ROLES                       AGE     VERSION
10.10.11.111   Ready    control-plane,etcd,master   4d22h   v1.33.9+k3s1
10.10.12.242   Ready    control-plane,etcd,master   4d22h   v1.33.9+k3s1
10.10.13.253   Ready    control-plane,etcd,master   4d22h   v1.33.9+k3s1
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

Please refer to `/var/log/k3s.log` for troubleshooting any issues with the management cluster creation.
{% endhint %}

***

## Install Management Plane

Now that the management cluster is created, run the following commands to install and configure the <code class="expression">space.vars.product\_name</code> self-hosted management plane.

{% tabs %}
{% tab title="Bash" %}

```bash
airctl start

# Sample output:
 INFO  pcd-virt management plane creation started
 SUCCESS  generating certs and config...
 SUCCESS  setting up base infrastructure...
▀  starting consul...Secret consul-gossip-encryption-key in namespace default not found, creating new...
 INFO  bork setup done, creating management plane
 INFO  starting pcd-virt deployment...
 SUCCESS  pcd-virt deployment now complete
 INFO  pcd-virt management plane created - the services will take a while to start
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

It may take up to 45 minutes for all services to be deployed.
{% endhint %}

### Monitor Deployment Progress

You can track the progress of the management plane deployment by checking the logs of the `du-install` pod as a `root` user:

{% tabs %}
{% tab title="Bash" %}

```bash
export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
kubectl get pod -A | grep du-install
kubectl logs -n <ns> <pod name> -f
```

{% endtab %}
{% endtabs %}

Alternatively, to check specific pods:

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl get pods -n foo-bork | grep du-install
# Sample output:
du-install-foo-bmqqd                        0/1     Completed   0          108m
du-install-foo-region1-f7fdw                0/1     Completed   0          98m
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

Please refer to `airctl-logs/airctl.log` for logs in case of any issues with `airctl` start command.
{% endhint %}

### Check Management Plane & Region Status

To verify the status of the management plane and its regions, run:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl status
```

{% endtab %}
{% endtabs %}

## Obtain UI Credentials

Once the installation is complete, you can retrieve the credentials for the <code class="expression">space.vars.product\_name</code> UI by running the following command:.

{% tabs %}
{% tab title="Bash" %}

```bash
airctl get-creds
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**DNS Considerations**

If you do not have a working internal DNS that resolves the management plane FQDN to its IP address, you must manually update the `/etc/hosts` file on your local machine and any new host you want to add to the management cluster.
{% endhint %}

Use the following command to check if the necessary entry exists:

{% tabs %}
{% tab title="Bash" %}

```bash
cat /etc/hosts | grep foo
<VIP for externalIP>         foo-region1.bar.io
<VIP for externalIP>         foo.bar.io
<VIP for externalIP>         foo-bork.bar.io
```

{% endtab %}
{% endtabs %}

After updating the hosts file, open your web browser and log into the UI using the provided credentials. Then, follow the steps in the [Getting Started](/private-cloud-director/getting-started/getting-started) guide to configure your hosts and set up your <code class="expression">space.vars.product\_name</code> environment.

## Important Files & Directories

The following directories contain various log and state files related to the <code class="expression">space.vars.product\_name</code> self-hosted deployment.

#### Directories\*\*:\*\*

* `/opt/pf9/airctl` – Contains all binaries, offline installers, Docker image tar files, miscellaneous scripts, and configuration files for airctl.
* `/opt/pf9/pf9-kube` – Managed by Nodelet. Stores binaries and scripts used to manage the management cluster.
* `/etc/rancher/k3s/` – Contains configuration files for K3s and custom registries used by it.
* `/var/lib/rancher/k3s/` — stores K3s runtime data including certificates, containerd data, cluster state, manifests, and other internal components required for the cluster to run.
* `~/airctl-logs/` – Stores all logs related to the deployment.

#### Files\*\*:\*\*

* `/opt/pf9/airctl/conf/airctl-config.yaml` – Contains configuration information for the management plane.
* `/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml` – Stores configuration required to bootstrap the management cluster.
* `/var/log/k3s.log` – Log file for troubleshooting issues with management cluster creation.
* `~/airctl-logs/airctl.log` – Log file for troubleshooting issues with management plane creation.


# Install using the helper script

Deploy a self-hosted Private Cloud Director on-premises using the go.pcd.run/onprem helper script.

## Overview

The `go.pcd.run/onprem` helper script automates the on-premises deployment of self-hosted Private Cloud Director. It handles environment preparation tasks that the [Install](/private-cloud-director/getting-started/self-hosted/self-hosted-install) requires you to complete separately, including OpenSSL installation, swap configuration, kernel settings, and package dependencies, so you can deploy with fewer steps.

Use this method if you want a faster path to a running management plane on a network-connected environment. For air-gapped deployments, see [Air Gapped Installation](/private-cloud-director/getting-started/self-hosted/self-hosted-airgap-install).

### Prerequisites

Before running the script, ensure your environment meets the following requirements. Most prerequisites from the [Pre-requisites](/private-cloud-director/getting-started/self-hosted/self-hosted-pre-requisites) are handled automatically by the script. Only the following are required upfront:

* **Passwordless sudo**: The user running the script must have passwordless sudo configured.
* **SSH keys**: SSH key-based access must be set up so that the host running the script can connect to all master and worker nodes without a password prompt.

You also need a Platform9 user agent key. Your Platform9 sales engineer provides this key before installation.

#### Step 1: Set environment variables

The script requires a set of environment variables to run. Export the required variables and any optional ones that apply to your environment.

Replace all values in `<>` with your environment-specific values.

```bash
# Required
export USER_AGENT_KEY="<YOUR_USER_AGENT_KEY>"

# Optional — defaults shown
export SSH_KEY_PATH="<SSH_KEY_PATH>"           # Default: "${HOME}/.ssh/id_rsa"
export SSH_USER="<SSH_USER>"                   # Default: current user
export PCD_VERSION="<PCD_VERSION>"             # Default: "latest"

# Optional — proxy settings
export HTTPS_PROXY=""
export HTTP_PROXY=""
export NO_PROXY=""

# Optional — custom storage and options overrides
export AIRCTL_CUSTOM_CSI_FILES_PATH="<PATH_TO_DIR>"             # Required for custom storage class
export AIRCTL_OPTIONS_OVERRIDES_FILE_PATH="<PATH_TO_JSON_FILE>" # Required for options.json overrides
```

#### Step 2: Deploy <code class="expression">space.vars.product\_acronym</code>

Choose one of the following methods. Both methods require the environment variables from Step 1 to be exported first.

**Method 1: Environment variable-based installation**

Export your deployment parameters as environment variables, then run the script. Replace all values in `<>` with your environment-specific values.

```bash
export AIRCTL_MASTER_IPS="<MASTER_IP_1>,<MASTER_IP_2>,<MASTER_IP_3>" && \
export AIRCTL_MASTER_VIP4=<MASTER_VIP_IPV4> && \
export AIRCTL_REGIONS="<REGION_NAME>" && \
export AIRCTL_DU_FQDN=<DU_FQDN> && \
export AIRCTL_EXTERNAL_IP4=<EXTERNAL_IPV4> && \
export AIRCTL_STORAGE_PROVIDER="hostpath-provisioner" && \
export AIRCTL_NFS_IP=<NFS_SERVER_IP> && \
export AIRCTL_NFS_SHARE="<NFS_SHARE_PATH>" && \
export AIRCTL_IPV4_ENABLED=true && \
export AIRCTL_MASTER_VIP_INTERFACE=<INTERFACE_NAME>
```

To derive the full list of available environment variables, refer to the `airctl configure` flag list in the [Airctl Reference](/private-cloud-director/getting-started/self-hosted/airctl). Each long flag maps to an environment variable by prepending `AIRCTL_`, converting to uppercase, and replacing hyphens with underscores. For example, `--du-fqdn` becomes `AIRCTL_DU_FQDN`.

After exporting your variables, run the script:

```bash
curl -sS https://go.pcd.run/onprem | bash
```

**Method 2: Flag-based installation**

Pass deployment parameters directly to the script as flags. Replace all values in `<>` with your environment-specific values.

```bash
curl -sS https://go.pcd.run/onprem | bash -s -- \
  --du-fqdn <DU_FQDN> \
  --external-ip4 <EXTERNAL_IPV4> \
  --ipv4-enabled \
  --master-ips <MASTER_IP_1>,<MASTER_IP_2>,<MASTER_IP_3> \
  --master-vip-interface <INTERFACE_NAME> \
  --master-vip4 <MASTER_VIP_IPV4> \
  --nfs-ip <NFS_SERVER_IP> \
  --nfs-share <NFS_SHARE_PATH> \
  --storage-provider hostpath-provisioner \
  --regions "<REGION_NAME>"
```

For the full list of supported flags, see [Airctl Reference](/private-cloud-director/getting-started/self-hosted/airctl#configuration-commands) in the Airctl Reference.

{% hint style="info" %}
**NOTE**

The basic environment variables from Step 1, including `USER_AGENT_KEY` are required regardless of which method you use
{% endhint %}

### Proxy configuration (optional)

If your environment routes traffic through a network proxy, export the proxy variables before running the script. Replace all values in `<>` with your environment-specific values.

The `NO_PROXY` list must include all master node IPs, worker node IPs, `EXTERNAL_IP4`, and `MASTER_VIP4`, along with any other addresses that should bypass the proxy.

```bash
export HTTPS_PROXY="http://<PROXY_HOST>:<PROXY_PORT>"
export HTTP_PROXY="http://<PROXY_HOST>:<PROXY_PORT>"
export NO_PROXY="<MASTER_IP_1>,<MASTER_IP_2>,<MASTER_IP_3>,<EXTERNAL_IPV4>,<MASTER_VIP_IPV4>,127.0.0.1,10.20.0.0/22,localhost,::1,.svc,.svc.cluster.local,10.21.0.0/16,10.20.0.0/16,.cluster.local,.default.svc"
```

### Custom CSI storage backend (optional)

By default, the script uses `hostpath-provisioner` as the storage provider. To use a custom CSI storage backend instead, complete the following steps before running the script.

1. Create a directory and place your custom `StorageClass` YAML files in it.
2. Export the path to that directory:

```bash
export AIRCTL_CUSTOM_CSI_FILES_PATH="<PATH_TO_DIR>"
```

3. Set the storage provider to `custom` using the appropriate variable or flag for your chosen method:
   * Environment variable: `AIRCTL_STORAGE_PROVIDER="custom"`
   * Flag: `--storage-provider custom`

{% hint style="warning" %}
**NOTE**

Your custom YAML files must include at least one storage class named `pcd-sc`, along with all storage classes referenced in `options.json` if you are using an overrides file.
{% endhint %}

### Overriding default options.json (optional)

To customize storage classes per component or adjust management plane deployment settings, you can override fields in the default `options.json`.

1. Create a JSON file containing only the fields you want to add or override.
2. Export the path to that file before running the script:

```bash
export AIRCTL_OPTIONS_OVERRIDES_FILE_PATH="<PATH_TO_JSON_FILE>"
```

{% hint style="info" %}
**NOTE**

This variable cannot be passed as a flag to the script. It must be exported as an environment variable.
{% endhint %}


# Air Gapped Installation

Deploy Platform9 Private Cloud Director (PCD) in an environment without direct internet access.

### Overview

Use this guide to deploy Platform9 <code class="expression">space.vars.product\_name</code> (<code class="expression">space.vars.product\_acronym</code>) in an air-gapped environment, where cluster nodes have no direct internet access. Before you start the installation, download all required packages, container images, and installation artifacts on a host with internet access, then transfer them into your environment.

In an air gapped deployment, the following components must be set up inside your environment before <code class="expression">space.vars.product\_acronym</code> can be installed. You will configure each of these as part of this guide.

| Component                     | Purpose                                    |
| ----------------------------- | ------------------------------------------ |
| Private APT repository        | Hosts required Ubuntu packages             |
| Private container registry    | Hosts required container images            |
| Airctl installation artifacts | Installer components downloaded externally |
| NTP server                    | Ensures time synchronization across nodes  |

### Prerequisites

Before you begin, ensure the following conditions are met.

#### Network requirements

All cluster nodes must have network connectivity to the following internal services:

* Private APT repository
* Private container registry
* NTP server

For management plane host configuration, follow the prerequisites guide, except for package updates and OpenSSL installation. Those steps are covered in this guide.

#### DNS requirements

Configure DNS resolution for the container registry hostname.

**Example:**

```
registry.pf9.io → 10.10.13.145
```

If DNS is not available, add the registry hostname to `/etc/hosts` on every node instead.

**Example:**

```
10.10.13.145 registry.pf9.io
```

### Step 1: Configure NTP synchronization

All nodes must synchronize time with an NTP server inside your network to prevent clock skew across the cluster. In this step, you configure each node to point to your internal NTP server, restart the time synchronization service, and verify that synchronization is active before proceeding. If NTP is already configured on your nodes, skip this step.

1. Configure the NTP client on each node to point at your internal NTP server.

```bash
    sudo mkdir -p /etc/systemd/timesyncd.conf.d
    echo "[Time]
    NTP=<ntp-server-ip-or-fqdn>" | sudo tee /etc/systemd/timesyncd.conf.d/custom.conf
```

2. Restart the time synchronization service and enable it to start on boot.

bash

```bash
    sudo systemctl restart systemd-timesyncd
    sudo systemctl enable systemd-timesyncd
```

3. Verify that the node is actively syncing with the NTP server before moving on.

bash

```bash
    timedatectl status
    timedatectl show-timesync --all
```

### Step 2: Set up a private APT repository

The private APT repository hosts the Ubuntu packages required by the platform. You create this repository on a server with internet access and then make it accessible to the air-gapped cluster. To complete this step, you will download the required scripts and dependency list, download the package dependencies, and then initialize the repository.

1. On a host with internet connectivity, download the repository scripts and the dependency list.

```bash
    # Script to create a repo
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3.us-west-1.amazonaws.com/latest/sample_scripts/create_apt_repo.sh

    # Script to download package dependencies
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3.us-west-1.amazonaws.com/latest/sample_scripts/download_all_deps.sh

    # Dependency list
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3.us-west-1.amazonaws.com/latest/dependency_list.txt
```

2. Make both scripts executable before running them.

```bash
    chmod +x create_apt_repo.sh download_all_deps.sh
```

3. Download all package dependencies required for the installation.

```bash
    ./download_all_deps.sh <path-to-dependency-list.txt>
```

{% hint style="info" %}
**NOTE**

If you are using HTTPS with a custom CA, install the CA into `/usr/local/share/ca-certificates` and run `update-ca-certificates` before proceeding.
{% endhint %}

4\. Initialize the APT repository. Two options are available: HTTPS and HTTP. HTTPS is recommended because it encrypts package transfers and prevents tampering in transit. Use HTTP only if your environment does not support TLS.

{% hint style="info" %}
**NOTE**

Creating the repository typically takes approximately 20 minutes, depending on hardware performance.
{% endhint %}

**Option 1: HTTPS (recommended)**

If you do not have an existing certificate, generate a self-signed certificate and initialize the repository.

```bash
    sudo ./scripts/create_apt_repo.sh init-https repo.local
    sudo ./create_apt_repo.sh add-bulk ./deb_packages
```

If you already have a certificate and key, provide the paths directly.

```bash
    sudo ./scripts/create_apt_repo.sh init-https repo.example.com /path/to/cert.pem /path/to/key.pem
    sudo ./create_apt_repo.sh add-bulk ./deb_packages
```

**Option 2: HTTP (insecure)**

Use this option only if TLS is not available in your environment. Package transfers over HTTP are unencrypted and not verified.

```bash
    sudo ./create_apt_repo.sh init
    sudo ./create_apt_repo.sh add-bulk ./deb_packages
```

### Step 3: Configure the APT repository on cluster nodes

Update the APT configuration on all management and compute nodes so they can access the private repository.

1. If you are using a self-signed certificate, distribute the CA to each node and update the certificate store.

```bash
    sudo cp /path/to/apt_repo.crt /usr/local/share/ca-certificates/
    sudo update-ca-certificates
```

2. Back up your existing APT sources.

```bash
    sudo cp /etc/apt/sources.list /etc/apt/sources.list.bak
    sudo mkdir -p /etc/apt/sources.list.d.bak
    sudo mv /etc/apt/sources.list.d/*.list /etc/apt/sources.list.d.bak/ 2>/dev/null || true
    sudo rm /etc/apt/sources.list
```

3. Add the private repository and update the package index.

```bash
    echo "deb [trusted=yes] http://<repo-host>/ stable main" | sudo tee /etc/apt/sources.list.d/private-repo.list
    sudo apt update
```

### Step 4: Set up a private container registry

The private registry stores all container images required by <code class="expression">space.vars.product\_acronym</code>. You set up the registry on a host with internet access and then make it reachable from your air-gapped cluster nodes. To complete this step, you will download the registry setup scripts, initialize the registry, and confirm it is ready to receive images.

1. On a host with internet connectivity, download the registry setup scripts and the image list.

```bash
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3-accelerate.amazonaws.com/latest/pcdv-images.txt
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3.us-west-1.amazonaws.com/latest/sample_scripts/setup_registry.sh
    curl --user-agent "<TOKEN>" -O https://pf9-airctl.s3.us-west-1.amazonaws.com/latest/push-images.sh
```

2. Make the setup script executable and run it to initialize the registry.

```bash
    chmod +x setup_registry.sh
    ./setup_registry.sh
```

3. When prompted, provide the following inputs to configure the registry.

   | Prompt           | Example               | Description                  |
   | ---------------- | --------------------- | ---------------------------- |
   | Registry host IP | `10.10.13.145`        | IP of the registry server    |
   | Registry domain  | `registry.pf9.io`     | DNS name used by the cluster |
   | Username         | `<registry-user>`     | Registry login username      |
   | Password         | `<registry-password>` | Registry login password      |

**Example session:**

```bash
    Enter the registry host IP address: 10.10.13.145
    Enter the registry domain: registry.pf9.io
    Enter the username: admin
    Enter the password: <registryPassword>
```

### Step 5: Push images to the private registry

In this step, you push all required <code class="expression">space.vars.product\_acronym</code> container images to the private registry you created in Step 4. Before Docker can communicate with a registry that uses a self-signed certificate, you must add the registry's CA certificate to Docker's trust store. Without this, Docker rejects the connection and the image push fails.

{% hint style="info" %}
**NOTE**

Docker must be installed on the host machine before you can push images.
{% endhint %}

1. Add the registry's CA certificate to Docker's trust store.

```bash
    sudo mkdir -p /etc/docker/certs.d/<registry_url>:443
    sudo cp /usr/local/share/ca-certificates/ca.crt /etc/docker/certs.d/<registry_url>:443/ca-cert.pem
```

2. Restart the Docker service so the certificate change takes effect.

```bash
    sudo systemctl restart docker
```

3. Push the images to the private registry using the `push-images.sh` script.

```bash
    ./push-images.sh \
      --images-file-path <image_list_path> \
      --private-registry-url <registry_url> \
      --private-registry-username=<registry_username> \
      --private-registry-password=<registry_password>
```

### Step 6: Install the OpenSSL dependency

<code class="expression">space.vars.product\_acronym</code> requires a specific build of OpenSSL that is not available through the standard Ubuntu package repositories. In this step, you download the pre-built package from the Platform9 artifact store, transfer it to each cluster node, and install it. The checksum verification confirms that the package was not corrupted in transit.

1. On a host with internet access, download the OpenSSL package.

```bash
    curl --user-agent "<YOUR_USER_AGENT_KEY>" \
      https://pf9-airctl.s3-accelerate.amazonaws.com/openssl-smcp-ubuntu/openssl_3.0.7-1_amd64.deb \
      --output openssl_3.0.7-1_amd64.deb
```

2. Transfer the downloaded package to each cluster node.
3. On each node, verify the MD5 checksum to confirm the package integrity.

```bash
    md5sum openssl_3.0.7-1_amd64.deb | grep 706caf \
      || { echo "MD5 checksum does not match, exiting."; exit 1; }
```

4. Install the OpenSSL package and configure the library path.

```bash
    # Install the OpenSSL package
    sudo dpkg -i openssl_3.0.7-1_amd64.deb \
      || { echo "Failed to install OpenSSL, exiting."; exit 1; }

    # Add the OpenSSL library path
    echo "/usr/local/ssl/lib64" | sudo tee /etc/ld.so.conf.d/openssl-3.0.7.conf

    # Refresh the dynamic linker cache
    sudo ldconfig -v

    # Create a symbolic link to the new OpenSSL binary
    sudo ln -sf /usr/local/ssl/bin/openssl /usr/bin/openssl
```

5. Verify that the correct version of OpenSSL is active on the node.

```bash
    openssl version | grep 3.0.7 \
      || { echo "OpenSSL version does not match, exiting."; exit 1; }
```

Confirm that the output matches the following before continuing.

```
    OpenSSL 3.0.7 1 Nov 2022 (Library: OpenSSL 3.0.7 1 Nov 2022)
```

### Step 7: Install required packages

`cgroup-tools` is a Linux utility that <code class="expression">space.vars.product\_acronym</code> requires to manage and interact with control groups on each cluster node. Run the following command on each node to update the package index and install `cgroup-tools` in a single operation.

```bash
apt-get update -y && apt-get install cgroup-tools -y
```

### Step 8: Deploy <code class="expression">space.vars.product\_acronym</code>

With the private registry and APT repository in place, you can now run the <code class="expression">space.vars.product\_acronym</code> installer. Complete the following steps on a management node.

{% hint style="info" %}
**NOTE**

If the registry hostname is not resolvable via DNS in your environment, add the hostname to `/etc/hosts` on each node before proceeding.
{% endhint %}

1. On a system with internet connectivity, download the installer artifacts.

```bash
    curl --user-agent "<YOUR_USER_AGENT_KEY>" \
      https://pf9-airctl.s3-accelerate.amazonaws.com/latest/index.txt | \
      awk '{print "curl -sS --user-agent \"<YOUR_USER_AGENT_KEY>\" \"https://pf9-airctl.s3-accelerate.amazonaws.com/latest/" $NF "\" -o ${HOME}/" $NF}' | bash
```

2. Copy all downloaded artifacts to the management node, then make the installer executable.

```bash
    chmod +x ./install-pcd.sh
```

3. Run the installer using the version number from `version.txt`.

```bash
    ./install-pcd.sh $(cat version.txt)
```

4. Create a symlink so that `airctl` is available globally on the management node.

```bash
    sudo ln -s /opt/pf9/airctl/airctl /usr/bin/airctl
```

5. Generate the configuration file for your management cluster. Use single-master for a POC environment or multi-master for production.

{% hint style="info" %}
**NOTE**

Before running this command, copy the `ca.crt` file from the registry to the management node. Provide its path using the `--custom-registry-ca-cert-path` flag.
{% endhint %}

```bash
    airctl configure \
      -4 \
      -f <fqdn> \
      -e <du-public-ip> \
      -i <comma-separated-master-node-IPs> \
      --master-vip4 <vip-for-nodelet-cluster> \
      -v ens3 \
      -r <du-region-name> \
      -k <nfs-host-ip-for-hostpath-provisioner> \
      -n /mnt/gnocchi \
      -p hostpath-provisioner \
      --custom-registry-url https://registry.pf9.io \
      --custom-registry-username <registry_username> \
      --custom-registry-password "<registry_password>" \
      --custom-registry-ca-cert-path <path-to-registry-ca.crt> \
      --custom-registry-path-overrides \
      --enable-pcd-chart-bundle \
      --verbose
```

### Next steps

After completing the air-gapped setup, proceed to [Install](/private-cloud-director/getting-started/self-hosted/self-hosted-install) to deploy the management cluster and install the <code class="expression">space.vars.product\_acronym</code> management plane using `airctl`.


# Upgrade Management Plane and Hosts

{% hint style="info" %}
**Upgrading from 2026.1 to 2026.4?** An in-place upgrade is not supported for this version transition. Customers upgrading from **2026.1 to 2026.4** must follow the migration guide:

[Migrating from nodeletd to k3s-Based Clusters](/private-cloud-director/getting-started/self-hosted/migrating-from-nodeletd-to-k3s-based-clusters)
{% endhint %}

The upgrade of the self-hosted <code class="expression">space.vars.product\_name</code> is a two-phase, sequential process:

* **Management Plane:** Upgrading the Management Plane updates core services, management APIs, region-specific configurations, and orchestration components.
* **Hosts**: Once the management plane has been upgraded, hosts in each region are being upgraded to ensure compatibility, seamless communication, and completion of the overall upgrade process

Before you begin the upgrade, ensure you meet all prerequisites.

{% hint style="info" %}
**Detailed upgrade runbook and verification checklist**

This page documents the `airctl` commands to upgrade a self-hosted deployment. For deployment-model-agnostic operational guidance — pre-upgrade environment checks, host upgrade ordering, Ubuntu 22.04-to-24.04 OS upgrade caveats, and failure recovery — see the [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook).

After upgrading, use the [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification) checklist to confirm all services are healthy and to re-enable VM HA and DRR.
{% endhint %}

## Prerequisites for Upgrade

Before you upgrade, complete the following checks to ensure your environment is ready.

### Verify Management Plane Health

Before you upgrade, verify that the Management Plane is healthy. Run `airctl status` and confirm that `region health` shows as `Ready`. If any regions are not ready, resolve all issues before continuing. For example, pods may be in a `CrashLoopBackOff` state.

### Manual Backup

You must perform a [Manual Backup Procedure](/private-cloud-director/getting-started/self-hosted/backup-and-restore#manual-backup-procedure) of your management plane infrastructure before starting the upgrade. This backup is critical for recovery.

## Upgrade Management Plane

Upgrade your <code class="expression">space.vars.product\_name</code> environment to access the latest features, security patches, and performance improvements through this two-step process.

### Upgrade Management Plane

Updating the Management Plane running on the Management Cluster ensures multiple components are automatically upgraded as part of an `airctl upgrade` process.

Here are the components that will be affected.

* Core services and applications.
* Management APIs and user interfaces.
* Region-specific configurations and services.
* Service coordination and orchestration components.

#### The Impact of Management Plane Upgrade

{% hint style="info" %}
**NOTE**

Temporary service interruptions may occur during the upgrade process.
{% endhint %}

* The user interface will be unavailable during the upgrade.
* All API services will be temporarily unavailable.
* Create, Read, Update, and Delete operations will not be possible during this time.
* **Running workloads on VMs will not be impacted.**

#### Step 1: Download and run the Installer Artifacts

`airctl` is the command-line installer for <code class="expression">space.vars.product\_name</code>. Run the following commands **only on a verified management cluster node** to download `airctl` with the required installation artifacts.

{% hint style="info" %}
**Air-gapped environments only**

If the management node does not have internet access, download the installer artifacts on a jump host with internet connectivity and transfer them to the management node.

Follow the artifact download steps from the air-gapped installation guide, then skip this step and continue using the copied artifacts.
{% endhint %}

1. Download the Installer Script and Artifacts using the following command.

```bash
curl --user-agent "<YOUR-USER-AGENT-KEY>" https://pf9-airctl.s3-accelerate.amazonaws.com/latest/index.txt | awk '{print "curl -sS --user-agent \"<YOUR_USER_AGENT_KEY>\" \"https://pf9-airctl.s3-accelerate.amazonaws.com/latest/" $NF "\" -o ${HOME}/" $NF}' | bash
```

You can choose to download a specific version of the installer script and artifact.

For example, you can replace `latest` with a specific version, such as `v-2026.1.1-4444155`

Here is an example of a modified command:

{% tabs %}
{% tab title="Example" %}

```bash
curl --user-agent "<YOUR_USER_AGENT_KEY>" https://pf9-airctl.s3-accelerate.amazonaws.com/v-2026.1.1-4444155/index.txt | \
awk '{print "curl -sS --user-agent \"<YOUR_USER_AGENT_KEY>\" \"https://pf9-airctl.s3-accelerate.amazonaws.com/v-2026.1.1-4444155/" $NF "\" -o ${HOME}/" $NF}' | bash
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**NOTE**

Replace `<YOUR_USER_AGENT_KEY>` with your user key.
{% endhint %}

2. Make the Installer Executable using the following command.

```bash
chmod +x ./install-pcd.sh
```

3. Run the Installation Script using the following command.

```bash
./install-pcd.sh `cat version.txt`
```

The command installs the specific version of `airctl` based on the `version.txt` file.

4. Create symbolic links for `airctl` and its configuration file:

```bash
sudo rm -f /usr/bin/airctl # delete any existing file or symlink
sudo ln -s /opt/pf9/airctl/airctl /usr/bin/airctl
ln -s /opt/pf9/airctl/conf/airctl-config.yaml $HOME/airctl-config.yaml
```

#### Proxy Configuration (Optional)

If your environment uses a network proxy, update the values `/opt/pf9/airctl/conf/helm_values/bork.template.yml` again, as the changes would be lost after a new build is installed. Update it again as shown below:

```bash
cat /opt/pf9/airctl/conf/helm_values/bork.template.yml | grep proxy
# Sample output:
https_proxy: "http://squid.platform9.horse:3128"
http_proxy: "http://squid.platform9.horse:3128"
no_proxy: "10.149.106.11,10.149.106.12,10.149.106.13,10.149.106.14,10.149.106.15,10.149.106.16,127.0.0.1,10.20.0.0/22,localhost,::1,.svc,.svc.cluster.local,10.21.0.0/16,10.20.0.0/16,.cluster.local,.platform9.localnet,.default.svc"
```

The `no_proxy` list should include the `master-ips`, `worker-ips`, `external-ip4`, `master-vip4`, and any other addresses where traffic should not be routed through a proxy server.

{% hint style="info" %}
**NOTE**

The upgrade-cluster step is skipped in this release, as the Kubernetes version remains unchanged at **v1.30**.
{% endhint %}

{% hint style="info" %}
**Air-gapped environments only**

Ensure that all required container images for the target version are already available in the private registry accessible by the cluster nodes before running the upgrade.

If not already done, download the image list for the target version and push the images to your private registry. Refer to the air-gapped installation guide for detailed steps.
{% endhint %}

#### Step 2: Upgrade all regions

To upgrade **all** regions set up on your <code class="expression">space.vars.product\_name</code>, execute the following command.

```bash
airctl upgrade
```

**Optionally**, you can upgrade a specific region by replacing `<REGION_NAME>` with your target region.

```bash
airctl upgrade --region <REGION_NAME>
```

Here is the modified example command.

{% tabs %}
{% tab title="Example" %}

```json
# airctl upgrade --region Region1
 INFO  rollback state directory /tmp/airctl_ddu_backup_Region1_240020739                           
 INFO  Saving the helm revisions to the state file                                        
 INFO  --- backing up region--- Region1                                                             
 INFO  Archive created successfully for Region1. Backup backup.tar.gz saved to /tmp/airctl_ddu_backup_Region1_240020739                                                        
 INFO  --- moving old state file to /tmp/airctl_ddu_backup_Region1_240020739/state.yaml ---        
 INFO  --- moving old bork_values.yaml file to /tmp/airctl_ddu_backup_Region1_240020739/bork_values.yaml ---                                     
 SUCCESS  Upgrading region Region1                                                      
upgrade done
```

{% endtab %}
{% endtabs %}

To monitor and diagnose the upgrade logs, add `--verbose`.

Here is the sample command.

```bash
airctl upgrade --verbose
```

#### Verify: Upgrade Success on Management Plane

After completing the upgrade for the management plane, verify that you can continue to access the user interface, and then verify the deployment status across all regions.

Run the following command:

```bash
airctl status
```

Here is the sample output.

{% tabs %}
{% tab title="Example" %}

```bash
------------- deployment details ---------------
fqdn:                airctl-1-4206457-802.platform9.localnet
region:              airctl-1-4206457-802
deployment status:   ready
region health:       ✅ Ready
version:              PCD v-2026.1.1-4444155
-------- region service status ----------
desired services:     30
ready services:       30


------------- deployment details ---------------
fqdn:                airctl-1-4206457-802-mel.platform9.localnet
region:              airctl-1-4206457-802-mel
deployment status:   ready
region health:       ✅ Ready
version:              PCD v-2026.1.1-4444155
-------- region service status ----------
desired services:     84
ready services:       84
```

{% endtab %}
{% endtabs %}

After a successful upgrade and verification, <code class="expression">space.vars.product\_name</code> it **does not support rolling back** to a previous version.

## Upgrade the Hosts

After the management plane is successfully upgraded, the hosts running in each of your regions must be updated. The agents running on the Host manage communication between hosts and the management plane.

### The Impact from Hosts Upgrade

* No `servHostavailability` is expected during the host upgrade.
* Workloads running on VMs will continue to operate without disruption.

#### Step 1: Record the Current Version of Packages on the Host

Before you begin the upgrade, execute the following command to record the current package versions:

{% tabs %}
{% tab title="Example" %}

```bash
$ airctl host-status --config /opt/pf9/airctl/conf/airctl-config.yaml
Getting host statuses...                                                                                                                                                                                                                                                      
Getting host statuses for region: [REGION_NAME]                                                                                                                                                                                                                                         
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐                                                   
| Host Name   | IP Addresses                              | Host ID      | Host Agent Version     | Status | Agent Status | Apps                                          |
| [HOSTNAME1] | 172.29.32.34, 192.168.122.1, 10.0.11.10   | [HOST1_UUID] | v-2026.1.1-4444155.6f598c0 | ok     | running      | pf9-cindervolume-config:2026.1.1-4444155|
|             |                                           |              |                        |        |              | pf9-cindervolume-base:2026.1.1-4444155     |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-controller:2026.1.1-4444155 |
|             |                                           |              |                        |        |              | pf9-ostackhost:2026.1.1-4444155           |
|             |                                           |              |                        |        |              | pf9-glance-role:2026.1.1-4444155        |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-metadata-agent:2026.1.1-4444155 |
|             |                                           |              |                        |        |              | pf9-neutron-base:2026.1.1-4444155            |
| [HOSTNAME2] | 172.29.32.110, 192.168.122.1, 10.0.11.102 | [HOST2_UUID] | v-2026.1.1-4444155.6f598c0 | ok     | running      | pf9-ostackhost:2026.1.1-4444155          |
|             |                                           |              |                        |        |              | pf9-neutron-base:2026.1.1-4444155            |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-controller:2026.1.1-4444155   |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-metadata-agent:2026.1.1-4444155 |
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
 SUCCESS  getting host statuses...
```

{% endtab %}
{% endtabs %}

#### Step 2: Upgrade Hosts

To upgrade the hosts in a specific region, execute the following commands on the management cluster node:

{% hint style="info" %}
**NOTE**

The `upgrade-hosts` command times out after **1800 seconds (30 minutes)** by default. You can override this default by adding the `hostupgradetimeout` parameter in the `airctl-config.yaml` file. To prevent timeouts when upgrading a large number of hosts, increase this value proportionally — for every **50 hosts**, add **30 minutes (1800 seconds)** to the `hostupgradetimeout` value.
{% endhint %}

```bash
airctl upgrade-hosts --region <REGION_NAME>
```

The command triggers the creation of a `host-upgrade-xxxx` pod in the corresponding region namespace. You can monitor the upgrade progress or verify its success by checking the pod status and logs:

```bash
# List the host upgrade pods
kubectl get pods -n <REGION_FQDN> | grep host-upgrade

# View logs of the host upgrade pod
kubectl logs -n <REGION_FQDN> <HOST_UPGRADE_POD_NAME>
```

Here is a sample output:

{% tabs %}
{% tab title="Example" %}

```bash
host-upgrade-1747919688-g5rk5   5/5   Completed   0   83m
kubectl logs -n example-onprem-region1 host-upgrade-1747919688-g5rk5
```

{% endtab %}
{% endtabs %}

If a host fails to upgrade, run the following command to rerun the upgrade that the user executes using **IP.**

```bash
airctl upgrade-hosts --region <REGION_NAME> --host-ips "<HOST_IP>"
```

#### Step 3: Verify Upgrade Status of the Hosts

After the upgrade, confirm that the host packages have been successfully updated by running the following command.

{% tabs %}
{% tab title="Example" %}

```bash
$ airctl host-status --config /opt/pf9/airctl/conf/airctl-config.yaml
Getting host statuses...                                                                                                                                                                                                                                                      
Getting host statuses for region: [REGION_NAME]                                                                                                                                                                                                                                         
┌─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┐                                                   
| Host Name   | IP Addresses                              | Host ID      | Host Agent Version     | Status | Agent Status | Apps                                          |
| [HOSTNAME1] | 172.29.32.34, 192.168.122.1, 10.0.11.10   | [HOST1_UUID] | v-2026.1.1-4444155.6f598c0 | ok     | running      | pf9-cindervolume-config:2026.1.1-4444155       |
|             |                                           |              |                        |        |              | pf9-cindervolume-base:2026.1.1-4444155          |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-controller:2026.1.1-4444155     |
|             |                                           |              |                        |        |              | pf9-ostackhost:2026.1.1-4444155                 |
|             |                                           |              |                        |        |              | pf9-glance-role:2026.1.1-4444155              |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-metadata-agent:2026.1.1-4444155 |
|             |                                           |              |                        |        |              | pf9-neutron-base:2026.1.1-4444155               |
| [HOSTNAME2] | 172.29.32.110, 192.168.122.1, 10.0.11.102 | [HOST2_UUID] | v-2026.1.1-4444155.6f598c0 | ok     | running      | pf9-ostackhost:2026.1.1-4444155                |
|             |                                           |              |                        |        |              | pf9-neutron-base:2026.1.1-4444155              |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-controller:2026.1.1-4444155     |
|             |                                           |              |                        |        |              | pf9-neutron-ovn-metadata-agent:2026.1.1-4444155 |
└─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┘
 SUCCESS  getting host statuses...
```

{% endtab %}
{% endtabs %}

Compare the output with the previous recording to ensure the packages were updated as expected.

## Recovery: Upgrade Failure

#### Step 1. Stop the Failed Management Cluster

```bash
airctl stop --verbose
```

This command shuts down all components of the existing management cluster. `--verbose` helps verify the detailed output of all stopped services.

#### Step 2. Delete the Management Cluster Configuration

```bash
airctl unconfigure-du --force --verbose
```

This command removes all existing configuration files and metadata associated with the management cluster. Using `--force` enables overriding any locked or incomplete states.

#### Step 3: Delete the Existing Management Cluster

```bash
airctl delete-cluster --verbose
```

This command permanently deletes the management cluster resources. Ensure that you have backed up all critical data before running this command

#### Step 4. Create a New Management Cluster

```bash
airctl --config /opt/pf9/airctl/conf/airctl-config.yaml create-cluster --verbose
```

The command creates a new management cluster with the same configuration as before

#### Step 5. Restore from the Backup

```bash
airctl restore --backupdir <BACKUP_DIRECTORY> --verbose
```

Replace `<BACKUP_DIRECTORY>`with the actual path of your stored backup. The command restores the environment from the backup.

{% hint style="warning" %}
**Warning**

After a management plane upgrade, a rollback to a pre-upgrade version cannot be performed if any hosts have been upgraded as well.
{% endhint %}


# Migrating from nodeletd to k3s-Based Clusters

### Overview

This guide describes the recommended procedure for migrating an existing 2026.1 **nodeletd-based** Self-hosted Private Cloud Director deployment to a **k3s-based** cluster ( release 2026.4). The migration uses a backup-and-restore approach and is the supported upgrade path for on-premises customers in this release.

The high-level flow is:

1. Install the new release artifacts
2. Back up the existing (2026.1) deployment
3. Unconfigure and delete the existing cluster
4. Reconfigure and create a new k3s cluster
5. Restore the backup

> **Note:** This process involves deliberate cluster destruction and recreation. Plan for a maintenance window before proceeding.

***

### Prerequisites

Before starting the migration, confirm the following:

* You have access to the airctl management host.
* You have reviewed and noted all values from your current `/opt/pf9/airctl/conf/airctl-config.yaml` file and your existing nodelet bootstrap configuration ( `/opt/pf9/airctl/conf/nodelet-bootstrap-config.yaml` ) . You will need these values when reconfiguring the deployment.
* You have confirmed that a valid **restore target** exists (an accessible backup location or S3-compatible store).
* The **new release artifacts and packages** have been installed on the management host exactly as you would during a standard upgrade. Refer to the Upgrade Guide for package installation steps.
* You have read and understood the Important Considerations section below.

***

### Important Considerations

> ⚠️ **This migration involves irreversible steps.** Specifically, `airctl unconfigure-du` and `airctl delete-cluster` are **destructive and cannot be undone**. Ensure your backup is complete and verified before proceeding past Step 3.

* **Management Plane Downtime is expected.** This migration requires the cluster to be fully destroyed before the new k3s cluster is created. Plan for a maintenance window sized to your environment's restore time. In most environments, expect **several hours of downtime**.
* **There is no in-place rollback.** Once the cluster is deleted, rollback requires restoring from the backup taken in this procedure. Ensure the backup completes successfully and is accessible from the restore target before continuing.
* **Configuration values must be preserved.** During reconfiguration, you will re-run `airctl configure` using the same or equivalent inputs as the original installation. Derive these values from:
  * Your current `airctl-config` file
  * Your existing nodelet bootstrap configuration
* **Preserve your `options.json` file.** The `options.json` file, located at `/opt/pf9/airctl/conf/options.json`, controls deployment-specific settings such as storage class configuration, audit settings, and other platform options. This file is generally carried over by a new installation, but it is good practice to verify its contents before triggering the restore. If it is missing or outdated, recreate or update it to reflect your intended configuration before running the restore command.
* **Verify the backup before proceeding.** A corrupted or incomplete backup will prevent successful restore. Confirm the backup artifact is intact and readable before deleting the existing cluster.

***

### Migration Procedure

#### Step 0 — Back Up Airctl Host Configuration Files

Before installing new release artifacts, back up `options.json` and any other customized configuration files from the airctl host. The install script may overwrite `options.json` — preserving a copy now ensures you can restore your site-specific settings in Step 5a.

```bash
mkdir -p /tmp/airctl-conf-backup

# Back up the main options file
cp /opt/pf9/airctl/conf/options.json /tmp/airctl-conf-backup/options.json

# If you have other customized configuration files in /opt/pf9/airctl/conf/,
# copy them here as well before proceeding.
```

Keep this backup accessible on the host. You will reference it in Step 5a (Restore options.json Customizations).

***

#### Step 1 — Install New Release Artifacts

Install the new packages as you would for a standard upgrade. Do not restart or reconfigure the deployment at this stage.

Refer to the [Upgrade Guide](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/upgrade#step-1-download-and-run-the-installer-artifacts) for the artifact installation step — specifically, the step that runs the install script. **Follow only the artifact installation step.** Do not run `airctl upgrade`, `airctl start`, or any subsequent steps from the Upgrade Guide. Return to this guide after the install script completes.

> ⚠️ Ensure you have completed Step 0 before running the install script.

***

#### Step 2 — Back Up the Existing Deployment

Take a full backup of the existing deployment before making any configuration changes.

```bash
/opt/pf9/airctl/airctl backup --outdir /tmp/backup-mgmt/ --verbose
```

> ⚠️ **Do not proceed until the backup completes successfully.** Verify that the backup artifact exists at the expected destination and that it is not empty or corrupt. Keep note of the backup path or identifier — you will need it in Step 7.

***

#### Step 3 — Unconfigure the Existing Deployment

Before proceeding, verify that the deployment is healthy. This is the last safe stopping point before an irreversible operation.

```bash
airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

> ⚠️ Do not proceed if any region is degraded. Resolve issues first. The next command cannot be undone without a valid backup.

Uninstall the current deployment. This step stops services and clears the active configuration state.

```bash
/opt/pf9/airctl/airctl unconfigure-du
```

***

#### Step 4 — Delete the Existing Cluster

Delete the existing nodeletd cluster. This step is **irreversible**.

```bash
/opt/pf9/airctl/airctl delete-cluster
```

***

#### Step 5 — Reconfigure the Deployment

Re-run `airctl configure` using the same inputs that were used during the original installation.

For a full description of the `airctl configure` parameters and prompts, refer to the [Installation Guide](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/self-hosted-install#configure-airctl). Use the values from your saved `airctl-config` and nodelet bootstrap configuration as reference. Have both files open in a separate terminal session before proceeding.

> **Note:** The configuration inputs required here are identical to those used during initial installation. Do not use default or example values — use the values specific to your environment.

**Proxy configuration:** If your environment requires a proxy, configure it after `airctl configure` completes. Refer to the [Proxy Configuration](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/self-hosted-install#proxy-configuration-optional) section of the Installation Guide for instructions.

**Custom certificates:** If your original deployment used custom SSL/TLS certificates, you must supply them again during reconfiguration. Pass the certificate and key paths as arguments to `airctl configure`:

```bash
/opt/pf9/airctl/airctl configure \
  --user-cert-path /path/to/service.crt \
  --user-key-path /path/to/service.key
```

Alternatively, export them as environment variables before running the command:

```bash
export USER_CERT_PATH="/path/to/service.crt"
export USER_KEY_PATH="/path/to/service.key"
```

After `airctl configure` completes, verify that `user_cert_path` and `user_key_path` are present in `/opt/pf9/airctl/conf/airctl-config.yaml`. For full details, see [Using Custom Certificates](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/using-custom-certificates).

***

#### Step 5a — Restore options.json Customizations

After `airctl configure` completes, `options.json` is regenerated with default values. Verify the new file against your backup and merge back any site-specific customizations.

```bash
# Review differences between your backup and the newly generated file
diff /tmp/airctl-conf-backup/options.json \
     /opt/pf9/airctl/conf/options.json
```

For each value that differs and should reflect your site configuration, update the new file. You can edit it directly or use `jq` to restore specific keys. For example:

```bash
jq --slurpfile bk /tmp/airctl-conf-backup/options.json \
  '.rabbitmq_clustering_enabled = $bk[0].rabbitmq_clustering_enabled' \
  /opt/pf9/airctl/conf/options.json > /tmp/options.merged.json \
  && mv /tmp/options.merged.json /opt/pf9/airctl/conf/options.json
```

If you backed up other customized configuration files in Step 0, restore those now before continuing.

***

#### Step 5b — Verify and Recreate the airctl Symlink

After reconfiguration, confirm that the `airctl` binary symlink at `/usr/local/bin/airctl` points to the newly installed version.

```bash
# Verify the symlink
ls -la /usr/local/bin/airctl

# Recreate if missing or pointing to an old path
sudo ln -sf /opt/pf9/airctl/airctl /usr/local/bin/airctl

# Confirm the version is correct
airctl version
```

If your site uses a convenience symlink to the configuration file, recreate it now:

```bash
sudo ln -sf /opt/pf9/airctl/conf/airctl-config.yaml /etc/airctl-config.yaml
```

***

#### Step 6 — Validate Post-Configure State

After `airctl configure` completes, verify that the environment is correctly initialized before creating the cluster.

**Verify the updated** `/opt/pf9/airctl/conf/airctl-config.yaml` file.

Confirm that the configuration reflects the values you provided and that there are no unexpected changes.

**Verify the k3s bootstrap configuration:**

Confirm that the k3s bootstrap configuration has been created at the expected path: `/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml`

***

#### Step 7 — Create the New k3s Cluster

Create the new k3s-based cluster.

<pre class="language-bash"><code class="lang-bash"><strong>/opt/pf9/airctl/airctl create-cluster --verbose
</strong></code></pre>

Wait for the cluster creation to complete. Monitor the output for errors. Do not interrupt this process.

***

#### Step 8 — Restore the Backup

Once the new cluster is running, restore your data from the backup taken in Step 2.

{% hint style="info" %}
**RabbitMQ clustering:** The restore process enables RabbitMQ clustering by default. If you wish to disable it, update `/opt/pf9/airctl/conf/options.json` before proceeding with the restore by setting:

```json
"rabbitmq_clustering_enabled": "false"
```

{% endhint %}

Refer to the [Backup and Restore Guide](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/backup-and-restore) for detailed restore instructions.

**Quick reference — restore command:**

```bash
/opt/pf9/airctl/airctl restore --backupdir <backup-path> --verbose
```

Replace `<backup-path>` with the path or identifier recorded in Step 2.

> **Note:** The restore process may take significant time depending on the size of your deployment. Do not interrupt it once started.

***

### Validation Steps

After the restore completes, verify that the migration was successful:

1. **Confirm deployment health** — The deployment may take up to 10–20 minutes to fully initialize after restore. Once the initialization period has passed, confirm all regions are healthy:

```bash
airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

If any region remains degraded after 20 minutes, collect diagnostics and contact Platform9 Support.

2. **Confirm cluster node health** — Check that all cluster nodes report as healthy:

```bash
kubectl --kubeconfig /etc/rancher/k3s/k3s.yaml get nodes
```

3. **Confirm management plane availability** — Log in to the Private Cloud Director UI and confirm that the dashboard loads and all services are visible.
4. **Verify workloads and data** — Confirm that previously managed hosts, clusters, and configurations are present and accurate.

***

### **Upgrade hosts**

Once the restore is complete, the management plane upgrade is considered done. You must then upgrade the hosts in each region to ensure compatibility and complete the overall upgrade process. Refer to the [Upgrade Management Plane and Hosts](https://docs.platform9.com/private-cloud-director/getting-started/self-hosted/upgrade#upgrade-the-hosts) guide for host upgrade instructions.

***

### Post-Migration Notes

* The migration replaces the underlying cluster runtime from **nodeletd** to **k3s**. Application-level behavior and the management plane interface remain unchanged for end users.
* **kubeconfig path change:** The management cluster kubeconfig has moved. Update any scripts, `.bashrc` aliases, or CI pipelines that reference the previous path.

  | Cluster type                | kubeconfig path                                   |
  | --------------------------- | ------------------------------------------------- |
  | nodeletd (before migration) | `/etc/nodelet/airctl-mgmt/certs/admin.kubeconfig` |
  | k3s (after migration)       | `/etc/rancher/k3s/k3s.yaml`                       |

  ```bash
  export KUBECONFIG=/etc/rancher/k3s/k3s.yaml
  kubectl --kubeconfig /etc/rancher/k3s/k3s.yaml get nodes
  ```
* **kubectl binary path change:** The kubectl binary is now located at `/opt/pf9/airctl/bin/kubectl`. Update any scripts or aliases that invoke kubectl using an explicit path.

  ```bash
  /opt/pf9/airctl/bin/kubectl version --client
  ```
* If you encounter issues during or after restore, contact Platform9 Support and provide the `airctl-config` in use, and any relevant log output.
* Review the [Release Notes](https://docs.platform9.com/release-notes/april-2026-release) for known issues or post-migration configuration recommendations specific to this release.

***

### Rollback Considerations

There is no automated rollback for this migration. If the migration fails after the cluster has been deleted (Step 4 or later), recovery requires:

1. Re-running `airctl unconfigure-du --` with `--force` flag
2. Restoring from the backup created in Step 2.

If the backup cannot be restored, contact Platform9 Support immediately. **Do not delete or overwrite the backup artifact under any circumstances.**


# Cluster Scaling & Other Operations

This document describes steps to scale your K3s management cluster that is part of your self-hosted <code class="expression">space.vars.product\_name</code> deployment.

## Scale Up Management Cluster

Let's consider that we have 3 master nodes as part of your management cluster, with IP addresses ``1.1.1.1, 2.2.2.2 `and` 3.3.3.3``.

{% tabs %}
{% tab title="K3s Bootstrap Config File" %}

```bash
$ cat /opt/pf9/airctl/conf/k3s-bootstrap-config.yaml
...
masterNodes:
- nodeName: 1.1.1.1
- nodeName: 2.2.2.2
- nodeName: 3.3.3.3
```

{% endtab %}
{% endtabs %}

Lets say that you want to scale up the management cluster to 5 nodes, and that `4.4.4.4 & 5.5.5.5` are the IP addresses of the 2 new nodes to be added.

To scale up the number of cluster nodes to 5:

* Configure the [pre-requisites](/private-cloud-director/getting-started/self-hosted/self-hosted-pre-requisites) on the two new nodes.
* Edit the cluster bootstrap configuration file `/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml` and add the two new IP addresses to the section called `masterNodes` in the file.
* Finally, run the `airctl` command shown below to scale up the cluster.

{% hint style="info" %}
**Note**

`airctl` expects the node count to be always odd in the cluster bootstrap configuration file `/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml`
{% endhint %}

{% tabs %}
{% tab title="k3s Bootstrap Config File" %}

```bash
$ cat /opt/pf9/airctl/conf/k3s-bootstrap-config.yaml
...
masterNodes:
- nodeName: 1.1.1.1
- nodeName: 2.2.2.2
- nodeName: 3.3.3.3
- nodeName: 4.4.4.4
- nodeName: 5.5.5.5
```

{% endtab %}
{% endtabs %}

{% tabs %}
{% tab title="Scale Command" %}

```bash
airctl scale-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml --verbose
```

{% endtab %}
{% endtabs %}

Now verify that the management cluster has scaled up by querying the cluster nodes.

{% tabs %}
{% tab title="Cluster State Post Scale Up" %}

```bash
$ kubectl get nodes
NAME           STATUS   ROLES                        AGE      VERSION
1.1.1.1        Ready    control-plane,etcd,master    4d23h    v1.33.9+k3s1
2.2.2.2        Ready    control-plane,etcd,master    4d23h    v1.33.9+k3s1
3.3.3.3        Ready    control-plane,etcd,master    4d23h    v1.33.9+k3s1
4.4.4.4        Ready    control-plane,etcd,master    5m42s    v1.33.9+k3s1
5.5.5.5        Ready    control-plane,etcd,master    5m40s    v1.33.9+k3s1
```

{% endtab %}
{% endtabs %}

## Management Cluster Status

To check the status of your management cluster, run the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml --region <REGION_NAME>
```

{% endtab %}
{% endtabs %}

Sample output:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml --region foo-region1
# Sample output:
------------- deployment details ---------------
fqdn:                foo-region1.bar.io
cluster:             foo-bork.bar.io
region:              foo-region1
task state:          ready
version:             v-5.12.0-3479469
-------- region service status ----------
desired services:  45
ready services:    45
```

{% endtab %}
{% endtabs %}

## Scale Down Management Cluster

Now let's assume that we want to remove nodes `2.2.2.2` and `3.3.3.3` from the management cluster. To scale down the cluster:

* Edit the cluster bootstrap configuration file and remove the IP addresses of the two nodes.
* Then run the `airctl` command as shown below to scale down the cluster.

{% tabs %}
{% tab title="K3s Bootstrap Config File" %}

```bash
$ cat /opt/pf9/airctl/conf/k3s-bootstrap-config.yaml
...
masterNodes:
- nodeName: 1.1.1.1
- nodeName: 4.4.4.4
- nodeName: 5.5.5.5
```

{% endtab %}
{% endtabs %}

{% tabs %}
{% tab title="Scale Command" %}

```bash
airctl scale-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml --verbose
```

{% endtab %}
{% endtabs %}

* Cluster state post scale down operation completion:

{% tabs %}
{% tab title="Cluster State Post Scale Down" %}

```bash
$ kubectl get nodes
NAME            STATUS   ROLES                        AGE      VERSION
1.1.1.1         Ready    control-plane,etcd,master    4d23h   v1.33.9+k3s1
4.4.4.4         Ready    control-plane,etcd,master    5m42s    v1.33.9+k3s1
5.5.5.5         Ready    control-plane,etcd,master    5m40s    v1.33.9+k3s1
```

{% endtab %}
{% endtabs %}

## Uninstall Self-Hosted Deployment

To uninstall the specific region of your self-hosted deployment, run the following command. If you want to uninstall all regions, just remove `--region` flag.

{% tabs %}
{% tab title="Bash" %}

```bash
airctl unconfigure-du --config /opt/pf9/airctl/conf/airctl-config.yaml --region <REGION_NAME> --force
```

{% endtab %}
{% endtabs %}

This command will uninstall and remove configured regions along with all infrastructure software like consul, vault, percona, k8sniff, etc.

If you plan to reuse the same nodes to deploy a new self-hosed <code class="expression">space.vars.product\_name</code> environment, make sure to also run the following command on all nodes first.

{% tabs %}
{% tab title="Bash" %}

```bash
rm -rf airctl* install-pcd.sh nodelet* options.json pcd-chart.tgz /opt/pf9/airctl/ .airctl/
```

{% endtab %}
{% endtabs %}


# Backup and Restore Management Plane

This guide provides steps for backing up and restoring the self-hosted <code class="expression">space.vars.product\_name</code> management plane in disaster recovery scenarios. The procedures include both manual and automated backup methods, as well as manual restoration process.

{% hint style="info" %}
**NOTE**

When restoring the management plane, ensure it's done on a K3s management cluster that is separate from the cluster where the backup was generated.
{% endhint %}

## Prerequisites

#### System Requirements

* Access to the K3s management cluster
* Installed and configured `airctl` binary
* Valid `airctl` configuration file at `/opt/pf9/airctl/conf/airctl-config.yaml`
* Root or sudo access to the management node

#### For S3 Backup Storage

* AWS credentials with S3 bucket access
* Existing S3 bucket for backup storage
* AWS CLI configured (for verification purposes)

## Important Considerations

1. The restoration process must be performed on a separate K3s management cluster that is different from the management cluster where the backup was generated.

### Manual Backup Procedure

1. Create a backup directory:

{% tabs %}
{% tab title="Bash" %}

```bash
mkdir -p /tmp/backup-mgmt/
```

{% endtab %}
{% endtabs %}

2. Execute the airctl backup command:

{% hint style="info" %}
**Info**

Execute the following commands as a non-root user.
{% endhint %}

{% tabs %}
{% tab title="Bash" %}

```bash
airctl backup --outdir /tmp/backup-mgmt/ --config /opt/pf9/airctl/conf/airctl-config.yaml  --verbose
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
Use --region \<region\_name> parameter if you intend to back up only a specific region. If not specified, all the regions will be included in the backup.
{% endhint %}

3. Verify backup contents:

{% tabs %}
{% tab title="Bash" %}

```bash
tar tvf /tmp/backup-mgmt/backup.tar.gz
```

{% endtab %}
{% endtabs %}

The backup archive should contain:

* `state_backup.yaml`: System state configuration
* `bork_values_backup.yaml`: Kubernetes management cluster configuration
* `consul.snap`: Consul snapshot
* `mysql_dump_Infra.sql`: Infrastructure database backup
* `mysql_dump_Region1.sql`: Region-specific database backup
* `ovn-north-backup` & `ovn-south-backup` : Ovn database backup

## Automated Backup Configuration

The <code class="expression">space.vars.product\_name</code> management plane includes an automated backup system that protects your data and configuration. This system creates regular backups and can store them both locally and in Amazon S3.

Learn how to verify backup operations and configure S3 storage for your backups.

### Understanding the backup system

When you install <code class="expression">space.vars.product\_name</code> management plane, the system automatically sets up backup protection for you. During installation, it creates a cronjob called `mgmt-plane-backup` in the `pf9-utils` namespace that runs every hour to back up your system.

Your backups get stored in a dedicated storage area called `mgmt-plane-backup-pvc` on your K3s cluster. This storage persists even if pods restart, keeping your backup data safe and accessible.

#### Step 1: Verify backup operations

You can check the status of your backup system at any time using kubectl commands.

1. Run the following command to view the backup cronjob status.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl get cronjob mgmt-plane-backup -n pf9-utils
```

{% endtab %}
{% endtabs %}

This displays when you ran the last backup, confirming your system is working properly.

#### Step 2: Check backup logs

To troubleshoot backup operations or verify successful completion, you can view the backup logs.

1. List the backup pods to find the most recent operation.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl get po -n pf9-utils | grep mgmt-plane-backup
```

{% endtab %}
{% endtabs %}

2. From the output, copy the pod name with the most recent timestamp.
3. View the logs for that specific pod.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl logs mgmt-plane-backup-29167800-l294f -n pf9-utils
```

{% endtab %}
{% endtabs %}

Replace `mgmt-plane-backup-29167800-l294f`with your actual pod name from step 1.

The logs show detailed information about the backup operation, including any errors or success messages.

#### Step 3: Configure S3 backup storage

To enable storing backups in an S3 bucket, you need to configure S3 credentials in a secret named `aws-credentials` in the`pf9-utils` namespace.

Before you begin, consider the following points.

* The backup system stores data in the PVC named `mgmt-plane-backup-pvc` on the K3s cluster and will also upload to your configured S3 location.
* The `mgmt-plane-backup`cronjob runs hourly to ensure regular system backups to both local storage and S3.

1. Run the following command to open and edit `aws-credentials` on the editor.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl edit secret aws-credentials -n pf9-utils
```

{% endtab %}
{% endtabs %}

2. On the `aws-credentials` edit section `data:`

{% tabs %}
{% tab title="Bash" %}

```bash
# Make the edits to the yaml file

data:
  AWS_ACCESS_KEY_ID: "<YOUR_ACCESS_KEY>"
  AWS_SECRET_ACCESS_KEY: "<YOUR_SECRET_KEY>"
  AWS_REGION: "<YOUR_AWS_REGION>"
  AWS_S3_PATH: "s3://<YOUR_BUCKET_NAME/PATH/>"
```

{% endtab %}
{% endtabs %}

3. Replace the placeholder values with your actual AWS credentials.

| Placeholder              | Replace with                             |
| ------------------------ | ---------------------------------------- |
| `YOUR_ACCESS_KEY`        | Your AWS access key ID                   |
| `YOUR_SECRET_KEY`        | Your AWS secret access key               |
| `YOUR_AWS_REGION`        | Your AWS region (for example, us-west-2) |
| `YOUR_BUCKET_NAME/PATH/` | Your S3 bucket name and optional path    |

3. Save and close the editor.

Once configured, your backups will be stored both locally in the PVC and in your specified S3 bucket location. This provides enhanced data protection and allows for disaster recovery scenarios.

You have successfully configured automated backups for your <code class="expression">space.vars.product\_name</code> management plane. Your system now creates regular backups and stores them securely both locally and in Amazon S3.

## Manual Restore Procedure

When you need to restore your <code class="expression">space.vars.product\_name</code> management plane, you can access backups stored locally in the PVC or from Amazon S3. This section walks you through both restoration methods.

### Standard Restore (from local PVC)

To restore from local backups, you need to access the backup files stored in the mgmt-plane-backup-pvc. This process involves finding the PVC, locating the underlying storage, and mounting it to access the backup files.

#### Step 1: Locate the backup PVC

Run the following command to find the backup PVC in the `pf9-utils` namespace.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl get pvc -n pf9-utils
```

{% endtab %}
{% endtabs %}

The output displays your backup PVC details. Here is a sample example.

{% tabs %}
{% tab title="Example" %}

```bash
NAME                    STATUS   VOLUME                                     CAPACITY   ACCESS MODES   STORAGECLASS   VOLUMEATTRIBUTESCLASS   AGE
mgmt-plane-backup-pvc   Bound    pvc-cd010ade-87a2-4af1-bfce-ab7ce1308632   100Gi      RWO            pcd-sc         <unset>                 34h
```

{% endtab %}
{% endtabs %}

#### Step 2: Find the underlying persistent volume

Get the volume name for the backup PVC, by running the following command.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl get pvc mgmt-plane-backup-pvc -n pf9-utils -o jsonpath='{.spec.volumeName}'
```

{% endtab %}
{% endtabs %}

The output returns the persistent volume name. Here is an example.

{% tabs %}
{% tab title="Bash" %}

```bash
pvc-cd010ade-87a2-4af1-bfce-ab7ce1308632
```

{% endtab %}
{% endtabs %}

#### Step 3: Get the persistent volume configuration

Describe the persistent volume to find the NFS share information by running the following command.

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl describe pv pvc-cd010ade-87a2-4af1-bfce-ab7ce1308632
```

{% endtab %}
{% endtabs %}

In the output, look for the `Source:` section with `VolumeHandle` field. This contains the NFS server and path information that you need for mounting.

#### Step 4: Install NFS utilities

Install the required NFS packages on your system by running the following command.

{% tabs %}
{% tab title="Bash" %}

```bash
ubuntu@sample-vm:~$ sudo apt update && sudo apt -y install nfs-common
```

{% endtab %}
{% endtabs %}

#### Step 5: Mount the NFS share

Create a local directory and mount the NFS share:

{% tabs %}
{% tab title="Bash" %}

```bash
ubuntu@sample-vm:~$ mkdir -p /home/ubuntu/backups
ubuntu@sample-vm:~$ sudo mount -t nfs 10.149.106.253:/mnt/gnocchi /home/ubuntu/backups
```

{% endtab %}
{% endtabs %}

Replace `10.149.106.253:/mnt/gnocchi` with the server and share path from your `VolumeHandle` field.

#### Step 6: Access the backup files

List the available backup files by running the following command

{% tabs %}
{% tab title="Bash" %}

```bash
ubuntu@sample-vm:~$ ls backups/pvc-cd010ade-87a2-4af1-bfce-ab7ce1308632
```

{% endtab %}
{% endtabs %}

The backup files will display with timestamps. Here is a sample output.

{% tabs %}
{% tab title="Bash" %}

```bash
backup_20250608_210055.tar.gz  backup_20250608_230057.tar.gz  backup_20250609_010057.tar.gz
backup_20250609_030054.tar.gz  backup_20250609_050055.tar.gz  backup_20250608_220057.tar.gz
```

{% endtab %}
{% endtabs %}

Choose the backup file you want to restore based on the timestamp that matches your desired restore point.

#### Step 7: Execute the restore command

Extract and specify the backup directory path, then run the restore command:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl restore --backup-dir <backup_dir_path> --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

Replace `<backup`*`dir`*`path>`with the path to your extracted backup directory.

{% hint style="info" %}
**NOTE**

Execute the restore command as a non-root user, similar to the backup procedure.
{% endhint %}

### Restore from S3 Backup

Create and configure the `/etc/default/airctl-backup` file with required AWS parameters, making sure that `AWS_S3_PATH` points specifically to the backup file you want to restore, not just the S3 bucket:

{% tabs %}
{% tab title="Bash" %}

```bash
AWS_ACCESS_KEY_ID=<YOUR_ACCESS_KEY>
AWS_SECRET_ACCESS_KEY=<YOUR_SECRET_KEY>
AWS_REGION=<YOUR_AWS_REGION>
AWS_S3_PATH=s3://<YOUR_BUCKET_NAME/PATH/SPECIFIC_BACKUP_FILE>
```

{% endtab %}
{% endtabs %}

Execute the S3 restore command:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl restore --s3backup --config /opt/pf9/airctl/conf/airctl-config.yaml --verbose
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**NOTE**

For complete disaster recovery, manually restore Gnocchi metrics data from the original `pcd-sc` persistent volume
{% endhint %}

Verification Steps

Check backup file integrity using MD5 checksum::

{% tabs %}
{% tab title="Bash" %}

```bash
# Generate MD5 checksum for the backup file
md5sum <BACKUP_FILE>.tar.gz

# Optional: Compare with a pre-recorded checksum
# You can save the MD5 checksum when initially creating the backup
md5sum /root/backup-mgmt/backup.tar.gz > backup-checksum.txt

# Later, verify the backup file matches the original checksum
md5sum -c backup-checksum.txt
```

{% endtab %}
{% endtabs %}

Verify S3 uploads (if configured):

{% tabs %}
{% tab title="Bash" %}

```bash
aws s3 ls s3://<BUCKET_NAME>/<BACKUP_PATH>
```

{% endtab %}
{% endtabs %}

Monitor restore progress:

{% tabs %}
{% tab title="Bash" %}

```bash
kubectl logs -f <RESTORE_POD_NAME>
```

{% endtab %}
{% endtabs %}

## Common Issues

* If AWS credentials are not properly configured, automated S3 backups will continue locally but skip S3 upload.
* Restore operations may take significant time depending on data volume.
* Services may take additional time to start after restore completion.


# Airctl Reference

Airctl is an on-premises Kubernetes orchestrator by Platform9. Comprising a useful command-line utility and related functions that manage the lifecycle of the Management Server on the Management Plane of <code class="expression">space.vars.product\_name</code>.

This powerful tool serves as your control center for managing <code class="expression">space.vars.product\_name</code> on-premises environment. Airctl provides comprehensive management capabilities that ensure your Kubernetes environment runs smoothly while automatically handling critical infrastructure components.

### What Airctl Does?

* **Cluster Lifecycle Management**: Create, configure, scale, and upgrade K3s management clusters.
* **Backup & Recovery**: Protect your cluster configuration and MySQL database.

### Prerequisites

Before using Airctl, ensure your environment meets these requirements:

* **Operating System**: Ubuntu 22.04 with `sudo` access.
* **Storage**: NFS server (for shared storage scenarios).
* **Resources**: Minimum hardware requirements for your target deployment type.

### Airctl Core Functions

Airctl offers functions that simplify on-premises management.

#### State management

Initializes and configures the Management Plane, creating necessary components automatically when needed.

#### Database Backup

Facilitates MySQL database backup for the Platform9 data store. Optimized for typical three-node cluster configurations.

#### Configuration Management

Manages cluster state through configuration files containing node IP addresses and credentials. These files describe the desired state of your cluster and serve as the foundation for deployment and management operations.

#### CLI Utilities

Provides comprehensive command-line utilities for backup and recovery of cluster state. These utilities ensure you can protect, restore, and maintain your entire management cluster configuration, enabling disaster recovery and environment replication across different deployments.

## Command Reference

### Cluster Management Commands

| **Command**       | **Description**                                 | **Use Case**               |
| ----------------- | ----------------------------------------------- | -------------------------- |
| `create-cluster`  | Create a k3s-based management cluster           | Initial cluster deployment |
| `delete-cluster`  | Delete a headless management cluster            | Cluster decommissioning    |
| `scale-cluster`   | Scale up or down a k3s-based management cluster | Capacity adjustment        |
| `upgrade-cluster` | Upgrade a k3s-based management cluster          | Version updates            |
| `start`           | Start the Management Plane and auto-configure   | Normal operations          |
| `stop`            | Stop the Management Plane                       | Maintenance windows        |
| `status`          | Report Management Plane status                  | Health monitoring          |
| `upgrade`         | Upgrade the management plane                    | Normal operations          |
| `add-region`      | Add a new region to the PCD deployment          | Advanced multi tenancy     |

### Configuration Commands

| **Command**             | **Description**                         | **Example** |
| ----------------------- | --------------------------------------- | ----------- |
| `configure`             | Generate management plane configuration |             |
| `update-admin-password` | Set airctl admin password               |             |
| `get-creds`             | Display admin credentials               |             |

### Backup and Recovery Commands

These CLI utilities provide comprehensive backup and recovery capabilities for both MySQL database and cluster state management.

| **Command**          | **Description**                                  | **Output**                |
| -------------------- | ------------------------------------------------ | ------------------------- |
| `backup`             | Back up management plane state and configuration | Creates tar.gz file       |
| `restore`            | Restore from backup (includes cluster state)     | Restores from tar.gz file |
| `gen-support-bundle` | Generate support bundle                          | Diagnostic package        |

### Maintenance Commands

| **Command**       | **Description**                                                       | **When to Use**                  |
| ----------------- | --------------------------------------------------------------------- | -------------------------------- |
| `check`           | Verify k3s pre-requisites on all management cluster nodes (see below) | Before installation              |
| `upgrade`         | Upgrade management plane apps                                         | Version updates                  |
| `reset-taskstate` | Reset deployment state after resolving errors                         | Management plane upgrade failure |

#### What `airctl check` Verifies

`airctl --config <airctl-config.yaml> check` runs a pre-requisites pass on every node listed in the airctl config file. Per node, it confirms:

* The node is **Ready** and reachable.
* **Disk Space** — at least 60 GB free.
* **Swap** is disabled.
* **IPv6 Support** is enabled.
* **Kernel and VM Panic Settings** are configured.
* **Port Connectivity (K3s)** — the ports k3s requires are not already in use: `2379`, `2380`, `6443`, `10250`.
* The **Firewalld** service is in the expected state.
* **Default Route Weights** are correct.
* An **NTP Service** is installed and the node clock is synchronized (see [Pre-requisites](/private-cloud-director/getting-started/pre-requisites)).

The command reports `✓` (pass) or `x` (fail) per check on each node, and a per-node summary at the end. Address any failures before running `airctl create-cluster` or `airctl upgrade`.

### Advanced Commands

| **Command**      | **Description**                                 | **Use Case**           |
| ---------------- | ----------------------------------------------- | ---------------------- |
| `unconfigure-du` | Delete the Management Plane configuration state | Clean slate deployment |

### Key Configuration Options

#### General Configuration

| **Flag**                        | **Decription**                          | **Default**                 |
| ------------------------------- | --------------------------------------- | --------------------------- |
| `-d, --cluster-deployment-type` | Choose 'onpremk3s' or 'k3s'             | onpremk3s                   |
| `-f, --du-fqdn`                 | Management plane (DU) FQDN              | airctl-1.platform9.localnet |
| `-r, --regions`                 | Specify regions for the DU (at least 1) | Region1                     |
| `-a, --enable-k8s`              | Enable the Kubernetes management plane  | false                       |

#### Node Configuration

| **Flag**           | **Decription**                             | **Default** |
| ------------------ | ------------------------------------------ | ----------- |
| `-i, --master-ips` | Master node IP addresses (comma-separated) | -           |
| `-w, --worker-ips` | Worker node IP addresses (comma-separated) | -           |
| `--ssh-port` int   | SSH port for nodes                         | 22          |

#### Network Configuration

| **Flag**             | **Decription**                                                                                    | **Default** |
| -------------------- | ------------------------------------------------------------------------------------------------- | ----------- |
| `-4, --ipv4-enabled` | <p>Enable IPv4 networking.<br>To enable dual-stack networks, set this flag and "ipv6-enabled"</p> | true        |
| `-6, --ipv6-enabled` | <p>Enable IPv6 networking.<br>To enable dual-stack networks, set this flag and "ipv4-enabled"</p> | false       |
| `-e, --external-ip4` | Set the external IPv4 for the DU                                                                  | -           |
| `-x, --external-ip6` | Set the external IPv6 for the DU                                                                  | -           |

#### VIP Configuration

| **Flag**                     | **Decription**                                               | **Default** |
| ---------------------------- | ------------------------------------------------------------ | ----------- |
| `--master-vip4`              | Virtual IP for Management Cluster (IPv4)                     | -           |
| `--master-vip6`              | Virtual IP for Management Cluster (IPv6)                     | -           |
| `-v, --master-vip-interface` | Specify the masterVipVrouterId since more than 1 master node | -           |
| `-z, --master-vip-vrouterid` | Specify the masterVipVrouterId since more than 1 master node | -           |

#### Storage Configuration

| **Flag**                 | **Decription**                                                                                                                                     | **Default**          |
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- |
| `-p, --storage-provider` | StorageProvider (Either 'hostpath-provisioner' or 'custom') (For custom, save provider-specific YAMLs to /opt/pf9/airctl/conf/ddu/storage/custom/) | hostpath-provisioner |
| `-k, --nfs-ip`           | Set the NFS server IP address for DU                                                                                                               | -                    |
| `-n, --nfs-share`        | Specify the NFS share path for DU                                                                                                                  | -                    |

#### K3s Configuration

| **Flag**                        | **Decription**                               | **Default**  |
| ------------------------------- | -------------------------------------------- | ------------ |
| `-c, --k3s-pod-cidr`            | Specify the Pod CIDR for the k3s cluster     | 10.20.0.0/16 |
| `-s, --k3s-service-cidr`Service | Specify the service CIDR for the k3s cluster | 10.21.0.0/16 |

#### Registry Configuration

| **Flag**                                                   | **Decription**                                                             | **Default** |
| ---------------------------------------------------------- | -------------------------------------------------------------------------- | ----------- |
| <p><code>--registry-type</code><br>privateRegistryType</p> | Type of container image registry to use. Possible values: DU, custom, none | DU          |

## Global Flags

These flags can be used with any Airctl command:

| **Flag**        | **Decription**                    | **Example**                  |
| --------------- | --------------------------------- | ---------------------------- |
| `--config`      | config file                       | `$HOME/airctl-config.yaml`   |
| `--quiet`       | Disable spinners                  |                              |
| `--verbose`     | Print verbose logs to the console | `airctl configure --verbose` |
| `-h, --help`    | Show command help                 | `airctl configure --help`    |
| `-v, --version` | Version for airctl                |                              |

Configure your management cluster options

Select from various installation types tailored to your specific needs and environmental requirements.

* **Single-master management cluster** - Ideal for proof-of-concept (POC) installations and testing environments where high availability isn't critical.
* **Multi-master management cluster** - Required for production installations where you need high availability and redundancy for your workloads.

## Airctl Configuration Reference

{% tabs %}
{% tab title="YAML" %}

```yaml
# airctl-config.yaml Schema Documentation
# This file documents all possible configuration options for airctl-config.yaml
# Format: YAML with JSON Schema-like descriptions

type: object
properties:
  # Installation and Deployment Configuration
  installDriver:
    type: string
    enum: ["DDU", "openstack"]
    default: "openstack"
    description: "Type of installation driver to use for deployment"

  clusterDeploymentType:
    type: string
    enum: ["onpremk3s", "k3s"]
    description: "Type of cluster deployment (onpremk3s or k3s) - k3s is supported in Community Edition"

  # DU Configuration
  duFqdn:
    type: string
    required: true
    example: "airctl-1.pf9.localnet"
    description: "Fully qualified domain name for the DU"

  duTenant:
    type: string
    default: "service"
    description: "DU tenant identifier"

  duRegion:
    type: string
    default: "Infra"
    example: "Infra RegionOne"
    description: "DU region(s) - space-separated for multiple regions"

  duUser:
    type: string
    example: "admin@airctl.localnet"
    description: "DU admin user email/username"

  duPassword:
    type: string
    description: "DU admin user password"

  # SSH Configuration
  sshUser:
    type: string
    required: true
    example: "ubuntu"
    description: "SSH username for node access"

  sshPort:
    type: integer
    default: 22
    description: "SSH port for node access"

  sshPublicKeyFile:
    type: string
    default: "~/.ssh/id_rsa.pub"
    description: "Path to SSH public key file"

  sshPrivateKeyFile:
    type: string
    default: "~/.ssh/id_rsa"
    description: "Path to SSH private key file"

  # Network Configuration
  externalIP:
    type: string
    example: "10.128.135.5"
    description: "External IP (VIP) to access the control plane on the UI. Needs to be in same subnet as cluster nodes"

  externalIpV6:
    type: string
    description: "External IP (VIP) to access the control plane on the UI. Needs to be in same subnet as cluster nodes"

  IpV4:
    type: boolean
    default: true
    description: "Enable IPv4 support"

  IpV6:
    type: boolean
    default: false
    description: "Enable IPv6 support"

  # Storage Configuration
  storageProvider:
    type: string
    enum: ["hostpath-provisioner", "custom"]
    default: "hostpath-provisioner"
    description: "Storage provider type"

  nfsServerIP:
    type: string
    example: "10.10.10.10"
    description: "NFS server IP address"

  nfsShare:
    type: string
    example: "/mnt/pcd-data-share"
    description: "NFS share path"

  # Installation Paths and Directories
  installDir:
    type: string
    default: "/opt/pf9/airctl"
    description: "Airctl installation directory"

  logDir:
    type: string
    default: "~"
    description: "Log directory path (defaults to user home)"

  # Configuration File Paths
  kubeCfgPath:
    type: string
    default: "/etc/rancher/k3s/k3s.yaml"
    description: "Path to Kubernetes configuration file"

  bootstrapCfgPath:
    type: string
    default: "/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml"
    example: "/opt/pf9/airctl/conf/k3s-bootstrap-config.yaml"
    description: "Path to bootstrap configuration file to deploy the kubernetes cluster"

  # Certificate Configuration
  clusterFQDN:
    type: string
    description: "Cluster fully qualified domain name"

  caCertPath:
    type: string
    description: "Path to CA certificate file"

  caKeyPath:
    type: string
    description: "Path to CA key file"

  certPath:
    type: string
    description: "Path to certificate file"

  certKeyPath:
    type: string
    description: "Path to certificate key file"

  # Component Version Configuration
  consulVersion:
    type: string
    default: "1.2.0"
    description: "Consul version"

  vaultVersion:
    type: string
    default: "0.21.1"
    description: "Vault version"

  kplaneVersion:
    type: string
    default: "0.3.5"
    description: "Kplane version"

  pxcOperatorVersion:
    type: string
    default: "1.16.1"
    description: "Percona XtraDB Cluster Operator version"

  pxcDbVersion:
    type: string
    default: "1.16.1"
    description: "Percona XtraDB Cluster Database version"

  # Features and Options
  enableK8s:
    type: boolean
    default: true
    description: "Enable Kubernetes (K8s/Kaapi) support - default true for On-Prem, false for CE"

  analytics:
    type: boolean
    default: false
    description: "Enable Platform9 telemetry and analytics"

  pcdChartBundleEnabled:
    type: boolean
    default: false
    description: "Enable PCD chart bundle for airgap mode"

  # Timeouts
  svcDeploymentTimeout:
    type: integer
    default: 600
    description: "Service deployment timeout in seconds"

  nodeShutdownTaintTimeout:
    type: integer
    default: 3600
    description: "Node shutdown taint timeout in seconds"

  # Administrative Configuration
  adminUserEmail:
    type: string
    default: "admin@airctl.localnet"
    description: "Admin user email address"

  # NTP Configuration
  ntpServers:
    type: string
    example: "time.google.com time1.google.com"
    description: "Space-separated list of NTP servers for DU time synchronization"

  # Options Path
  optionsPath:
    type: string
    default: "/opt/pf9/airctl/conf/options.json"
    description: "Path to additional options JSON file"

  # Error Collection Configuration (Root Level)
  error_collection_enabled:
    type: boolean
    default: false
    description: "Enable error collection and debugging information upload"

  error_collection_upload_endpoint:
    type: string
    example: "https://ce-support-bundle-dev.s3.us-west-1.amazonaws.com"
    description: "Endpoint URL for uploading error collection data"
# Additional Notes:
# 1. Required fields vary based on installDriver value (DDU vs openstack)
# 2. For custom storage provider, nfsServerIP and nfsShare are required
# 3. Multiple regions can be specified in duRegion as space-separated values
# 4. Version fields may be automatically determined based on airctl version
```

{% endtab %}
{% endtabs %}


# Enable Kubernetes Management Plane

This section explains how to enable the Kubernetes management plane on an existing Self-Hosted <code class="expression">space.vars.product\_name</code> deployment using the `airctl` CLI tool.

## What the Kubernetes Management Plane Is

The Kubernetes management plane runs on the <code class="expression">space.vars.product\_name</code> management cluster and lets you provision, upgrade, scale, and manage <code class="expression">space.vars.product\_name</code> Kubernetes clusters.

For what <code class="expression">space.vars.product\_name</code> Kubernetes provides and how clusters work once this is enabled, see the [<code class="expression">space.vars.product\_name</code> Kubernetes Overview](/private-cloud-director/kubernetes-clusters/k8s-overview).

The Kubernetes management plane is optional and opt-in. If you did not enable it during installation (see [Install](/private-cloud-director/getting-started/self-hosted/self-hosted-install)), you can add it later to a running deployment with the steps below.

{% hint style="info" %}
The base management plane must be fully deployed and healthy before you enable the Kubernetes management plane. To confirm, run `airctl status`.
{% endhint %}

## Enable on an Existing Deployment

Run the following command on one of the management cluster hosts to install the Kubernetes management plane on the existing deployment:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl start --k8s-only
```

{% endtab %}
{% endtabs %}

The `--k8s-only` flag installs only the Kubernetes management plane and leaves the rest of the running management plane untouched.

{% hint style="warning" %}
The Kubernetes management plane uses additional resources on the management plane. Ensure your management cluster nodes meet the recommended sizing before enabling it. See [Pre-requisites](/private-cloud-director/getting-started/self-hosted/self-hosted-pre-requisites).
{% endhint %}

## Verify the Installation

Once the command completes, confirm that the Kubernetes management plane is deployed:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl status
```

{% endtab %}
{% endtabs %}

The **Kubernetes management plane** should report `Ready` in the output.


# Adding a New Region to an Existing Deployment

You can add a new region to an already deployed setup using the `airctl`. This allows you to expand your environment without redeploying the entire system.

### Steps to Add a New Region

1. Update the Configuration File:

Edit the `airctl-config.yaml` file and append the new region name under the `duRegion` field. Example:

{% tabs %}
{% tab title="Bash" %}

```bash
# cat /opt/pf9/airctl/conf/airctl-config.yaml | grep -i region
duRegion: Infra RegionOne NewRegion
```

{% endtab %}
{% endtabs %}

2. Run the add-Region Command:

{% tabs %}
{% tab title="Bash" %}

```bash
airctl add-region --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

3. Verify the Region Addition:

Once the command completes, verify that the new region has been successfully added. You should now see the newly added region reflected in both the **state file** and the **`airctl status`** output.

{% tabs %}
{% tab title="Bash" %}

```bash
airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

Here is the sample output.

{% tabs %}
{% tab title="Bash" %}

```bash
------------- deployment details ---------------
fqdn:                airctl-1-4108764-426.platform9.localnet
region:              airctl-1-4108764-426
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-2561
-------- region service status ----------
desired services:     30
ready services:       30


------------- deployment details ---------------
fqdn:                airctl-1-4108764-426-regionone.platform9.localnet
region:              airctl-1-4108764-426-regionone
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-2561
-------- region service status ----------
desired services:     77
ready services:       77

------------- deployment details ---------------
fqdn:                airctl-1-4108764-426-newregion.platform9.localnet
region:              airctl-1-4108764-426-newregion
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-2561
-------- region service status ----------
desired services:     77
ready services:       77
```

{% endtab %}
{% endtabs %}

{% tabs %}
{% tab title="Bash" %}

```bash
# cat .airctl/state.yaml
...
regions:
    - name: Infra
      fqdn: airctl-1-4108764-426.platform9.localnet
      dbHostname: percona-db-pxc-db-haproxy.foo.svc.cluster.local
      dbPassword: *********
    - name: RegionOne
      fqdn: airctl-1-4108764-426-regionone.platform9.localnet
      dbHostname: percona-db-pxc-db-haproxy.foo-regionone.svc.cluster.local
      dbPassword: *********
    - name: NewRegion
      fqdn: airctl-1-4108764-426-newregion.platform9.localnet
      dbHostname: percona-db-pxc-db-haproxy.foo-regionone.svc.cluster.local
      dbPassword: *********
```

{% endtab %}
{% endtabs %}

This process ensures a seamless addition of new regions to your deployment while maintaining consistency across configuration and runtime state.


# Viewing and Checking Region Status

This section explains how to list all deployed regions in <code class="expression">space.vars.product\_name</code> and check their current deployment status using the `airctl` CLI tool.

## List All Deployed Regions

To view the list of all regions currently deployed in the Self-Hosted version of <code class="expression">space.vars.product\_name</code> , run:

{% tabs %}
{% tab title="Bash" %}

```bash
/opt/pf9/airctl/airctl get-regions --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

This command retrieves region names directly from the state file and lists all currently active regions. Use it to confirm which regions are available before performing operations such as backups, deletions, or status checks.

Example output:

{% tabs %}
{% tab title="Bash" %}

```bash
$ /opt/pf9/airctl/airctl get-regions --config /opt/pf9/airctl/conf/airctl-config.yaml 
Deployed Regions:
• Infra
• LON
```

{% endtab %}
{% endtabs %}

## Check Region Status

The `airctl status` command provides detailed information about the health and service status of each region.

By default, it shows the status of all deployed regions, including service counts and readiness indicators.

### Status Command Flags

The `airctl status` command supports the following flags to customize output:

* \--region : Specifies which region’s status to display. **This flag is mandatory** when using `--ready` or `--desired`.
* \--ready: Displays **only the number of Ready services** for the specified region.
* \--desired: Displays **only the number of Desired services** for the specified region

{% hint style="info" %}
**Info**

* The `--ready` and `--desired` flags are **mutually exclusive** — only one can be used at a time.
* When used, these flags print **only the number of services** in the specified state for the given region.
  {% endhint %}

{% tabs %}
{% tab title="Bash" %}

```bash
Usage:
  airctl status [flags]

Flags:
      --desired         Show only the number of desired services for each region
  -h, --help            help for status
      --ready           Show only the number of ready services for each region
  -r, --region string   Display status for a specific region

Global Flags:
      --config string   config file (default is $HOME/airctl-config.yaml)
      --json            json output for commands (configure-hosts only currently)
      --quiet           disable spinners
      --verbose         print verbose logs to the console
```

{% endtab %}
{% endtabs %}

### Examples:

View All Regions and their Status :

{% tabs %}
{% tab title="Bash" %}

```bash
$ /opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml
------------- deployment details ---------------
fqdn:                airctl-1-4161345-238.platform9.localnet
region:              airctl-1-4161345-238
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-69
-------- region service status ----------
desired services:     30
ready services:       30


------------- deployment details ---------------
fqdn:                airctl-1-4161345-238-lon.platform9.localnet
region:              airctl-1-4161345-238-lon
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-69
-------- region service status ----------
desired services:     84
ready services:       84
```

{% endtab %}
{% endtabs %}

View status for a specific region :

{% tabs %}
{% tab title="Bash" %}

```bash
$ /opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml --region Infra
------------- deployment details ---------------
fqdn:                airctl-1-4161345-238.platform9.localnet
region:              airctl-1-4161345-238
deployment status:   ready
region health:       ✅ Ready
version:              PCD 2025.10-69
-------- region service status ----------
desired services:     30
ready services:       30
```

{% endtab %}
{% endtabs %}

View Ready Services for a Region :

{% tabs %}
{% tab title="Bash" %}

```bash
$ /opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml --ready --region LON
84
```

{% endtab %}
{% endtabs %}

View Desired Services for a Region :

{% tabs %}
{% tab title="Bash" %}

```bash
$ /opt/pf9/airctl/airctl status --config /opt/pf9/airctl/conf/airctl-config.yaml --desired --region LON
84
```

{% endtab %}
{% endtabs %}


# Uninstall Self-Hosted Deployment

This section describes how to safely delete a **region** from <code class="expression">space.vars.self\_hosted\_product\_name</code> using the `airctl` CLI. Region deletion should be performed carefully, as it involves removing associated configuration and metadata.

### Prerequisites & Safety Checks <a href="#prerequisites--safety-checks" id="prerequisites--safety-checks"></a>

Before proceeding with a region deletion, ensure the following:

1. **Take a complete backup** using the steps outlined in the [Backup Guide](https://platform9.com/docs/private-cloud-director/private-cloud-director/backup-and-restore#manual-backup-procedure).
   * This ensures that critical region data and configurations can be restored if needed.
2. **Verify the current regions** managed by your setup: `airctl get-regions`
   * Review the output and confirm the target region name(s) you plan to delete.
3. Ensure no ongoing deployments, upgrades, or workloads are running in the target region.
4. Gracefully terminate or migrate workloads if needed.

Example :

```bash
$ /opt/pf9/airctl/airctl get-regions --config /opt/pf9/airctl/conf/airctl-config.yaml 
Deployed Regions:
• Infra
• LON
```

### Command Syntax <a href="#command-syntax" id="command-syntax"></a>

#### Delete a specific region

To delete a specific region, use the following `airctl` command

```bash
airctl unconfigure-du \
  --config /opt/pf9/airctl/conf/airctl-config.yaml \
  --region <REGION_NAME>
```

This command will uninstall and remove configured regions along with all infrastructure software like consul, vault, percona, k8sniff, etc.

Options:

* `--region <REGION_NAME>` – Deletes the specified region.
* **Omit the `--region` flag** to delete *all* configured regions.
* Use the `--force` flag if the region deployment was halted or failed midway. This forces cleanup even when the state is inconsistent.

Use the `--force` flag if the region deployment was halted or failed midway. This forces cleanup even when the state is inconsistent.

{% hint style="info" %}
Use the `--force` flag if the region deployment was halted or failed midway. This forces cleanup even when the state is inconsistent.
{% endhint %}

Example:

```bash
airctl unconfigure-du --config /opt/pf9/airctl/conf/airctl-config.yaml --region LON
```

#### Verification & Cleanup <a href="#verification--cleanup" id="verification--cleanup"></a>

After deletion, verify that the region has been successfully removed. The deleted region should no longer appear in the list of active regions.

```bash
$ /opt/pf9/airctl/airctl get-regions --config /opt/pf9/airctl/conf/airctl-config.yaml 
Deployed Regions:
• Infra
```

If you plan to reuse the same nodes to deploy a new self-hosed <code class="expression">space.vars.product\_name</code> environment, make sure to also run the following command on all nodes first.

```bash
rm -rf airctl* install-pcd.sh options.json pcd-chart.tgz /opt/pf9/airctl/ .airctl/
```

#### Delete the Management Cluster <a href="#delete-the-management-cluster" id="delete-the-management-cluster"></a>

The final step is to delete the cluster to complete the cleanup or redeploy from scratch.

```bash
/opt/pf9/airctl/airctl delete-cluster --config /opt/pf9/airctl/conf/airctl-config.yaml
```


# Airctl Support Bundle

A support bundle is a comprehensive archive created specifically for the purpose of submitting pertinent information to the Platform9 support team.

The support bundle mechanism is designed to be pluggable, allowing for the easy addition of commands, extra logs, or other relevant data to be captured during the archiving process. The steps outlined below will guide you through this procedure.

## Management Plane

For the Management Plane (k3s cluster), the command `airctl gen-support-bundle` is instrumental in generating a comprehensive support bundle. This bundle contains crucial diagnostic information that can be utilized for troubleshooting and support purposes.

Options specific to the `gen-support-bundle` command are:

{% tabs %}
{% tab title="Bash" %}

```bash
Usage:
  airctl gen-support-bundle [flags]

Flags:
      --all               collect logs from all the hosts in the airctl config file
      --cluster-dump      retrieve cluster-info dump
      --cluster-state     Get cluster state (default true)
  -h, --help              help for gen-support-bundle
      --host-ips string   collect host side logs. Enter unique host ip's in comma-separated format.                           Example: "10.128.243.113, 10.128.242.112"
      --out string        the output file name (default "/tmp/supportbundle.tgz")

Global Flags:
      --config string   config file (default is $HOME/airctl-config.yaml)
      --json            json output for commands (configure-hosts only currently)
      --quiet           disable spinners
      --verbose         print verbose logs to the console
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

This process will collect logs from all nodes within the management plane cluster, ensuring comprehensive data retrieval for analysis and troubleshooting.

`airctl gen-support-bundle` command should be executed from a Management plane host
{% endhint %}

To generate the support bundle, you can use the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo /opt/pf9/airctl/airctl gen-support-bundle --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

By default, the archived contents will be located in the file `/tmp/supportbundle.tgz`. You can share this archived file with the Platform9 support team for further assistance.

To specify a custom path for the archived support bundle, utilize the `--out` option. This allows you to define the exact location where the support bundle will be saved, ensuring that it is easily accessible for future reference or analysis.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo /opt/pf9/airctl/airctl gen-support-bundle --out <PATH_NAME>.tgz --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endtab %}
{% endtabs %}

By default, the support bundle collects essential logs from the nodes of the management plane cluster specified in the `bootstrapCfgPath` configuration file.

Support bundle logs includes:

* `dduLogs` : The management plane cluster's (Kubernetes control plane) service logs from the pod/container on the node it's running.
* `/var/log/pf9`: Logs related to the Platform9 services and Kubernetes control plane services running on management plane cluster.
* `kubectl get [pods|nodes|events]` outputs.
* Per-node `airctl-k3s-support.*/` archives — one per management plane node — containing the k3s and containerd logs (`logs/k3s.log`, `logs/containerd.log`, `logs/k3s.journal.log`), k3s systemd unit state (`systemd/k3s.units.txt`, `systemd/k3s.status.txt`), node-local k3s configuration (`config/config.yaml`, `config/k3s.yaml`), and DDU fluentbit logs (`dduLogs/fluentbit/`).
* A `k8s-diagnostics/` directory with a snapshot of Kubernetes API state: `cluster-info.txt`, `events-recent.txt`, `nodes.yaml`, `namespaces.yaml`, `version.yaml`, plus `raw-livez_verbose.txt` and `raw-readyz_verbose.txt` from the `kube-apiserver` health endpoints.

{% hint style="info" %}
**Info**

`-- cluster-dump` option is not functional at the moment.

If the Platform9 support team requests the entire management plane cluster logs, including logs from management plane components, along with the support bundle, please follow the steps below after exporting the KUBECONFIG:

`$ kubectl cluster-info dump --all-namespaces --output-directory=/path/to/dump`

Cluster Dump a set of namespaces

`$ kubectl cluster-info dump --namespaces default,kube-system --output-directory=/path/to/dump`

Next, compress this directory along with the support bundle into a tar archive.
{% endhint %}

## Host Support Bundle

Generate a support bundle for hosts that are connected or authorized to the Management Plane, which we need to troubleshoot the hosts in case of any issues.

Steps for collecting the support bundle from hosts are outlined here: [Hosts Support Bundle](/private-cloud-director/reference/pcdctl-command-line#support-bundle).


# Troubleshooting And Log Files

This section covers log file locations and steps to troubleshoot issues with your self-hosted install.


# Pcddebugger Utility

Learn about PCDdebugger - a command-line Python script designed to simplify and accelerate PCD troubleshooting

## PCDdebugger

PCDdebugger is a command-line Python script designed to simplify and accelerate the troubleshooting process for **PCD Virtualization** environments. It automates the collection of diagnostic information for various services and resources, consolidating the output into a structured directory for easy analysis and sharing.

## Features

* **Comprehensive Data Collection:** Gathers detailed information for key PCD resources including VMs (Nova), images (Glance), networks (Neutron), ports, volumes (Cinder), stacks (Heat), and users (Keystone).
* **Kubernetes Integration:** Performs a complete MySQL dump from a specified Kubernetes namespace, essential for debugging control plane issues.
* **Dependency Traversal:** Automatically discovers and collects data for resources related to a specified VM, such as its ports, volumes, network, subnets, image, and flavor.
* **Organized Output:** Saves all collected data into a timestamped directory, with subfolders for each service, making the information easy to navigate.
* **Archive Option:** Includes a --zip flag to automatically create a compressed archive of the collected data, ready for sharing.

## Prerequisites

Before running the script, ensure the following are installed and configured on your machine:

* **pcd client:** Authenticated and configured to connect to your PCD cloud. (**Ensure your rc file is sourced**).
* **kubectl:** Authenticated and configured to connect to your Kubernetes cluster. (This is required only in case of mysql dump)

## Installation

{% stepper %}
{% step %}

#### Mac-OS

* Download the binary:

{% code title="Command" %}

```bash
curl -LO https://github.com/platform9/PCDDebugger/releases/download/v2.0.1/PCDdebugger-v2.0.1-macos
```

{% endcode %}

* Rename the binary file (Optional):

{% code title="Command" %}

```bash
mv PCDdebugger-v2.0.1-macos PCDdebugger
```

{% endcode %}

* Make the script executable:

{% code title="Command" %}

```bash
chmod +x ./PCDdebugger
```

{% endcode %}

* To run PCDdebugger from any directory, move it to a location in your system's PATH (Optional):

{% code title="Command" %}

```bash
sudo mv PCDdebugger /usr/local/bin/
```

{% endcode %}

{% hint style="info" %}
Mac-OS will ask to allow the permission to run the PCDDebugger binary. Allow it from the System Settings → Privacy and Security.
{% endhint %}
{% endstep %}

{% step %}

#### Linux

* Download the binary:

{% code title="Command" %}

```bash
curl -LO https://github.com/platform9/PCDDebugger/releases/download/v2.0.1/PCDdebugger-v2.0.1-linux
```

{% endcode %}

* Rename the binary file (Optional):

{% code title="Command" %}

```bash
mv PCDdebugger-v2.0.1-linux PCDdebugger
```

{% endcode %}

* Make the binary executable:

{% code title="Command" %}

```bash
chmod +x PCDdebugger
```

{% endcode %}

* To run PCDdebugger from any directory, move it to a location in your system's PATH (Optional):

{% code title="Command" %}

```bash
sudo mv PCDdebugger /usr/local/bin/
```

{% endcode %}
{% endstep %}

{% step %}

#### Windows

* Download the binary:

{% code title="Powershell" %}

```powershell
curl -L https://github.com/platform9/PCDDebugger/releases/download/v2.0.1/PCDdebugger-v2.0.1-windows.exe -o PCDdebugger.exe
```

{% endcode %}

* Place it in a folder (example): Move the downloaded `PCDdebugger.exe` file to a memorable location, for example, C:\Tools.
* (Optional) Add to PATH so you can run it from any command prompt:
  1. Search for "Edit the system environment variables" in the Start Menu.
  2. Click the "Environment Variables..." button.
  3. Under "System variables", find and select the Path variable, then click "Edit...".
  4. Click "New" and add the path to the folder (e.g., C:\Tools).
  5. Click OK on all windows to save.
* Run `PCDdebugger.exe` from PowerShell or Command Prompt.
  {% endstep %}
  {% endstepper %}

## Usage

The basic command structure is:

{% tabs %}
{% tab title="command" %}
{% code title="Command" %}

```bash
./PCDdebugger [RESOURCE_FLAG] [OPTIONS] [--insecure]
```

{% endcode %}
{% endtab %}
{% endtabs %}

{% hint style="info" %}
The "--insecure" parameter is mandatory if SSL is not configured; otherwise it will throw an "SSL Certificate" error.
{% endhint %}

### Examples

* Collect all information for a specific VM (gathers the VM, its ports, volumes, network, subnets, image, and flavor):

{% code title="Command" %}

```bash
./PCDdebugger --vm <VM_UUID>
```

{% endcode %}

* Collect details for a specific Glance image:

{% code title="Command" %}

```bash
./PCDdebugger --image <IMAGE_UUID>
```

{% endcode %}

* Collect details for a Neutron network and its subnets:

{% code title="Command" %}

```bash
./PCDdebugger --network <NETWORK_UUID>
```

{% endcode %}

* Perform a MySQL dump from a Kubernetes cluster (requires --namespace):

{% code title="Command" %}

```bash
./PCDdebugger --mysql-dump --namespace <WORKLOAD_REGION>
```

{% endcode %}

{% hint style="info" %}
The "--mysql-dump" parameter can only be run on the Self Hosted PCD Virtualization.
{% endhint %}

* Combine multiple flags and create a zip archive:

{% code title="Command" %}

```bash
./PCDdebugger --vm <VM_UUID> --mysql-dump --namespace <K8S_NAMESPACE> --zip
```

{% endcode %}

* Specify a custom output directory and create a zip archive:

{% code title="Command" %}

```bash
./PCDdebugger --vm <VM_UUID> --output <PATH_TO_DIRECTORY> --zip
```

{% endcode %}

Example:

{% code title="Command" %}

```bash
./PCDdebugger --vm aa1bb1cc3-dd4ee5ff6-gg7hh8-ii9jj10kk --output /tmp --zip
```

{% endcode %}

Command-Line Flags:

* `--vm <ID_OR_NAME>`: Collect details for a specific Nova VM and its related resources.
* `--image <ID_OR_NAME>`: Collect details for a specific Glance image.
* `--network <ID_OR_NAME>`: Collect details for a specific Neutron network and its subnets.
* `--port <ID_OR_NAME>`: Collect details for a specific Neutron port.
* `--volume <ID_OR_NAME>`: Collect details for a specific Cinder volume.
* `--stack <ID_OR_NAME>`: Collect details for a specific Heat stack.
* `--user <ID_OR_NAME>`: Collect details for a specific Keystone user.
* `--mysql-dump`: Perform a MySQL dump of all databases. **Requires** `--namespace`.
* `--namespace <NAMESPACE>`: The Kubernetes namespace to use for the MySQL dump.
* `--output <DIRECTORY>`: Specify a custom directory for the output files.
* `--zip`: Compress the final output directory into a `.zip` file.
* `--help`: Show the help message and exit.

{% hint style="info" %}
All output is saved to a directory named PCDdebugger- by default.
{% endhint %}

## Additional Information

For more details, visit the GitHub page: <https://github.com/platform9/PCDDebugger>


# Collecting Cluster Dump For Management Cluster

Learn about pcddump - a powerful utility for collecting comprehensive PCD Management cluster information for offline troubleshooting.

**PCD Dump** is a powerful utility for collecting comprehensive PCD Management cluster information for offline troubleshooting. This script gathers detailed cluster dump information from <code class="expression">space.vars.self\_hosted\_product\_name</code> Control Plane, providing administrators with comprehensive diagnostics and troubleshooting data for effective cluster management and issue resolution.

PCD Dump is available here - <https://github.com/platform9/PCDDump>

## Prerequisites

* Internet connectivity to download the script.
* `curl` installed
* `kubconfig` should be exported and `kubectl` installed with sufficient permissions:

{% tabs %}
{% tab title="Command" %}
{% code title="Verify kubeconfig and internet connectivity" %}

```bash
export KUBECONFIG=</path/to/your/pcd-management-cluster-kubeconfig.yaml>

# Verify connectivity
kubectl get nodes
```

{% endcode %}
{% endtab %}
{% endtabs %}

## Installation

{% hint style="warning" %}
Before running the script, ensure that the prerequisites are met.
{% endhint %}

* Run this script to initiate the PCD cluster dump generation:

{% tabs %}
{% tab title="Quick Run" %}
{% code title="One-line quick run" %}

```bash
bash <(curl -Ls https://raw.githubusercontent.com/platform9/PCDDump/refs/heads/main/pcddump.sh)
```

{% endcode %}
{% endtab %}
{% endtabs %}

* For manual execution:

{% tabs %}
{% tab title="Bash" %}
{% code title="Manual download and run" %}

```bash
# Download the script
curl -L https://raw.githubusercontent.com/platform9/PCDDump/refs/heads/main/pcddump.sh -o pcddump.sh

# Make it executable
chmod +x pcddump.sh

# Execute the script
./pcddump.sh
```

{% endcode %}
{% endtab %}
{% endtabs %}

{% hint style="info" %}
The tar output file is saved under /tmp as /tmp/pcddump-$(date +%F\_%H-%M-%S).tar.gz
{% endhint %}

## Additional Information

You can upload the `/tmp/pcddump-$(date +%F_%H-%M-%S).tar.gz` file as suggested by the Platform9 support team. More information in <https://github.com/platform9/PCDDump>


# Pods Are Displaying FailedAttachVolume or FailedMount Errors

You may run into a problem where the PCD management plane pods are not functioning or running due to PVC mount failures, specifically displaying "FailedAttachVolume" or "FailedMount" errors.

{% tabs %}
{% tab title="Pod Event Section" %}

```ruby
Warning  FailedMount         5h32m (x31004 over 6h27m)  kubelet                  MountVolume.WaitForAttach failed for volume "pvc-[pvc-id]" : volume attachment is being deleted

Warning  FailedAttachVolume  5h30m (x36 over 6h27m)     attachdetach-controller  AttachVolume.Attach failed for volume "pvc-[pvc-id]" : volume attachment is being deleted
```

{% endtab %}
{% endtabs %}

## Most common causes

* The storage backend is unreachable.
* The underlying host does not have sufficient resources to run these pods.
* CSI Driver itself is not configured correctly or has some errors.
* Calico network pods are not working as expected.

## Steps to Troubleshoot

{% stepper %}
{% step %}

#### Check pods in init state

Check the number of pods in the `Init` state to identify any pod stuck in initialization. The failure of PVC mount attachments can cause pods to remain in an `Init` state.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl get pods -A | grep -i "init"
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Get CSI drivers

Run the following command to list CSI drivers (storage backend).

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl get csidrivers
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Verify CSI driver pods are running

Verify if the CSI driver pods are running. The pods can either be in a dedicated namespace or inside the `kube-system` namespace. In this example the NetApp backend uses the `trident` namespace to host its storage backend pods.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl get pods -n <CSI-driver-namespace>/<kube-system>

E.g.
$ kubectl get pods -n trident
trident           trident-controller-pod    0/6     ContainerCreating   0                 6h44m
trident           trident-node-linux-pod    0/2     CrashLoopBackOff    20 (5h33m ago)    23m
trident           trident-node-linux-pod    0/2     CrashLoopBackOff    15 (5d4h ago)     23d
trident           trident-node-linux-pod    0/2     CrashLoopBackOff    34 (5d3h ago)     23d
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Inspect Calico pods

As Calico provides pod networking, review all Calico pods and determine why these pods are in a "CrashLoopBackOff/ContainerCreating/OOMKilled/Pending/Error" state; check the events from the describe output.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl describe <pod-name> -n <calico-namespace>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Check logs for failing pods

Get more information on the failure from the pod logs.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl logs <pod-name> -n <CSI-driver-namespace>/<kube-system>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### When to escalate

If these steps don't resolve the issue, please contact your Backend Storage Provider or reach out to the Platform9 Support Team for additional assistance.

* Platform9 Support Team: <https://support.platform9.com/hc/en-us>
  {% endstep %}
  {% endstepper %}


# Delete Orphaned Virtual Machine Entries

When a virtual machine is deleted / migrated / evacuated, sometimes its corresponding allocation on the source hypervisor host is not deleted from the PCD database.

## Root Cause

* If the compute service on the source hypervisor host is stopped abruptly while a VM is being deleted / migrated / evacuated, then the records for the VM in the PCD management database may not be deleted fully.
* As a result, the host-side service cannot communicate with nova-conductor on the controller, so details are not shared and nova-compute may remain under the impression it still holds the virtual machine.

## Resolution

Use the following steps to locate and remove orphaned allocations. Run the commands from a hypervisor host.

{% stepper %}
{% step %}

#### 1. List stale allocations

Run nova-manage placement audit to list allocations for virtual machines that are either deleted or moved to other hypervisor hosts.

{% code title="List stale allocations (run from a pod with nova-manage)" %}

```bash
kubectl exec -it deploy/nova_api_osapi -n <NS> -- bash
nova-manage placement audit --verbose
```

{% endcode %}

Verify the output to identify suspected orphaned allocations.
{% endstep %}

{% step %}

#### 2. Verify VM existence / location

Confirm whether the VM still exists or is hosted on another hypervisor host.

{% code title="List servers on source host" %}

```bash
pcdctl server list --all-projects --host <Source Host>
```

{% endcode %}

{% code title="Show instance host" %}

```bash
pcdctl server show -c OS-EXT-SRV-ATTR:host <instance-ID>
```

{% endcode %}
{% endstep %}

{% step %}

#### 3. Identify the resource provider ID (if needed)

List resource providers to find the provider associated with the allocation.

{% code title="List resource providers" %}

```bash
pcdctl resource provider list
```

{% endcode %}
{% endstep %}

{% step %}

#### 4. Delete a single orphaned allocation

After validating the allocation is stale and the instance is deleted or moved to another compute node, delete the orphaned allocation.

{% code title="Delete specific allocation" %}

```bash
kubectl exec -it <nova_api_osapi-pod_name> -n <NS> -- bash
nova-manage placement audit --verbose --delete <instance-ID>
```

{% endcode %}
{% endstep %}

{% step %}

#### 5. Delete multiple/all orphaned allocations

If there are multiple orphaned allocations, delete them all at once.

{% code title="Delete all orphaned allocations" %}

```bash
nova-manage placement audit --verbose --delete
```

{% endcode %}
{% endstep %}

{% step %}

#### 6. Heal allocations

After deletions, run heal\_allocations to ensure Placement entries are consistent.

{% code title="Heal allocations" %}

```bash
nova-manage placement heal_allocations
```

{% endcode %}
{% endstep %}

{% step %}

#### 7. Validate there are no more orphaned allocations

Re-run the audit to ensure no orphaned allocations remain.

{% code title="Validate no orphaned allocations remain" %}

```bash
nova-manage placement audit --verbose
```

{% endcode %}
{% endstep %}
{% endstepper %}


# Using Custom Certificates

Starting from this release, you can configure custom SSL/TLS certificates for your Self-hosted Private Cloud Director deployment. Previously, the system only used self-signed certificates generated during the deployment process.

### Overview

By default, the system generates self-signed certificates during installation. You can now:

* Use your own custom-signed certificates
* Continue using automatically generated self-signed certificates

You can apply custom certificates during a new installation or update existing deployments using the `renew-certs` command.

### Configure Custom Certificates During Installation

Prerequisites

* A valid SSL/TLS certificate file (`.crt`)
* A corresponding private key file (`.key`)
* Read access to both files for the user running the configuration

{% stepper %}
{% step %}

#### Prepare your certificate files

Place your certificate and key files in an accessible location on your system.

Example:

```bash
service.crt   # Certificate file
service.key   # Private key file
```

{% endstep %}

{% step %}

#### Specify certificate paths

You can provide the certificate and key paths using either of these methods.

Option A: Export as environment variables

```bash
export USER_CERT_PATH="_path_to_service.crt"
export USER_KEY_PATH="_path_to_service.key"
```

Option B: Pass as command-line arguments

```bash
airctl configure {other_flags} \
--user-cert-path _path_to_service.crt \
--user-key-path _path_to_service.key
```

{% endstep %}

{% step %}

#### Verify the configuration

After you run the configuration command, verify that the certificate paths appear in the configuration file.

Check `/opt/pf9/airctl/conf/airctl-config.yaml`:

```yaml
...
user_cert_path: _path_to_service.crt
user_key_path: _path_to_service.key
...
```

Expected outcome: The configuration file contains your specified certificate paths.
{% endstep %}

{% step %}

#### Updating the certs

Run the renew-certs command which will update the certs:

```bash
airctl renew-certs --config _opt_pf9_airctl_conf_airctl-config.yaml
```

{% endstep %}
{% endstepper %}

### Update Certificates on Existing Deployments

If you have an existing PCD deployment, you can replace the current certificates using the `airctl renew-certs` command.

Prerequisites

* You have an existing PCD deployment
* You have the new certificate and key files ready
* You have access to the configuration file at `/opt/pf9/airctl/conf/airctl-config.yaml`

{% stepper %}
{% step %}

#### Update the configuration file

Set the new certificate paths using one of these methods:

* Export environment variables (as shown in the installation section)
* Use the `airctl configure` command with certificate path arguments
* Edit `/opt/pf9/airctl/conf/airctl-config.yaml` directly

> Keep all other configuration fields unchanged.
> {% endstep %}

{% step %}

#### Verify the updated configuration

Check that `/opt/pf9/airctl/conf/airctl-config.yaml` reflects the new certificate paths:

```yaml
certPath: _path_to_new_service.crt
certKeyPath: _path_to_new_service.key
```

{% endstep %}

{% step %}

#### Renew the certificates

Run the certificate renewal command:

```bash
airctl renew-certs --config _opt_pf9_airctl_conf_airctl-config.yaml
```

Expected outcome: The command updates the certificates in your deployment. You can switch from custom to self-signed certificates or from self-signed to custom certificates.
{% endstep %}
{% endstepper %}

{% hint style="info" %}
Important Notes

* Ensure that your `DU_FQDN` environment variable or the `duFqdn` field in the airctl configuration file matches the domain specified in your certificates.
* The user running configuration commands must have read access to the certificate files.
* You can switch between custom and self-signed certificates at any time using the `airctl renew-certs` command.
  {% endhint %}


# Teleport Installation

This guide provides comprehensive instructions for installing Teleport agents on Platform9 Private Cloud Director environments.

## Introduction

Teleport enables secure, role-based access to Kubernetes clusters and individual nodes for debugging, monitoring, and management operations. It is a solution for:

* **Secure Remote Access**: SSH and web-based access to cluster resources
* **Role-Based Access Control**: Granular permissions for different user roles
* **Audit and Compliance**: Full logging and session recording
* **Cross-Platform Access**: Unified access to multiple environments

### Prerequisites

* Self-hosted Private cloude Directer is installed and configured
* kubectl access to the **Management cluster**
* Valid Teleport join token from Platform9 Support

#### Installation Options

#### Option 1: Kubernetes Agent Installation

Install Teleport agent into a Kubernetes namespace for namespace-scoped access.

```bash
./install-teleport.sh --token <JOIN_TOKEN> \
  --namespace <NAMESPACE> \
  --proxy-addr <PROXY_ADDR> \
  --cluster-name <CLUSTER_NAME> \
  --version <VERSION> \
  --access-namespace <NAMESPACE[,NAMESPACE]> \
  --access-level <admin|readonly> \
  --resource kube
```

**Parameters:**

* `--token`: Join token (required)
* `--namespace`: Target namespace where the Teleport agent pod runs (default: `teleport`)
* `--proxy-addr`: Teleport proxy address (default: `platform9.teleport.sh:443`)
* `--cluster-name`: name for identification in Teleport interface(default: `example-k8s-cluster`)
* `--version`: Teleport version (default: `18.1.4`)
* `--access-namespace`: Namespace(s) to grant access (comma-separated)
* `--access-level`: Access level - `admin` or `readonly` (default: `admin`)
* `--resource`: Resource type - `kube` (default)

**Examples:**

```bash
# Basic installation with admin access to teleport namespace
./install-teleport.sh --token abc123 --namespace my-app --access-level admin

# Cluster-wide admin access
./install-teleport.sh --token abc123 --access-level admin

# Read-only access to specific namespaces
./install-teleport.sh --token abc123 --access-namespace app1,app2 --access-level readonly

# Custom namespace and cluster name
./install-teleport.sh --token abc123 --namespace production --cluster-name prod-k8s-cluster
```

#### Option 2: Node Installation

Install Teleport node agent on individual Kubernetes nodes for node-level access.

```bash
./install-teleport.sh --token <JOIN_TOKEN> \
  --proxy-addr <PROXY_ADDR> \
  --version <VERSION> \
  --resource node
```

**Parameters:**

* `--token`: Join token (required)
* `--proxy-addr`: Teleport proxy address (default: `platform9.teleport.sh:443`)
* `--version`: Teleport version (default: `18.1.4`)
* `--resource`: Resource type - `node` (required)

**Example:**

```bash
# Install Teleport on a worker node
./install-teleport.sh --token abc123 --resource node
```

### Access Levels

#### Admin Access

* **Cluster Role**: `cluster-admin` (full cluster access)
* **Namespace Role**: `admin` (full namespace access)
* **Capabilities**: Create, read, update, delete resources

#### Read-Only Access

* **Cluster Role**: `teleport-view-only` (cluster-wide read-only)
* **Namespace Role**: `teleport-namespace-view` (namespace read-only)
* **Capabilities**: Get, list, watch resources only

### What Gets Installed

#### Kubernetes Agent

* **Helm Release**: `teleport-agent` in specified namespace
* **RBAC**: ClusterRole and RoleBindings for specified access level
* **Service Account**: Teleport service account with appropriate permissions
* **Proxy Configuration**: Automatic connection to Platform9 Teleport proxy

#### Node Agent

* **System Service**: Teleport node service installed and started
* **Configuration**: Node configured with join token and proxy settings
* **System Integration**: Integrated with systemd for automatic startup

### Proxy Configuration

The script automatically configures connection to Platform9's Teleport proxy:

* **Default Proxy**: `platform9.teleport.sh:443`
* **Custom Proxy**: Use `--proxy-addr` for alternative proxy
* **Environment Variables**: HTTP\_PROXY, HTTPS\_PROXY, NO\_PROXY automatically configured

### Security Considerations

#### Token Security

* Join tokens are single-use and expire after first use
* Store tokens securely and never commit to version control
* Obtain fresh tokens for each installation

#### Access Control

* **Principle of Least Privilege**: Use read-only access when possible
* **Namespace Scoping**: Limit access to specific namespaces when no cluster-wide access is needed

#### Network Security

* **Proxy Only**: All Teleport connections go through Platform9 proxy
* **TLS Encryption**: All communications encrypted via TLS
* **Firewall Rules**: Ensure outbound connections to proxy are allowed

### Troubleshooting

#### Common Issues

**Token Issues:**

```bash
# Verify token format
echo $JOIN_TOKEN | wc -c  # Should be reasonable length

# Check token expiration
# Contact Platform9 Support if token fails
```

**Connection Issues:**

```bash
# Test proxy connectivity
curl -I https://platform9.teleport.sh:443

# Check network connectivity
ping platform9.teleport.sh
```

#### Log Locations

* **Kubernetes Agent**: Check pod logs: `kubectl logs -n <NAMESPACE> teleport-agent-xxx`
* **Node Agent**: Check system logs: `sudo journalctl -u teleport -f`
* **Installation Logs**: Script outputs detailed progress and error messages

#### Cleanup

**Remove Kubernetes Agent:**

```bash
helm uninstall teleport-agent -n <NAMESPACE>
kubectl delete rolebinding,role -n <NAMESPACE> -l app.kubernetes.io/name=teleport
```

**Remove Node Agent:**

```bash
sudo systemctl stop teleport
sudo systemctl disable teleport
sudo rm -f /etc/teleport.yaml
```

### Advanced Configuration

#### Multiple Namespace Access

Grant access to multiple namespaces with comma-separated list:

```bash
./install-teleport.sh --token abc123 \
  --access-namespace dev,staging,production \
  --access-level readonly
```

This creates RoleBindings in each specified namespace with read-only permissions.

### Support

For issues or questions:

* Contact Platform9 Support
* Verify network connectivity to Teleport proxy


# Availability Zones

Starting from this release, you can configure availability zones for your Private Cloud Director deployment.

### Overview

By default, the nodes were neither labeled nor was the workload distributed across them. With this enhancement, nodes can now be labeled through a YAML configuration, ensuring workloads are scheduled correctly based on geo-location or any other deployment-specific requirements.

This also improves high availability by ensuring that workloads are evenly distributed across all zones during deployment. As a result, even if certain nodes fail, the services remain available and continue operating without disruption.

### Configure AZ's

The AZ's need to be configured before setup creation while running the configure command

```bash
airctl configure \ 
  --master-ips <master-ips> \ 
  --master-vip4 <master-vip4> \ 
  --master-vip-interface <master-vip-interface> \ 
  --ipv4 <type> \ 
  --external-ip4 <external-ip4> \ 
  --du-fqdn <fqdn> \ 
  --storage-provider <storage-provider> \ 
  --node-az-labels /path/to/node-az-labels.yaml
```

The node-az-labels.yaml file should have values in this type

```yaml
# node IP: zone name example
10.10.11.173: asia
10.10.15.162: asia
10.10.15.173: europe
10.10.15.63: europe
10.10.20.11: us
```

Once the above steps are completed, create the cluster. You will then be able to observe that the workload is distributed across the nodes as expected.


# Platform9 OS Beta Installation

{% hint style="warning" %}
**Platform9 OS is currently in closed beta.** Access is by invitation only. To request an invitation or learn more, contact your Platform9 account manager.
{% endhint %}

## Overview

Platform9 OS is powered by Rocky Linux by CIQ (RLC) and <code class="expression">space.vars.product\_name</code>. It provides a seamless way to deploy the operating system and <code class="expression">space.vars.product\_acronym</code> dependencies in a single workflow using an ISO image. No separate OS installation or manual agent setup is required.

Customers interested in joining the beta program should contact their Platform9 account manager.

## Prerequisites

All prerequisites noted on the [hypervisor configuration prerequisites page](/private-cloud-director/getting-started/pre-requisites#general-hypervisor-configuration-pre-requisites) continue to apply to this installation mode.

## Installation walkthrough

### Step 1: Boot from the ISO

Boot your target server from the Platform9 OS ISO. The Rocky Linux by CIQ (RLC) 10.1 installer screen appears, showing the **Installation Summary** view organized into four categories: **Localization**, **Software**, **System**, and **User Settings**.

Select **Installation Destination** to configure the target disk. On the **Installation Destination** screen, choose the target disk under **Local Standard Disks**. Optionally configure **Storage Configuration** — select **Automatic** for a standard layout or **Custom** to define your own partition scheme. Click **Done** and then **Begin Installation** to start the process.

### Step 2: OS installation and reboot

The installer displays installation progress, including post-installation script execution. When the process completes, a **Reboot System** prompt appears. Allow the system to reboot.

### Step 3: PCD Installer wizard

After the reboot, a four-step text-based installer titled <code class="expression">space.vars.product\_name</code> **Installer** launches automatically. Navigate using the keys shown in the footer: `Tab` to navigate between fields, `Space` to edit a selection, `Enter` to confirm, `Esc` to go back, and `Ctrl+C` to exit.

#### Network Configuration (Step 1 of 4)

Configure the network interfaces that this host will use. You can optionally bond multiple interfaces for redundancy using the **Create Bond** action.

If your environment requires an outbound proxy to reach the management plane, enable the **Outbound Proxy** field and enter the proxy address in the format `http://<user>:<password>@<host>:<port>`.

After configuring your network settings, run the built-in network test to verify connectivity before proceeding.

#### Root user password (Step 2 of 4)

Enter and confirm a root user password. This password is used for direct server administration access.

#### GPU Configuration (Step 3 of 4)

This step is optional. If the host has GPUs installed, you can configure them now. The following options are available:

* **Skip GPU configuration** — do not configure GPUs during this installation.
* **PCI Passthrough** — pass discrete GPUs directly to guest VMs for exclusive, bare-metal GPU access.
* **vGPU pre-configure** — prepare the host for NVIDIA vGPU host driver setup.
* **vGPU SR-IOV configure** — configure SR-IOV vGPU mode for hardware-partitioned GPU sharing.
* **Validate Passthrough** — verify that a previously configured passthrough setup is functioning correctly.
* **Validate vGPU** — verify that a previously configured vGPU setup is functioning correctly.

{% hint style="info" %}
For in-depth details on GPU Passthrough and vGPU setup, configuration options, and supported GPU models, see [GPU Support in <code class="expression">space.vars.product\_name</code>](/private-cloud-director/gpu/gpu-support-pcd).
{% endhint %}

#### PCD Connection Details (Step 4 of 4)

Enter the management plane connection details so the host can register with your <code class="expression">space.vars.product\_acronym</code> instance:

* **Account URL:** the URL of your <code class="expression">space.vars.product\_acronym</code> management plane.
* **Username:** your <code class="expression">space.vars.product\_acronym</code> account username.
* **Password:** your <code class="expression">space.vars.product\_acronym</code> account password.
* **Region** and **Tenant** details as applicable to your environment.

You can find these values in the <code class="expression">space.vars.product\_acronym</code> UI under **Infrastructure > Cluster hosts > Add New Hosts**.

{% hint style="info" %}
If you are setting up a traditional software-only self-hosted management plane rather than connecting to an existing one, see [Self-Hosted Install](/private-cloud-director/getting-started/self-hosted/self-hosted-install) for that path.
{% endhint %}

## Post-install Host Boot Console

Once <code class="expression">space.vars.product\_acronym</code> components are installed and the host has registered, the screen presents the <code class="expression">space.vars.product\_name</code> **— Host Boot Console**. This console displays the host specifications:

* **CPU** — processor model and core count (for example, "8 x 12th Gen Intel(R) Core(TM) i9-12900KS").
* **Memory** — total available memory (for example, "14.99Gi").
* **Network** — the primary interface name and its assigned IP address (for example, "ens3 — 172.16.122.88/24").

Two footer actions are available: press `F2` to open the **System Configuration** menu, or press `F12` to drop to a shell.

### System Configuration menu (F2)

Pressing `F2` opens the **System Configuration** menu with two options:

* **Configure Management Network** — update network settings after the initial installation.
* **View System Service Status** — open the service status view to inspect and manage <code class="expression">space.vars.product\_acronym</code> daemons.

### Viewing and restarting services

The **View System Service Status** screen lists each <code class="expression">space.vars.product\_acronym</code>-managed daemon with a status indicator. The services shown include **Network**, **Comms**, **Hostagent**, and **Sidekick**. Each service displays its current status (for example, **Active**).

If a service is not running, select it and use the **Restart** action to bring it back up without rebooting the host.


# Accept Certificate Authority

In the current release of <code class="expression">space.vars.product\_name</code> , a few services are configured with self-signed certificate. You must trust the certificate by accepting the self-signed certificate for proper operation of these services.

Following services require you to trust the certificate:

1. Image Library Service - You need to trust the certificate before you can upload any Images using <code class="expression">space.vars.product\_name</code> UI or API.
   1. When using the OpenStack CLI, you can use `--insecure` option to upload images to the Image Library Service without trusting the certificate.
2. Kubernetes Service - You need to trust the cluster endpoint certificate for each Kubernetes cluster that you create, before you can view any cluster specific data in the <code class="expression">space.vars.product\_name</code> UI.

To accept the certificate:

1. Click on the link in the UI where you see a notification that says "Action Required: Trust Certificate"
2. The link will open in a new tab. Expand on the "Advance" option in the tab and proceed to click on the link to trust the certificate.
3. Note that in some cases you may need to change the security settings for your browser before you can trust the certificate.


# Host Operating System Support Matrix

This page describes which host operating systems are supported for each Private Cloud Director release, along with deprecation timelines and end-of-support dates.

| PCD Release      | Ubuntu 22.04   | Ubuntu 24.04 | Platform9 OS powered by Rocky Linux from CIQ (RLC) 10.2 |
| ---------------- | -------------- | ------------ | ------------------------------------------------------- |
| **January 2026** | Supported      | Supported    | --                                                      |
| **April 2026**   | Supported      | Supported    | --                                                      |
| **August 2026**  | Deprecated     | Supported    | Supported                                               |
| **October 2026** | Deprecated     | Supported    | Supported                                               |
| **January 2027** | Deprecated     | Supported    | Supported                                               |
| **April 2027**   | End of support | Supported    | Supported                                               |

**Legend:**

* Supported = fully tested and maintained
* Deprecated = still functional, but no longer recommended for new deployments; plan to migrate
* End of support = no longer tested, patched, or supported

{% hint style="info" %}
**Ubuntu Support**

Ubuntu host operating systems follow a bring-your-own-OS model. You provide, install, and maintain Ubuntu yourself, and Platform9 supports Private Cloud Director on the Ubuntu versions marked Supported in the table above. Platform9 does not package, harden, or manage the Ubuntu operating system, and there is no Platform9-delivered Ubuntu image equivalent to Platform9 OS.
{% endhint %}

{% hint style="info" %}
**Rocky Linux Support**

Private Cloud Director is supported only on Platform9 OS powered by Rocky Linux from CIQ (RLC) 10.2 and installed through the Platform9 OS ISO. Platform9 OS is a controlled, hardened configuration: it is hardened to the CIS Level 1 benchmark and is configured to pull packages and updates only from the Platform9 validated package repository. Platform9 builds, tests, patches, and supports this operating system together with PCD as a single integrated stack.

Community Rocky Linux is not a supported PCD host operating system. This applies regardless of the version number: a host running community Rocky Linux 10.2 is still unsupported, because it does not carry the Platform9 OS hardening and does not use the Platform9 validated package repository, so it is not the configuration PCD is built and tested against. Private Cloud Director installed on a customer-provided Rocky Linux host is an unsupported configuration.
{% endhint %}

### Important dates

**August 2026:** Ubuntu 22.04 is deprecated. New cluster deployments should use Ubuntu 24.04 or Platform9 OS powered by Rocky Linux 10.2.

**January 2027:** Last PCD release that includes Ubuntu 22.04 support.

**April 2027:** Ubuntu 22.04 reaches end of support. Hosts running Ubuntu 22.04 must be migrated before upgrading to this release.

### Migration between host operating systems

Workload migration support depends on whether the source and destination hosts run the same host operating system.

| Migration type | Between Ubuntu and Platform9 OS hosts | Between Platform9 OS hosts |
| -------------- | ------------------------------------- | -------------------------- |
| Live migration | Not supported                         | Supported                  |
| Cold migration | Supported                             | Supported                  |

For details on migrating workloads from Ubuntu hosts to Platform9 OS hosts, see [Cross-OS Cold Migration: Ubuntu Hosts to Platform9 OS Hosts](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#cross-os-cold-migration-ubuntu-hosts-to-platform9-os-hosts).


# Configuration Maximums

This page lists the configuration limits Platform9 has validated for Private Cloud Director (PCD). Stay at or below these values to ensure full Platform9 support.

**These are tested limits, not the architectural maximums of the underlying components.** The individual components often support higher values. If your requirements exceed the published values, contact Platform9 support for guidance.

Actual achievable limits depend on hardware selection, workload characteristics, and network topology, and it may not be possible to hit every maximum simultaneously. Consult the individual subsystem sections (compute, network, storage) when sizing your environment. To scale beyond the per-Region maximums, deploy multiple Regions.

## Virtual Machine Maximums

Limits that apply to a single VM running on PCD.

### Storage <a href="#storage" id="storage"></a>

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Virtual disks per VM (root + attached volumes)</td><td>8</td></tr><tr><td>Boot volume size</td><td>4TB</td></tr><tr><td>Ephemeral disk size</td><td>No hard maximums exist</td></tr></tbody></table>

### Networking <a href="#networking" id="networking"></a>

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Security groups per port</td><td>10</td></tr><tr><td>Security group rules per security group</td><td>100</td></tr><tr><td>Floating IPs per VM</td><td>50</td></tr></tbody></table>

## Hypervisor Host Maximums

Limits that apply to a single PCD hypervisor host.

### Host Compute

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Maximum vCPU oversubscription ratio</td><td>16×</td></tr><tr><td>Maximum RAM oversubscription ratio</td><td>~1.5×</td></tr></tbody></table>

### Region Maximums <a href="#region-maximums" id="region-maximums"></a>

Limits that apply to a single PCD Region (the management plane controlling a set of hypervisors). Customers requiring scale beyond these values should plan for multiple Regions.

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Clusters per Region</td><td>8</td></tr><tr><td>Hypervisors per Region</td><td>591 active; 629 total registered</td></tr><tr><td>VMs per Region</td><td>9,673 powered on VMs</td></tr><tr><td>Tenants per Region</td><td>74</td></tr><tr><td>Users per Region</td><td>136</td></tr><tr><td>Host Aggregates per Region</td><td>9</td></tr><tr><td>Flavors per Region</td><td>15</td></tr><tr><td>Images per Region</td><td>7</td></tr><tr><td>Concurrent VM create operations per Region</td><td>50</td></tr><tr><td>Concurrent VM migrate operations per Region</td><td>5</td></tr></tbody></table>

## High Availability (VM HA) and DRR Maximums

Limits that apply to PCD VM HA clusters.

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Hypervisors per cluster</td><td>100</td></tr></tbody></table>

## Networking Maximums

Limits that apply to the software-defined networking layer.

### Networks and Subnets

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Tenant networks per Region</td><td>14</td></tr><tr><td>Subnets per Region</td><td>13</td></tr><tr><td>Subnets per network</td><td>100</td></tr><tr><td>Ports per Region</td><td>19,992</td></tr><tr><td>Ports per network</td><td>10,000</td></tr></tbody></table>

### Addressing and Isolation

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Security groups per Region</td><td>450</td></tr><tr><td>VLAN ID range per physical network</td><td>-</td></tr><tr><td>VXLAN / Geneve tunnel endpoints per Region</td><td>591</td></tr></tbody></table>

### Layer 2 Networks <a href="#layer-2-networks" id="layer-2-networks"></a>

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Ports per Layer 2 network</td><td>500</td></tr></tbody></table>

## Block Storage Maximums <a href="#block-storage-maximums" id="block-storage-maximums"></a>

Limits that apply to the block storage service.

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>Volumes per Region</td><td>10,000</td></tr><tr><td>Volumes per tenant</td><td>10,000</td></tr><tr><td>Backends per Region</td><td>6</td></tr><tr><td>Volume types per Region</td><td>6</td></tr></tbody></table>

## Tenant Maximums <a href="#tenant-maximums" id="tenant-maximums"></a>

Default quotas applied to each tenant.

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Criteria</th><th>Value</th></tr></thead><tbody><tr><td>VMs per tenant</td><td>10,000</td></tr><tr><td>vCPUs per tenant</td><td>10,000</td></tr><tr><td>RAM per tenant</td><td>8,000 GB</td></tr><tr><td>Volumes per tenant</td><td>10,000</td></tr><tr><td>Volume storage per tenant</td><td>10,000 GB</td></tr><tr><td>Networks per tenant</td><td>100</td></tr><tr><td>Routers per tenant</td><td>10</td></tr><tr><td>Floating IPs per tenant</td><td>50</td></tr><tr><td>Security groups per tenant</td><td>500</td></tr><tr><td>Key pairs per user</td><td>100</td></tr></tbody></table>


# Upgrade Overview

This section covers upgrading a <code class="expression">space.vars.product\_name</code> environment — preparing your hosts, sequencing the upgrade safely, and verifying that every service is healthy afterward.

<code class="expression">space.vars.product\_name</code> supports two deployment models:

* **SaaS** — Platform9 hosts and operates the management plane.
* **Self-Hosted** — you operate the management plane on-premise.

The runbooks below apply to both models. Steps that are specific to the on-premise (Self-Hosted) management plane — primarily the `airctl` commands — are called out separately within each runbook.

## In this section

* [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook) — pre-upgrade environment checks, host upgrade ordering across multi-role hosts, Ubuntu 22.04-to-24.04 OS upgrade caveats, and failure recovery.
* [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification) — the checklist to confirm region health, host roles, GPU, storage, and networking are healthy, then re-enable VM HA and DRR.

## Related pages

* **Self-Hosted management plane upgrade:** [Upgrade Management Plane and Hosts](/private-cloud-director/getting-started/self-hosted/upgrade) covers upgrading an on-premise (Self-Hosted) deployment with `airctl`.


# Host Upgrade Runbook

## Overview

This runbook guides you through upgrading the hosts in a <code class="expression">space.vars.product\_name</code> virtualized cluster. It covers the pre-upgrade environment checks you must complete before upgrading any host, the recommended ordering when hosts carry multiple roles, Ubuntu 22.04-to-24.04 host OS upgrade considerations, and how to recover from common upgrade failures.

{% hint style="info" %}
**Applies to both deployment models.** <code class="expression">space.vars.product\_name</code> supports two deployment models — **SaaS** (Platform9 hosts and operates the management plane) and **Self-Hosted** (you operate the management plane on-premise). The host-level guidance on this page applies to both models. Steps that act on the on-premise management plane — for example, the `airctl` commands and management-plane health checks — apply to **Self-Hosted deployments only** and are called out in boxes like the ones below. In SaaS deployments, Platform9 operates and upgrades the management plane.
{% endhint %}

In this guide, you will prepare your environment for a host upgrade, sequence the upgrade safely across multi-role hosts, handle Ubuntu OS upgrades without losing network connectivity, and recover if an upgrade stops partway through.

## Pre-Upgrade Checklist <a href="#pre-upgrade-checklist" id="pre-upgrade-checklist"></a>

Complete every item in this checklist before upgrading any host. Skipping checks is the most common cause of upgrade failures and extended maintenance windows.

### Check Region Health

The management plane must be healthy before any host upgrade begins. A degraded management plane cannot coordinate the host upgrade process reliably.

From the <code class="expression">space.vars.product\_name</code> UI, navigate to **Infrastructure > Regions** and confirm that all regions show a healthy status. Resolve any failing services before continuing.

{% hint style="info" %}
**Self-Hosted deployments only.** From the management cluster node, confirm region health from the command line:

```bash
airctl status
```

Confirm that every region shows `region health: ✅ Ready` and that `desired services` matches `ready services`. Resolve any failing pods or services before continuing.
{% endhint %}

### Verify Per-Host Free Disk Space

Insufficient disk space on a host causes the upgrade package installation to fail partway through, leaving the host in a partially upgraded state that is difficult to recover.

SSH to each host you plan to upgrade and check the following mount points:

```bash
df -h /var /var/log /opt /tmp
```

Recommended minimums before starting a host upgrade:

| Mount point | Recommended free space |
| ----------- | ---------------------- |
| `/var`      | At least 5 GB          |
| `/var/log`  | At least 2 GB          |
| `/opt`      | At least 3 GB          |
| `/tmp`      | At least 1 GB          |

If `/var/log` is nearly full, rotate or archive old logs before proceeding:

```bash
journalctl --vacuum-size=1G
find /var/log -name "*.gz" -mtime +30 -delete
```

### Confirm All Hosts Are Authorized and Healthy

Every host you intend to upgrade must show an `applied` role status and an `online` connection status. Hosts in `failed`, `error`, `converging`, or `unknown` states should be resolved before the upgrade.

From the UI, navigate to **Infrastructure > Cluster Hosts** and confirm that no host shows a warning or error badge.

{% hint style="info" %}
**Self-Hosted deployments only.** You can also check host status from the management cluster node:

```bash
airctl host-status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

Look for any host where `Status` is not `ok` or `Agent Status` is not `running`. Investigate and resolve those hosts first.
{% endhint %}

### Verify Storage Backend Connectivity

If any host carries the Persistent Storage Service role (block storage) or the Image Library Service role, verify that those services are reachable and healthy before upgrading those hosts.

**Persistent Storage Service (block storage):**

```bash
pcdctl volume service list
```

All volume service endpoints should show `enabled` and `up`. If any endpoint is `down` or `disabled`, investigate before upgrading the host that carries that role.

**Image Library Service:**

```bash
pcdctl image-service list
```

Confirm that the Image Library Service endpoints are `enabled` and `up`.

You can also verify storage and image library health from the UI: navigate to **Infrastructure > Storage** and **Infrastructure > Image Library** respectively.

### Verify Network Connectivity

Confirm that all hosts are reachable over the management network. From any host (or, for Self-Hosted deployments, the management cluster node) that can reach the management network:

```bash
for ip in <host-ip-1> <host-ip-2> ...; do
  ping -c 2 "$ip" && echo "$ip OK" || echo "$ip UNREACHABLE"
done
```

Replace `<host-ip-1>`, `<host-ip-2>`, and so on with the management IP addresses of your hosts. If any host is unreachable, resolve the connectivity issue before upgrading.

### Disable VM HA and DRR Before Host OS Upgrades

{% hint style="warning" %}
**Important: disable VM HA and DRR before upgrading host operating systems.**

VM evacuation and live migration require matching operating system and KVM versions between source and destination hosts. If VM HA or DRR is active while some hosts run Ubuntu 22.04 and others run Ubuntu 24.04, evacuation attempts between those hosts will fail.
{% endhint %}

If you are upgrading the host OS (for example, from Ubuntu 22.04 to 24.04) — as distinct from upgrading the <code class="expression">space.vars.product\_name</code> host agent packages only — disable VM HA and DRR at the cluster level before starting:

1. Navigate to **Infrastructure > Clusters** in the UI.
2. Select the cluster whose hosts you are upgrading.
3. Toggle **VM High Availability** to off.
4. Toggle **Dynamic Resource Rebalancing** to off.

Re-enable both only after all hosts in the cluster have been upgraded to the same OS version. See [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification#re-enable-vm-ha-and-drr) for the re-enablement steps.

For full pre-conditions and behavior details, see [Virtual Machine High Availability](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) and [Dynamic Resource Rebalancing (DRR)](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr).

### Plan Your Maintenance Window

Estimate the upgrade window before you begin:

* Host agent-only upgrades (no OS change) typically complete in 5–10 minutes per host.
* Host OS upgrades (Ubuntu 22.04 to 24.04) require an additional 20–40 minutes per host plus reboot time.

Schedule the window to avoid peak workload hours. Communicate the window to application owners, because workloads on hosts entering maintenance mode will be live-migrated to other cluster hosts during the upgrade.

If any VMs remain on a host while its OVN controller is upgraded — for example, when maintenance mode is skipped or VM HA is disabled — those VMs briefly lose network packets while the controller restarts and re-establishes flows. Platform9 testing has observed roughly **4–6 seconds** of packet loss per host (4 seconds with 47 VMs on Ubuntu 24.04, 6 seconds with 15 VMs on Ubuntu 22.04). Existing TCP connections recover automatically. Following the maintenance-mode procedure below avoids this disruption by draining the host before its OVN controller restarts.

## Host Upgrade Ordering and Role Dependencies <a href="#host-upgrade-ordering" id="host-upgrade-ordering"></a>

When hosts carry multiple roles, the order in which you upgrade them affects service availability. Follow these sequencing rules. They apply to both deployment models.

### Recommended Upgrade Order

1. **Hosts with only the Hypervisor role** — upgrade these first. They carry no shared service responsibility, so their upgrade has the narrowest blast radius.
2. **Hosts with the Networking Service role** — upgrade these next. Put each host into maintenance mode before upgrading to migrate its VMs first.
3. **Hosts with the Image Library Service role** — upgrade image library hosts one at a time. Confirm the Image Library Service is healthy on the remaining hosts before upgrading the next one.
4. **Hosts with the Persistent Storage Service role** — upgrade block storage hosts one at a time. Confirm that volumes are accessible from other hosts before proceeding to the next block storage host.
5. **Multi-role hosts** — hosts that carry hypervisor, image library, or block storage roles simultaneously are the highest-risk to upgrade. Treat them like the most sensitive single role they carry (block storage > image library > hypervisor-only).

### Put Each Host Into Maintenance Mode Before Upgrading

Before upgrading any host, use maintenance mode to drain its VMs:

1. Navigate to **Infrastructure > Cluster Hosts**.
2. Select the host.
3. Click **Other > Enable Maintenance Mode** and follow the prompts.
4. Wait for the **Migration Status** banner to show all VMs migrated.

Then proceed with the upgrade for that host. After the upgrade is complete and the host is back online, disable maintenance mode before moving to the next host.

See [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) for full details on the enable and disable process.

### Persistent Storage Service Host Considerations

If your region has only one host with the Persistent Storage Service role, upgrading that host briefly makes block storage unavailable. Schedule that upgrade during a maintenance window when no block storage volume attach or detach operations are expected. If possible, add a second block storage host before the upgrade window so there is no single point of failure.

### Image Library Service Host Considerations

If your region has only one host with the Image Library Service role, new VM provisioning from images will fail during that host's upgrade. All existing running VMs continue operating normally. Schedule that upgrade during a window when no new VMs need to be deployed.

## Ubuntu 22.04 to 24.04 Host OS Upgrade Caveats <a href="#ubuntu-os-upgrade-caveats" id="ubuntu-os-upgrade-caveats"></a>

{% hint style="warning" %}
**Network connectivity risk on hypervisor hosts.**

Upgrading from Ubuntu 22.04 to 24.04 on a host with the Hypervisor role carries a risk of OVS (Open vSwitch) bridge misconfiguration during the OS upgrade. If the OVS bridge loses its interface binding after the reboot, the host loses network connectivity and cannot re-connect to the management plane. Plan for out-of-band console access (IPMI, iDRAC, or equivalent) before starting the OS upgrade on hypervisor hosts.
{% endhint %}

### Before the OS Upgrade

1. Record the current OVS bridge configuration on the host:

```bash
ovs-vsctl show
ip addr show
```

Save the output. You will need it to verify the configuration is intact after the reboot.

2. Confirm the host is in maintenance mode and all its VMs have been migrated off.
3. If the host has the Hypervisor role, ensure you have out-of-band console access in case network connectivity is lost after the OS upgrade.

### After the OS Upgrade and Reboot

Perform network re-validation immediately after the host comes back online:

1. Confirm the OVS bridge is present and the physical interface is still a member:

```bash
ovs-vsctl show
```

The output should show your bridge (for example, `br-ex` or `br-data`) with the physical interface listed as a port. If the physical interface is missing from the bridge, re-add it:

```bash
ovs-vsctl add-port <bridge-name> <physical-interface>
```

Replace `<bridge-name>` and `<physical-interface>` with the values from your pre-upgrade recording.

2. Confirm the host can reach the management plane:

```bash
curl -sk https://<management-plane-fqdn>/resmgr/v1/hosts | head -c 200
```

3. Confirm the host agent is running and connected:

```bash
systemctl status pf9-hostagent
```

The service should be `active (running)`. If it is stopped or failed, restart it:

```bash
systemctl restart pf9-hostagent
```

4. From the management plane, confirm the host shows `online` connection status in **Infrastructure > Cluster Hosts**.
5. After confirming network and host agent health, disable maintenance mode and allow the host to rejoin scheduling before proceeding to the next host.

### Mixed OS Versions in a Cluster

Do not leave a cluster in a mixed OS state longer than necessary. A cluster where some hosts run Ubuntu 22.04 and others run Ubuntu 24.04 has the following limitations:

* VM HA evacuations between mixed-OS hosts will fail (KVM version mismatch).
* DRR live migrations between mixed-OS hosts will fail.
* Manual VM migrations between mixed-OS hosts are not supported.

Keep VM HA and DRR disabled for the entire cluster until all hosts are on the same OS version.

## Host Upgrade Failure Recovery <a href="#failure-recovery" id="failure-recovery"></a>

The host agent and package-level recovery steps below apply to both deployment models. Steps that re-submit an upgrade through `airctl` or inspect the region's Kubernetes namespace apply to Self-Hosted deployments only and are marked. In SaaS deployments, if a host upgrade fails and the host-level steps below do not recover it, contact Platform9 Support.

### 409 Conflict Error During Host Upgrade

A `409 Conflict` response during a host upgrade typically means the management plane has a stale lock or an in-progress record for that host from a previous attempt. This prevents the upgrade from being re-submitted.

**Recovery steps:**

1. Verify the host agent is running and reports back to the management plane:

```bash
# On the affected host
systemctl status pf9-hostagent
```

{% hint style="info" %}
**Self-Hosted deployments only.** Check whether the host upgrade pod from the previous attempt is still running, and clear it before retrying:

```bash
kubectl get pods -n <region-fqdn> | grep host-upgrade
```

If a `host-upgrade-*` pod is in `Running` or `Pending` state from a prior attempt, wait for it to complete or delete it:

```bash
kubectl delete pod <host-upgrade-pod-name> -n <region-fqdn>
```

{% endhint %}

Once the host agent is confirmed running (and, for Self-Hosted deployments, the stale pod is cleared), retry the host upgrade. If the conflict persists, contact Platform9 Support.

### "Cluster Name Removed from Host" Error

This error appears when the host's local configuration no longer contains the cluster association record. It typically occurs if the host was deauthorized or had its roles removed while an upgrade was in progress.

**Recovery steps:**

1. Check the host's role status. From the UI, navigate to **Infrastructure > Cluster Hosts** and inspect the affected host.
2. If the host shows `unauthorized` or missing roles, re-authorize it and re-assign roles from the UI: select the host, click **Edit Roles**, and re-assign the appropriate roles.
3. Wait for the host to reach `applied` role status (the status transitions through `converging`).
4. Once the host is back to `applied` status, retry the host upgrade.

### Partial or Failed Upgrade Cleanup

If a host upgrade fails midway, the host may be in an inconsistent state with mixed package versions. Before retrying, perform the following cleanup on the affected host:

1. Check which <code class="expression">space.vars.product\_name</code> packages are installed on the host and their versions:

```bash
dpkg -l | grep pf9
```

2. If you see a mix of old and new package versions, force-reinstall the current target packages to bring the host to a consistent state:

```bash
apt-get install --reinstall pf9-ostackhost pf9-neutron-base pf9-neutron-ovn-controller
```

Add any other `pf9-*` packages shown by the `dpkg -l` output that are at the wrong version.

3. Restart the host agent after reinstalling:

```bash
systemctl restart pf9-hostagent
```

4. Wait for the host to return to `applied` status in the management plane, then retry the upgrade.

### Host Re-Sync and Retry

If a host is stuck in `converging` state for more than 10 minutes after a failed upgrade:

1. From the UI, navigate to **Infrastructure > Cluster Hosts**, select the host, and use **Other > Re-sync Host** (if available) to trigger a re-convergence.
2. Alternatively, restart the host agent on the host directly:

```bash
systemctl restart pf9-hostagent
```

3. Monitor the hostagent log for convergence progress:

```bash
tail -f /var/log/pf9/hostagent.log
```

Look for `converge successful` or `role application complete` messages.

4. If the host does not converge within another 10 minutes, contact Platform9 Support with the hostagent log from the affected host (and, for Self-Hosted deployments, the output of `airctl host-status`).

## Next Steps

After all hosts are upgraded, complete the [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification) checklist to confirm that all services are healthy and re-enable VM HA and DRR.


# Post-Upgrade Verification

## Overview

After upgrading all hosts, run through this verification checklist to confirm that every service is healthy and every role is operating correctly. Do not re-enable VM HA or DRR until you have completed the service and role verification steps below.

{% hint style="info" %}
**Applies to both deployment models.** The host, role, storage, networking, and GPU verification steps on this page apply to both **SaaS** and **Self-Hosted** deployments. Management-plane verification using `airctl` applies to **Self-Hosted deployments only** and is marked. In SaaS deployments, Platform9 operates and upgrades the management plane; verify region health from the UI.
{% endhint %}

In this guide, you will confirm region health, verify all host roles, validate GPU passthrough configuration, test storage volume attach, test Networking Service connectivity, and re-enable VM HA and DRR.

## Verify Region Health <a href="#verify-management-plane" id="verify-management-plane"></a>

From the <code class="expression">space.vars.product\_name</code> UI, navigate to **Infrastructure > Regions** and confirm all regions show a healthy status before proceeding.

{% hint style="info" %}
**Self-Hosted deployments only.** Run `airctl status` from the management cluster node and confirm every region is ready:

```bash
airctl status
```

Expected output for each region:

```
deployment status:   ready
region health:       ✅ Ready
desired services:    <N>
ready services:      <N>
```

The `desired services` and `ready services` counts must match. If any services are not ready, check the pod logs in the region namespace before proceeding:

```bash
kubectl get pods -n <region-fqdn> | grep -v Running | grep -v Completed
```

Investigate any pod in `CrashLoopBackOff`, `Error`, or `Pending` state.
{% endhint %}

## Verify Host Role Status <a href="#verify-host-roles" id="verify-host-roles"></a>

Every host in every region must show `Status: ok` and `Agent Status: running` after the upgrade.

From the UI, navigate to **Infrastructure > Cluster Hosts** and confirm:

* No host shows a warning or error badge.
* All hosts that had the Hypervisor role before the upgrade still show the Hypervisor role as `applied`.
* Hosts with the Persistent Storage Service role show that role as `applied`.
* Hosts with the Image Library Service role show that role as `applied`.

If a host's role is missing or shows as `unauthorized` after the upgrade, re-assign the role: select the host, click **Edit Roles**, assign the appropriate roles, and click **Update Role Assignment**. If a host is stuck in `converging`, select the host and choose **Other > Re-sync Host**.

For any host that is not healthy, check the host agent log on that host:

```bash
tail -n 100 /var/log/pf9/hostagent.log
```

{% hint style="info" %}
**Self-Hosted deployments only.** You can also verify host role status from the management cluster node:

```bash
airctl host-status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

{% endhint %}

## Validate GPU Passthrough <a href="#validate-gpu-passthrough" id="validate-gpu-passthrough"></a>

{% hint style="info" %}
Skip this section if your deployment does not use GPU passthrough.
{% endhint %}

After upgrading hosts that carry the Hypervisor role with GPU passthrough configured, re-validate that the GPU devices are still bound correctly.

1. On each GPU host, confirm the GPU device is bound to the `vfio-pci` driver:

```bash
lspci -k | grep -A 3 -i nvidia
```

The `Kernel driver in use` field should show `vfio-pci`. If it shows a different driver, the GPU binding was not preserved through the OS upgrade. Follow the [Set up GPU Passthrough](/private-cloud-director/gpu/gpu-support-pcd/set-up-gpu-passthrough) guide to rebind the device.

2. From the <code class="expression">space.vars.product\_name</code> UI, navigate to **Infrastructure > Cluster Hosts**, select a GPU host, and confirm that the GPU devices are listed under the host's hardware details.
3. Attempt to launch a test VM using a GPU-enabled flavor. Confirm the VM starts successfully and the GPU is accessible from within the VM.

For vGPU deployments, confirm that the vGPU profiles are still present and the vGPU driver service is running on the host:

```bash
systemctl status nvidia-vgpu-mgr
```

For full GPU troubleshooting steps, see [Troubleshooting GPU Support](/private-cloud-director/gpu/gpu-support-pcd/troubleshooting-gpu-support).

## Test Persistent Storage Volume Attach <a href="#test-volume-attach" id="test-volume-attach"></a>

Confirm that the Persistent Storage Service is operational by attaching a volume to a running VM.

1. List available volumes:

```bash
pcdctl volume list
```

2. Identify a test VM that is running:

```bash
pcdctl server list --status ACTIVE
```

3. Attach an existing available volume to the test VM:

```bash
pcdctl server add volume <server-id> <volume-id>
```

4. Confirm the attachment succeeded:

```bash
pcdctl server volume list <server-id>
```

The volume should appear with a `in-use` status. Detach the test volume after confirming:

```bash
pcdctl server remove volume <server-id> <volume-id>
```

If the attach fails, check the Persistent Storage Service endpoint status:

```bash
pcdctl volume service list
```

All endpoints should be `enabled` and `up`. If an endpoint is `down`, check the volume service logs on the block storage host.

## Test Networking Service Connectivity <a href="#test-networking" id="test-networking"></a>

Verify that the Networking Service is healthy and that VM network connectivity is intact after the upgrade.

### Verify Networking Service Endpoints

```bash
pcdctl network agent list
```

All agents should show `alive: True`. If any agent is not alive, check the networking agent status on the affected host:

```bash
systemctl status pf9-neutron-ovn-controller
```

### Test Network and Router Connectivity

1. List the networks in the region and confirm the expected networks are present:

```bash
pcdctl network list
```

2. List routers and confirm they are in `ACTIVE` state:

```bash
pcdctl router list
```

3. From a test VM, confirm outbound connectivity (if the VM has network egress):

```bash
# From within the test VM
ping -c 4 8.8.8.8
```

If any router is not `ACTIVE`, navigate to **Infrastructure > Networking > Routers** in the UI and inspect the router details for error messages.

### Test VM-to-VM Connectivity

Confirm that VMs on different hosts can communicate across the virtual network:

1. Identify two running VMs on different hypervisor hosts.
2. From one VM, ping the private IP address of the other:

```bash
ping -c 4 <other-vm-private-ip>
```

If ping fails between VMs on different hosts, check the OVS bridge status on the hypervisor hosts:

```bash
ovs-vsctl show
ovs-ofctl dump-flows br-int | head -20
```

## Re-Enable VM HA and DRR <a href="#re-enable-vm-ha-and-drr" id="re-enable-vm-ha-and-drr"></a>

Re-enable VM HA and DRR only after:

* All hosts in the cluster are running the same operating system version.
* All hosts show `Status: ok` and all roles are `applied`.
* The region health check passed.
* Storage and networking tests passed.

{% hint style="warning" %}
Do not re-enable VM HA if any hosts in the cluster are still running a different OS version. VM evacuation between hosts with different KVM versions will fail.
{% endhint %}

**Re-enable VM HA:**

1. Navigate to **Infrastructure > Clusters**.
2. Select the cluster.
3. Toggle **VM High Availability** to on.
4. Confirm the cluster shows `Protected` VM HA status after enabling. If it shows `Degraded` or `Not Protected`, hover over the VM HA status to see which prerequisite is not met, and resolve it before relying on VM HA for workload protection.

**Re-enable DRR:**

1. Navigate to **Infrastructure > Clusters**.
2. Select the cluster.
3. Toggle **Dynamic Resource Rebalancing** to on and confirm the frequency setting is correct for your environment.

For pre-conditions and full behavior details, see [Virtual Machine High Availability](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) and [Dynamic Resource Rebalancing (DRR)](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr).

## Upgrade Complete

Your <code class="expression">space.vars.product\_name</code> upgrade is complete once all verification steps above have passed and VM HA and DRR are re-enabled with healthy statuses. If you encounter issues during verification that you cannot resolve, contact Platform9 Support with the relevant pod or service logs. For Self-Hosted deployments, also include the output of `airctl status` and `airctl host-status`.


# Cluster Blueprint

A **Virtualized Cluster** in <code class="expression">space.vars.product\_name</code> is a group of hypervisor hosts that virtual machines get provisioned on.

You can create one or more virtualized clusters per region. You can also further slice a cluster into multiple subgroups of hosts using [Host Aggregates](/private-cloud-director/virtualized-clusters/host-aggregate). This allows you to combine hosts with similar characteristics into a cluster or a host aggregate within a cluster and target VM provisioning to that cluster or aggregate.

## What is Cluster Blueprint

A cluster blueprint allow you to describe common configure that all virtualized clusters will share in a declarative, prescriptive manner. Blueprint is designed to help you express your desired cluster architecture upfront, and ensure that cluster capacity that is added over time conforms to this desired architecture.

## Create a Cluster Blueprint

To deploy and use a **virtualized cluster**, your first step will be to create a cluster blueprint.

Navigate to Infrastructure -> Cluster Blueprints in the <code class="expression">space.vars.product\_name</code> UI to create your cluster blueprint.

Follow the details in [Networking Configuration](/private-cloud-director/virtualized-networking/networking-overview#networking-configuration) to find out more about configuring networking service as part of the cluster blueprint configuration.

Follow the details in [Block Storage Service Configuration](/private-cloud-director/storage/block-storage#block-storage-service-configuration) to learn about creating one or more storage types as part of cluster blueprint.

### Customize Cluster Defaults

Here you get to customize storage location for image library and virtual machine ephemeral storage.

### Image Library Storage Location

Specify the storage type and location where virtual machine images will be stored.

You can specify a **file system path or a volume type name** for this option.

If you specify a **volume type name** for the image library location, that volume type will be used to store images in the image library. If you specify a filesystem path instead, you must indicate whether the path corresponds to **local** or **shared storage**. Enable the **Shared Storage** toggle if the path is configured to use NFS or an equivalent setup that is shared and mounted on all image library hosts.

### Virtual Machine Storage Path

The virtual machine storage path specifies the disk location on each hypervisor that will be used to store:

* Ephemeral root disk files for VMs running on this hypervisor that are using ephemeral disks.
* Any additional VM metadata and files.

Read [Ephemeral Storage](/private-cloud-director/storage/ephemeral-storage#what-is-ephemeral-storage) for an in-depth understanding of what is ephemeral storage and it's advantages and disadvantages before you configure this for your setup.

The default value for this path is `/opt/data/instances`. You can modify the default to a different location based on your preference.

This can be a local disk or an NFS shared disk, depending on how the storage path is configured at each hypervisor level.

Note that you configure this path once at the blueprint level for all hosts across all clusters in your hypervisor.

### Virtual Machine Console IP

Specify an IP address or a domain name to be used to access the console of all VMs in the region. You can configure this in the blueprint. Ensure that proper routing or DNS resolution is in place so the console can be accessed reliably. Read more [details here](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint/setting-up-a-console-proxy) on how to set this up.

## Update Cluster Blueprint

Some configurations in the cluster blueprint may be updated after hosts are added to the clusters, others are subject to limitations depending on how they are being applied to existing hosts. Cluster Network Parameters and Image Library and VM storage locations cannot be edited once hosts are authorized to the <code class="expression">space.vars.product\_acronym</code> region.

You can still make changes to cluster blueprint to add a network interface to an existing in-use Host Configuration or update the key-value metadata for Volume Backend Configurations. These configuration changes do get propagated to existing hosts that are already part of a cluster. When adding a new network interface to an in-use Host Configuration, you would not be able to change the system traffic options to use this new interface, but you could create a new Physical Network using the Physical Network Label associated with this interface.

{% hint style="info" %}
**Editing Host Config**

Within the **Edit Roles** modal, the **Host Config** (network configuration) selection can only be modified when no roles are currently installed on the host. If roles (such as Hypervisor, Image Library, or Persistent Storage) are already active, the network configuration dropdown will be disabled. To change the network configuration for an existing host, the host must be fully reset. This process requires the removal of all assigned roles from the host within the blueprint, followed by a **Deauthorize** and **Decommission** action on the host. Once the host is cleared and successfully re-onboarded to the platform, the **Host Config** can be modified during the role assignment process.
{% endhint %}


# Setting Up a Console Proxy

<code class="expression">space.vars.product\_name</code> allows you to proxy all of the VM console connections in a region through one or more servers. You can configure this in the cluster blueprint using the field **VNC Proxy IP or Domain Name,** and set it to a domain name or an IP address. The proxy must always point to a hypervisor node in the same region. Using any other server or VM is not supported.

#### How console proxy traffic works <a href="#how-console-proxy-traffic-works" id="how-console-proxy-traffic-works"></a>

When a user opens a VM console, the connection does not go directly to the hypervisor running the VM. Every console session in the region first reaches the hypervisor that the console proxy endpoint resolves to (the **authoritative node),** and that host then proxies the session to the hypervisor that actually runs the VM.

In the example below, the endpoint resolves to **Hyp1**, which acts as the VNC proxy and forwards the session to **Hyp2**, where the VM runs. **Hyp3** is not involved in this session.

```
    ┌────────────────┐
    │     User /     │
    │    Browser     │
    └────────┬───────┘
             │  VNC console request
             ▼
   ┌──────────────────┐      ┌──────────────────┐      ┌──────────────────┐
   │       Hyp1       │      │       Hyp2       │      │       Hyp3       │
   │ (VNC proxy node) │      │                  │      │                  │
   │                  │      │   ┌──────────┐   │      │                  │
   │                  │─────▶│   │    VM    │   │      │                  │
   │                  │      │   └──────────┘   │      │                  │
   └──────────────────┘      └──────────────────┘      └──────────────────┘

Hyp1 is the resolved host and acts as the VNC proxy. It forwards the
session to Hyp2, which runs the requested VM. Hyp3 and any other
hosts can stay isolated from direct user access.
```

1. **Request.** The user opens a VM console from the UI or API. <code class="expression">space.vars.product\_name</code> returns a console URL that points to the console proxy endpoint — the IP or domain name set in the blueprint — on port `6080`, together with a one-time token that identifies the target VM and the hypervisor running it.
2. **Resolve the endpoint to a hypervisor.** The console proxy endpoint always resolves to a hypervisor in the region (never a standalone server or VM). Which host becomes the authoritative node depends on how you configured the proxy:
   * **A hypervisor's IP** — every session is routed through that one fixed host.
   * **A floating IP / VIP** — the session reaches whichever host currently holds the VIP (for example, managed by keepalived).
   * **A domain name** — the name can map to one or more hypervisors. See [Using a domain name](#using-a-domain-name) below.
3. **Forward to the correct hypervisor.** The noVNC proxy service on the authoritative node uses the token to identify the hypervisor running the target VM, opens a backend connection to it over the internal host-to-host network, and relays the console stream back to the user.

Only the authoritative node needs to be reachable by users. Every other hypervisor can have external access blocked, as long as hypervisor-to-hypervisor traffic on the console port is allowed.

**Using a domain name**

A domain name can map to more than one hypervisor. In that case your DNS, VPN, or `/etc/hosts` decides which host the request lands on, and that host then proxies the session to the hypervisor running the VM. In the example below the name resolves to either **Hyp1** or **Hyp3**, and from there the session reaches **Hyp2**:

```
                         ┌──────────────────┐
                         │      User /      │
                         │     Browser      │
                         └─────────┬────────┘
                                   │  console.example.com
                                   │  (DNS resolves to Hyp1 or Hyp3)
             ┌─────────────────────┴─────────────────────┐
             ▼                                           ▼
   ┌──────────────────┐                        ┌──────────────────┐
   │ Hyp1 (VNC proxy) │                        │ Hyp3 (VNC proxy) │
   └─────────┬────────┘                        └─────────┬────────┘
             │                                           │
             └─────────────────────┬─────────────────────┘
                                   ▼
                         ┌──────────────────┐
                         │       Hyp2       │
                         │                  │
                         │   ┌──────────┐   │
                         │   │    VM    │   │
                         │   └──────────┘   │
                         └──────────────────┘

The domain name maps to more than one hypervisor (here Hyp1 and Hyp3).
Your DNS, VPN, or /etc/hosts decides which one the request lands on;
that host then proxies the session to Hyp2, which runs the VM.
```

#### Options for configuring the console proxy <a href="#options-for-configuring-the-console-proxy" id="options-for-configuring-the-console-proxy"></a>

**1. Use a hypervisor’s IP address**

If you have multiple hypervisors, you can route all console connections through a single host by setting that hypervisor’s IP as the proxy IP in the blueprint. This allows you to restrict public access to all other hypervisors.

**2. Use a public /** **floating IP**

If your hypervisor nodes can route a public or floating IP, you can set that floating IP in the blueprint. Traffic will then flow through whichever host currently holds the floating IP. Tools like keepalived can be used to manage this.

**3. Use a domain name**

You can also provide a domain name instead of an IP. To make this work, map the domain name to one or more hypervisor hosts either:

* locally in your `/etc/hosts`, or
* through your organization’s DNS or VPN configuration.

#### Limitations <a href="#limitations" id="limitations"></a>

The console proxy is not a standalone, externally hosted gateway. It must be one (or more) of the hypervisors in the same region. Using any other server or VM is not supported. As a result:

* At least one hypervisor must remain reachable by users to act as the proxy (the jump point). A configuration where *all* hypervisors are isolated from direct user access, with only an external proxy exposed, is not supported today.
* The remaining hypervisors can be fully isolated from direct user access. They only need to accept console (host-to-host) traffic from the authoritative node.


# Virtualized Cluster

<code class="expression">space.vars.product\_name</code> enables you to manage and operate multiple virtualized clusters from a single <code class="expression">space.vars.product\_acronym</code> region.

## Use Cases

1. **Multi-Tenant Isolation** - You can assign different tenants to separate clusters to enforce resource boundaries and fault domain separation. The isolation prevents noisy neighbor issues and enhances security between tenant environments.
2. **Licensing Requirements -** Some specialized software, such as Oracle database, may require isolating hosts with the software license enabled. Creating a separate cluster will allow you to do this.
3. **Hardware Specialization -** You can group hosts with similar capabilities (such as GPU-enabled or high-memory machines) into dedicated clusters to optimize performance and resource utilization for specialized workloads.

## Understanding a Virtualized Cluster

A virtualized cluster is a grouping of interconnected physical servers or hosts, with certain cluster level features operating at the group level. Resources such as CPU, Memory, and GPU across all hosts in a virtualized cluster are presented as a single pool, so that when you create a new virtual machine in the virtualized cluster, the placement of the new virtual machine may be on any of the underlying physical servers, based on capacity and other provisioning constraints.

All clusters within a region operate under a **single cluster blueprint**, ensuring consistency in base configurations while allowing for cluster-specific customization.

Each virtualized cluster provides two key features that enable you to run production workloads.

### Create a Virtualized Cluster

To create a new virtualized cluster:

1. Navigate to **Infrastructure** > **Clusters** > **Add Cluster**
2. Provide a name for the cluster.
3. Choose desired settings for:
   1. VMHA (Virtual Machine High Availability)
   2. DRR (Dynamic Resource Rebalancing)
   3. GPU (Passthrough or vGPU)

### CPU Mode and Model

Defines how the hypervisor exposes CPU features to virtual machines. You can select one of the following CPU modes during cluster creation in the <code class="expression">space.vars.product\_name</code> UI.

* **Default:** Automatically selects the latest supported CPU model on the host.
* **Host Model:** Matches the physical CPU closely, balancing performance and migration flexibility. This mode is recommended for clusters with mostly homogeneous hardware (e.g., all Intel Skylake or all AMD EPYC of similar generation).
* **Host Passthrough:** Exposes the physical CPU directly for maximum performance (migration only possible between identical hosts). Host passthrough is not generally recommended due to migration limitations. This mode is suitable for:
  * Clusters with effectively identical host CPUs.
  * Clusters where the best possible performance is desired without any live migration needs.
* **Custom**: Specify a standardized CPU model to ensure consistent virtualized CPUs across hosts and enable live migration between different hardware. This option lets you pick a CPU model that matches the oldest or least capable CPUs in your fleet, allowing live migration between hosts even if they span several CPU generations. The trade-off is that you cannot fully leverage the features and performance of your newer CPU in favor of live migration compatibility.

### Virtual Machine High Availability (VM HA)

VM HA provides automatic fault tolerance for your workloads. When enabled:

* The system continuously monitors host health across the cluster.
* If a host failure is detected, affected VMs are automatically recovered on healthy hosts.
* Minimize downtime without requiring manual intervention.
* Business continuity is maintained even during infrastructure issues.

Read more here about [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha).

### Dynamic Resource Rebalancing (DRR)

DRR works as a continuous optimization engine that:

* Monitors allocated capacity and real-time utilization metrics (CPU and memory) across all hosts in the cluster.
* Analyzes resource distribution patterns to identify imbalances
* Intelligently migrates VMs across hosts in the cluster to optimize resource utilization.
* Prevents hotspots and resource contention before they impact performance

This proactive approach ensures that your clusters maintain optimal performance even as workload patterns change over time.

Read more here about [Dynamic Resource Rebalancing (DRR)](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr).

### GPU

Enabling GPU support allows you to group all hosts capable of GPU passthrough or vGPU-backed hypervisors into a single GPU-aware virtual cluster.

Before configuring GPU mode on any GPU hosts, ensure the following points are met.

* Supported GPU Modes: Passthrough or vGPU
* Use Passthrough to assign full physical GPUs to virtual machines.
* Use virtual GPUs (vGPUs) to virtually slice existing physical GPUs and share them across virtual machines.
* Non-GPU hosts can not be added to a GPU-enabled cluster.

This setting ensures that all GPUs in the cluster use the same GPU mode (Passthrough or vGPU) for uninterrupted GPU operations, such as live migration between GPU hardware, resizing, and scaling GPUs.

Read more here about [GPU Support in Private Cloud Director](/private-cloud-director/gpu/gpu-support-pcd).


# Dynamic Resource Rebalancing (DRR)

In this document you will learn about <code class="expression">space.vars.product\_name</code> Dynamic Resource Rebalancing (DRR), an automated feature that optimizes resource utilization of your virtualized clusters. DRR ensures efficient resource utilization via live migration of virtual machines, preventing performance bottlenecks.

## Introduction

In any dynamic virtualized environment, workloads fluctuate. Some virtual machines might become resource-intensive, while others remain idle. This can lead to imbalances across the physical hypervisor hosts within a cluster – some hosts become overloaded, causing performance degradation for their VMs, while other hosts sit underutilized. <code class="expression">space.vars.product\_name</code> addresses this challenge with its integrated **Dynamic Resource Rebalancing (DRR)** feature. DRR works continuously to ensure that resources like CPU and memory are utilized efficiently and that no single host becomes a performance bottleneck.

## Benefits of DRR

Enabling DRR in your <code class="expression">space.vars.product\_name</code> environment delivers tangible benefits:

* **Consistent VM Performance:** Helps prevent performance issues caused by resource contention on overloaded hosts.
* **Optimized Resource Utilization:** Ensures that your hardware investment is used more efficiently across the cluster.
* **Proactive Problem Avoidance:** Identifies and resolves potential resource bottlenecks before they negatively impact applications.
* **Reduced Operational Overhead:** Automates the complex task of monitoring and balancing VM workloads, freeing up administrator time.

## DRR Pre-requisites

* DRR requires at **least two hosts** in a virtualized cluster to operate.
* DRR has the same pre-requisites as [VM Live Migration Prerequisites](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration-prerequisites).
* We recommend that you configure the cluster hosts to have the same [CPU and Memory over-commitment ratio](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#host-overcommitment-allocation-ratios) best performance.
* Make sure to read the current [#known-issues-and-limitations](#known-issues-and-limitations "mention") section before proceeding to enable DRR on a cluster.

## Configure DRR for a Cluster

* **Scope** - DRR is configured as a property at **each virtualized cluster level**. When you create a new virtualized cluster, you have the option to enable DRR for the cluster.
  * You can also enable DRR on an existing virtualized cluster at a later point.
* **Frequency** - When you enable DRR, you are asked to choose a **value for frequency** at which DRR should run.
  * The default is **20 minutes**. Which means every 20 minutes, DRR will evaluate hosts within the virtualized cluster and identify targets for potential rebalancing.
  * You can change this value to **10 minutes or 30 minutes**.

## DRR and VM Migration Priority

### What Is Migration Priority

Migration priority is a key concept that is part of <code class="expression">space.vars.product\_name</code> DRR. Migration priority can be assigned at per virtual machine level. It determines:

* whether DRR should migrate a virtual machine and, if so,
* the priority in which it should be selected for migration.

### Migration Priority Values

Migration priority supports following values:

* **Unset** - This is the default state for migration priority value for virtual machines.
* **Normal** - Any virtual machines that do not have migration priority explicitly assigned, and that do not have soft affinity rules associated with them, will be treated with migration priority normal. The VMs are chosen for migration after the VMs with `high` priority value are migrated, but before the VMs with `low` priority value are migrated.
* **Low** - VMs with a priority value of `low` will be **selected** **last** for migration. You may therefore assign this value to virtual machines that incur a heavier tax on migration and that you would only want to live migrate if all other options are exhausted. VMs with soft affinity rules that do not have migration priority explicitly assigned will also be treated with low priority.
* **High** - VMs with a priority value of `high` will be **selected** **first** for migration. A good candidate for this category are dev/test VMs that may be smaller in size and / or their performance may not be impacted by live migration.
* **Excluded** - You can also set a migration priority value of `never` to a VM. VMs with this value will be not be migrated by the DRR service. You can use this for virtual machines that you do not want DRR to ever migrate.

**Note** that the migration priority values are only relevant in the context of DRR. These values are not taken into account by other <code class="expression">space.vars.product\_name</code> services such as [Virtual Machine High Availability](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) that may still evacuate the VM to a different host, if enabled at cluster level. An Administrator can also manually migrate a VM at any point independent of the migration priority value set for that VM.

{% hint style="warning" %}
**Important**

Migration priority values are only relevant in the context of DRR. They are not honored by other <code class="expression">space.vars.product\_name</code> services such as VM HA. Administrators can manually migrate a VM independent of the migration priority.
{% endhint %}

### Configure Migration Priority For VMs

You can change the migration priority value for any VM running on a virtualized cluster with DRR enabled.

* Select the VM in the virtual machines grid view by navigating to the 'virtual machines' menu in the <code class="expression">space.vars.product\_name</code> UI.
* From the action bar, choose 'Other actions', then choose 'Migration Priority'.
* You can now select the appropriate priority value for the VM.

## How DRR Works

DRR functions as an ongoing optimization engine for your cluster:

1. **Continuous Monitoring:** DRR continuously monitors key resource utilization metrics, specifically CPU and memory utilization, across all active hosts within the virtualized cluster.
2. **Imbalance Detection:** The system analyzes CPU and Memory utilization to identify imbalances in resource utilization across hosts.
   1. DRR looks for hosts that are overloaded in the cluster.
   2. **Host overload threshold**: When DRR finds a host with **CPU OR memory utilization of greater than 80%**, it determines that the host is overloaded.
   3. DRR then ensures that there are underutilized hosts with sufficient spare capacity in the cluster.
   4. Assuming it finds the required spare capacity, DRR will identify a suitable target host for migration. DRR will sort all compatible hosts in the cluster based on their utilization value, and **choose the host with lowest utilization**.
   5. DRR takes into account per host [overcommitment or over-allocation ratio](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#host-overcommitment-allocation-ratios) while finding a suitable candidate host.
   6. DRR then initiates VM migrations from overloaded hosts to other compatible hosts with spare capacity.
3. **Automated Live Migration:** DRR first groups VMs from the source host based on the VM Migration Priority.
   1. VMs with no priority value assigned to them are treated to have priority value of `default`
   2. VMs with priority value of `high` are chosen for migration first
   3. VMs with priority value of `default` are chosen next
   4. VMs with priority value of `low` are chosen last.
   5. VMs with priority value of `never` are skipped from migration.
   6. DRR will then initiate VM live migrations, **one VM at a time**. This ensures that there is *no downtime* for the virtual machine while it is being migrated.
   7. After each VM live migration, DRR re-evaluates if the source host is still overloaded. If true, it will continue with this process.

## DRR Interoperation with other services

This section describes how DRR interoperates with other services configured for your cluster.

### Host Aggregates

DRR only finds candidate target hosts for a VM that satisfy the specific VM's host aggregate requirement. If it does not find a suitable host, DRR will not migrate the VM.

### VM HA

DRR and Virtual Machine High Availability are designed to interoperate well together. It is possible that a host failure event might occur while DRR is actively rebalancing VMs, either from the same host or from other hosts in the cluster. When this happens:

1. VM HA will detect the host failure and initiate VM evacuations
2. The VM evacuations may result in cluster imbalance
3. DRR will then detect the imbalance during either the current run or the next run.
4. DRR will redistribute the load across the cluster to address the imbalance.

### VMs with Hard Affinity or Anti-Affinity Rules

1. **DRR will skip a VM with hard affinity rule** from migration.
   1. This behavior may change in the future once DRR has ability to initiate bulk VM migrations.
2. For a VM with hard anti-affinity rule, DRR will find suitable candidate host that satisfies the anti-affinity rule. If it can not find a suitable host, DRR will not migrate the VM.

For a VM with soft-affinity, DRR will try its best to find a target host that satisfies affinity, if not it will find another target host to migrate the VM. Also if migration priority is not set for the VM, DRR will consider the priority as low for the VM

For a VM with soft-anti-affinity, DRR will try its best to find a target host that satisfies anti-affinity, if not it will find another target host to migrate the VM. Also if migration priority is not set for the VM, DRR will consider the priority as low for the VM

### VM States

DRR **only live migrates VMs in Active or powered on state**. DRR will skip VMs in all other states.

### VMs with Special Properties

1. DRR will live migrate VMs that have hot added CPU or memory resources.

### Host States

DRR will **only operate on hosts in online state**. DRR will ignore hosts in offline or error state.

DRR will **only operate on hosts with hypervisor role assigned**. DRR will ignore all other hosts.

## Known Issues and Limitations

* DRR does not presently have the ability to live migrate VMs with [Virtual TPM](/private-cloud-director/virtualized-clusters/virtual-tpm) enabled today. This ability is coming soon.
* DRR will skip a VM with hard affinity rule from migration.
  * This behavior may change in the future once DRR has ability to initiate bulk VM migrations.
* For a VM with hard anti-affinity rule, DRR will find suitable candidate host that satisfies the anti-affinity rule. If it can not find a suitable host, DRR will not migrate the VM.


# Virtual Machine High Availability (VM HA)

In this document, you will learn about <code class="expression">space.vars.product\_name</code> **Virtual Machine High Availability (VM HA)**, a feature that automatically detects physical host failures within a cluster and restarts the affected VMs on other healthy hosts in the same cluster.

## Introduction

Hardware, firmware, or network issues can cause a virtualization host to go offline with little warning. Without an automated recovery mechanism, every VM on that host would remain down until an operator intervenes, violating most service-level objectives.

**Virtual Machine High Availability (VM HA)** protects against this risk. The process is designed to be automatic and requires minimal manual intervention during a failure event:

* Continuous Host Monitoring: The VM HA service continuously monitors the health and responsiveness of all hypervisor hosts in an HA-enabled virtualized cluster.
* Failure Detection: If a host stops responding (due to hardware failure, operating system crash, or certain network isolation scenarios), the system detects the failure
* Automatic VM Recovery: Upon confirmation of a host failure, which involves verifying the failure on both the management plane and the cluster hosts, VM HA automatically restarts any VMs running on the failed host. These VMs are powered on using available resources on the remaining healthy hosts within the cluster.

After recovery, complementary features such as **Dynamic Resource Rebalancing (DRR)** can redistribute load to restore optimal balance across the cluster, ensuring sustained performance.

## Benefits of VM HA

Enabling VM HA in your Private Cloud Director environment delivers the following benefits:

* **Minimized Downtime**: Automatically restarts or evacuates VMs when a host fails, reducing mean time to recovery from hours to minutes.
* **Service Continuity**: Keeps business-critical applications online even during unexpected infrastructure outages.
* **Operational Efficiency**: Eliminates the need for round-the-clock manual monitoring and intervention, freeing administrators to focus on higher-value tasks.
* **Policy-Driven Control**: Respects host aggregates, affinity/anti-affinity rules, and VM-specific settings, enabling you to determine how each workload is handled during failover.
* **Seamless Interoperation**: Works in concert with DRR and other <code class="expression">space.vars.product\_name</code> services to maintain both availability and resource efficiency after a failure event.

## VM HA Pre-Requisites

VM HA always operates within a virtualized cluster. You need to turn it on at the cluster level. Once enabled, VM HA applies to all virtual machines in the cluster.

* Shared storage is required for VM HA. [Shared storage requirements for VM HA](https://platform9.com/kb/pcd/storage/shared-storage-requirements-for-vmha) details the effect of various shared storage configurations on VM HA performance. For VM HA to work at the cluster level:
  * All VMs should be using a block storage volume as the root disk (non-ephemeral root disk), or
  * If any VMs use ephemeral storage for the root disk, [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) should be used for all hosts in the virtualized cluster, and the **This is Shared Storage** toggle is enabled in the **Cluster Blueprint** under the **Customize Cluster Defaults** tab.
  * When [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) is not used, any VMs using an ephemeral root disk will be rebuilt on another host during recovery.
* VM HA requires a minimum number of healthy hosts in a cluster to function correctly. A minimum of two hosts is required for HA activation.
* VM HA uses the VM Evacuation operation behind the scenes. VM Evacuation Prerequisites must be met for the operation to succeed.
* If any VMs in the cluster use a [Flavor](/private-cloud-director/virtualized-clusters/virtualmachine/vm-flavors) that assigns the VM to a host aggregate, then that host aggregate should have at least two hosts in the cluster for VM HA failover redundancy.
* If any VMs in the cluster use block storage, the block storage role must be assigned to at least two hosts in the cluster.
* If any VMs in the cluster use vTPM, then
  * The [virtual machine ephemeral storage directory](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint#virtual-machine-storage-path) and the [vTPM state file directory](/private-cloud-director/virtualized-clusters/virtual-tpm#vtpm-state-file) must be on shared storage (eg NFS) that is mounted on all of your hypervisor hosts in the cluster.
  * This is required as long as you are using vTPM as a feature for your VMs, even if the VMs do not otherwise use ephemeral storage for their root disk.
  * These directories must be owned by `pf9` user and `pf9group` group.
* The Image library role must be assigned to at least two hosts in the cluster, and the Image library must use shared storage.
* **Operating System Compatibility**: All hosts in the VM HA-enabled cluster must run the same operating system version.
  * Mixed Ubuntu versions (22.04 and 24.04) in the same cluster are not supported.
  * VM evacuation between different OS versions will fail due to the KVM version incompatibilities.
  * We recommend disabling VM HA before upgrading the operating system on any host of the cluster, then re-enabling it after the host(s) have been upgraded.
* Remember to read and understand the current VM HA before proceeding to configure it at the cluster level.

## Configure VM HA for a Cluster

VM HA configuration follows a two-step process to ensure cluster readiness:

#### Step 1: Create the Cluster

When you create a new cluster via **Infrastructure > Clusters > Add Cluster**, the **VM High Availability** toggle will appear but be disabled. You cannot enable VM HA during cluster creation.

#### Step 2: Enable VM HA After Adding Hosts

After creating the cluster:

1. Add at least two hosts with the hypervisor role to the cluster.
2. Verify that all hosts are assigned to the same host aggregate (if using host aggregates)
3. Navigate to the cluster settings and enable the **VM High Availability** toggle

{% hint style="info" %}
**NOTE**

The VM HA toggle is disabled until the minimum requirements are met. Hovering over the disabled toggle displays the message: "At least two hosts are required to enable VM HA."
{% endhint %}

### VM HA Enablement Rules

The system enforces these validation rules for VM HA:

* **Minimum Host Requirement**: At least 2 hypervisor hosts are required in the cluster.

{% hint style="info" %}
If fewer than two hosts are present, the toggle displays: *"At least two hosts are required to enable VM HA."*
{% endhint %}

* **Host Aggregate Consistency**: All hosts must either belong to the same host aggregate or have no aggregate assignment.

{% hint style="info" %}
If the hosts are part of different aggregates, the toggle displays the following message:*"VM HA cannot be enabled because hosts belong to different host aggregates, which prevents cross-host migration."*
{% endhint %}

On the **Clusters** grid view, you will see one of the following VM HA statuses for each cluster:

* **Disabled**: Shows the number of clusters for which VM HA is disabled.
* **Protected**: Shows the number of clusters with VM HA enabled that meet all prerequisites.
* **Degraded**: Shows the number of clusters with VM HA enabled that have a single point of failure (for example, only one host in the cluster has the persistent storage or image library role). Such clusters can absorb host failures as long as the host is not the single point of failure.
* **Not Protected**: Shows the number of clusters with VM HA enabled that have not met one or more prerequisites.

### Upgrading your Cluster Hosts

When upgrading your cluster hosts from one operating system version to another (for example, Ubuntu 22.04 to 24.04), follow the steps below:

* **Disable VM HA** before starting the upgrade process.
* **Evacuate all VMs** from the hosts that are scheduled for upgrade.
* **Upgrade hosts** one at a time to the new operating system version.
* **Do not attempt** to migrate VMs between upgraded and non-upgraded hosts.
* **Re-enable VM HA** only after all hosts have been upgraded to the same operating system version.

{% hint style="warning" %}
**Important**

VM HA-driven evacuation and live migration require matching operating system and KVM versions between source and destination hosts. Attempting to live migrate or evacuate VMs between Ubuntu 22.04 and 24.04 hosts, or between Ubuntu and Platform9 OS powered by Rocky Linux by CIQ hosts, will fail. **Cold migration** from Ubuntu to Platform9 OS hosts *is* supported; see [Cross-OS Cold Migration: Ubuntu Hosts to Platform9 OS Hosts](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#cross-os-cold-migration-ubuntu-hosts-to-platform9-os-hosts).
{% endhint %}

## How VM HA Works

<code class="expression">space.vars.product\_name</code> **Virtual Machine High Availability (VM HA)** provides automated protection against host failures by continuously monitoring host health and, when needed, evacuating affected virtual machines (VMs) to healthy hosts with minimal disruption. Three cooperating services enable the automated workflow:

1. **High Availability Manager**: Central coordinator that runs in the <code class="expression">space.vars.product\_name</code> management plane and watches for clusters with VHMA enabled, gathers health reports, confirms host failures, and issues evacuation requests.
2. **Host Agents: Daemons that run on every hypervisor host and probe their peers’ liveness and report findings to the High Availability Manager at regular intervals.**
3. **VM Evacuation Service:** Orchestrator that runs in the <code class="expression">space.vars.product\_name</code> management plane and receives confirmed *host‑down* events from **High Availability Manager** and migrates each impacted VM to a suitable target host.

### End-to-end Flow

1. **Cluster discovery:** The **High Availability Manager** polls the infrastructure to discover clusters where VHMA is enabled. For VM HA-enabled clusters, it verifies that the pf9-ha-slave role is present on every hypervisor host. Clusters with fewer than two hosts are ignored until additional hosts join.
2. **Peer list distribution:** Every 127  seconds, each Host Agent requests a *peer list* from the **High Availability Manager**. The list is a randomized subset of hosts in the same cluster, helping to spread traffic and avoid single‑point bias.
3. **Distributed health probing:** Using the peer list, the host agent performs a liveness check on its peer hosts via the libvirt exporter. This check is entirely agent‑to‑agent, with no central polling path.
4. **Status aggregation & reporting:** Every 59 seconds, the agent posts an aggregated view of peer health to the **High Availability Manager**.
5. **Failure detection and correlation:** When **the High Availability Manager** receives a report that a host appears to be *down*, it cross-checks the host’s status across reports from other agents to avoid false positives while ensuring a rapid response to true outages. Once a host is confirmed *down* via peer agent reports, a **150‑second cooldown** timer starts:
   * If the host recovers before the timer expires, the event is cleared, and no action is taken. This is to prevent routine reboots from triggering VM evacuations.
   * If the host remains *down* when the timer ends, **High Availability Manager** confirms the failure and emits a *host-down* notification to the **VM Evacuation Service**.
6. **Automated VM evacuation:** Upon receiving the *host‑down* notification, the **VM Evacuation Service** verifies the host’s state, collects all resident VMs, and evacuates them **one VM at a time** to a healthy host in the cluster. Target selection honors existing placement constraints (such as aggregates or affinity rules).
7. **Continuous protection loop:** After evacuation is complete, normal monitoring resumes. Host Agents continue probing both the recovered host (if it comes back online) and all remaining hosts, enabling VM HA to maintain an always-on protection loop with no manual intervention. Note that if a host that went offline comes back online, VM HA will not automatically migrate the VMs that were originally located on this host back to it. The Dynamic Resource Rebalancing (DRR) service will make VM-balancing decisions independently when resource contention occurs on any host.

## UI Observability

### VM HA Status Across Clusters

The PCD UI shows the current status of VM HA for each cluster at the following locations:

* The PCD home page displays a Cluster VM HA Status widget.
* The Infrastructure > Clusters page. Hovering over the VMHA status in the table displays the status for each prerequisite.
* The VM High Availability pane on the Cluster details page. Expanding the card will show the status for each individual prerequisite.

The possible VM HA statuses per cluster are:

* **Protected**: Shows the number of clusters with VM HA enabled that meet all prerequisites.
* **Degraded**: Shows the number of clusters with VM HA enabled that have a single point of failure (for example, only one host in the cluster has the persistent storage or image library role). Such clusters can absorb host failures as long as the host is not the single point of failure.
* **Not Protected**: Shows the number of clusters with VM HA enabled that have not met one or more prerequisites.
* **Disabled**: Shows the number of clusters for which VM HA is disabled.

### VM HA Events

When an active VM HA event is in progress, a banner message is displayed in the PCD UI with details about the event. The banner message will persist until evacuations are complete or the banner is manually dismissed.

The Host details page for a host that experienced an outage shows the evacuation status for all VMs on the host. The details will be displayed for 24 hours after the host outage. The VMHA Past Events table on the Host details page will show all historical VM HA events that occurred on the host. Clicking View Details displays detailed VM evacuation status for the selected event. From the detailed VM evacuation status page, you can retry failed evacuations (for all failed evacuations or for individual VMs) for the most recent VM HA event.

Note that VMs in error, unknown, rescued, or resized status (i.e., the VM has been resized but the resize has not been confirmed) will not be evacuated.

## VM HA Interoperation with Other Services

This section describes how VM HA interoperates with other services configured for your cluster.

### Host Aggregates

VM HA will honor [Host Aggregates](/private-cloud-director/virtualized-clusters/host-aggregate) and migrate VMs to another host from the same host aggregate.

Here are some limitations you might need to consider:

* VM HA requires that all hosts in a cluster belong to the same host aggregate or have no aggregate assignment. Mixed aggregate configurations will prevent VM HA from being enabled.
* All hosts from a host aggregate must belong to a single cluster and not span multiple clusters.
* If a host aggregate has only one host that goes down, VM HA cannot find a suitable migration target, causing VMs to enter an error state.

### DRR

DRR and Virtual Machine High Availability are designed to interoperate well together. A host failure event may occur while DRR is actively rebalancing VMs, either from the same host or from other hosts in the cluster. When this happens:

1. VM HA will detect the host failure and initiate VM evacuations
2. The VM evacuations may result in cluster imbalance
3. DRR will then detect the imbalance during either the current run or the next run.
4. DRR will redistribute the load across the cluster to address the imbalance.

### VMs with Hard Affinity or Anti-Affinity Rules

* **Hard affinity**: VM HA will temporarily break hard affinity to evacuate VMs with hard affinity. The evacuation process will attempt to locate a suitable target host to accommodate all VMs belonging to the hard affinity server group.
* **Soft affinity**: VM HA identifies a host with sufficient capacity to host all VMs in the affinity group and migrates all VMs sequentially to the target host. If VM HA cannot find a suitable target host, VM HA will migrate VMs to available hosts, potentially violating the soft affinity policy.
* **Hard anti-affinity**: VM HA will fail to evacuate VMs with hard anti-affinity, and VMs will land in `Error` state. This will be addressed in an upcoming release.
* **Soft anti-affinity**: VM HA will attempt to find target hosts that satisfy the anti-affinity policy for each VM to be migrated, ensuring that all VMs part of the soft anti-affinity policy are placed on separate hosts. If VM HA cannot find a suitable target host, VM HA will evacuate the VMs to available hosts, potentially violating the soft anti-affinity policy.

{% hint style="info" %}
**NOTE**

If VM evacuation fails in any of the scenarios listed above, the VM may still appear as active in the UI.
{% endhint %}

### VM States

VMs in a suspended or paused state will not be evacuated to another host.

### VMs with Special Properties

1. **Virtual TPM-enabled VMs:** VM HA will evacuate VMs with [Virtual TPM](/private-cloud-director/virtualized-clusters/virtual-tpm) enabled, provided the [vTPM HA prerequisites](#vm-ha-pre-requisites) are met.

**Resized VMs:** VM HA will skip evacuating a VM that has been resized but not confirmed yet.

## Supported Scale

VM HA is currently supported for up to 400 hosts per region.

## Known Issues and Limitations

* VM HA is presently not supported for clusters with GPUs enabled. This support will be added in a future release of Private Cloud Director
* VM HA relies on connectivity with the management plane. If network connectivity to the management plane is degraded, VM HA performance will be negatively affected. If the management plane experiences an outage, VM HA will not be operational for the duration of the outage.

## Troubleshooting VM HA

For step-by-step diagnostics when VM HA does not behave as expected, see the VM HA troubleshooting runbook:

[Troubleshoot VM HA](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshoot-vm-ha)

The runbook covers:

* A host failed but VMs were not evacuated — how to verify cluster settings, host detection, HA agent health, and libvirt exporter liveness.
* Consul health prerequisites — how to check and recover the distributed coordination service that VM HA depends on for failure detection.
* Shared and FC storage validation — how to confirm that storage is reachable on all hosts before relying on VM HA.
* Errors when enabling or disabling VM HA — prerequisites to verify and how to recover from 404 or 503 responses.
* Post-upgrade re-validation — how to confirm VM HA is functional after a host or management-plane upgrade.


# Host

Learn how to add and configure Hosts in Private Cloud Director for your virtualized cluster. Discover Host roles, management, and troubleshooting tips to ensure optimal performance and uptime for your

A Host is a physical machine that you supply to <code class="expression">space.vars.product\_name</code> as a hypervisor. Each Host contains the resources needed for your cluster, such as virtual machines, storage, and networking components. Once authorized and configured, you can deploy virtual machines on top of the Host.

You can add multiple Hosts to your <code class="expression">space.vars.product\_name</code> virtualized cluster. After you configure your Cluster Blueprint, <code class="expression">space.vars.product\_name</code> has the information it needs to configure Hosts that you add to the cluster.

Learn how you can add a Host to your virtualized cluster, assign roles, and configure it for production use. You will also learn how to manage the Host lifecycle and troubleshoot common issues.

## Hypervisor

<code class="expression">space.vars.product\_name</code> standardizes on and uses open source KVM hypervisor behind the scenes. KVM is a type 1 hypervisor that runs directly within the Linux kernel on the host hardware. KVM hypervisor relies on and leverages hardware virtualization extensions from x86 and/or AMD processors to provide full virtualization capabilities.

## Resource Management

Hypervisor resource management enables you to make the best use of CPU and Memory resources available across all your hypervisor nodes within a virtualized cluster, to ensure that:

* Your applications get the required resources when they need, to meet your business SLA.
* You make the most optimal utilization of available resources within your cluster, and prevent or reduce and waste.

Resource management in <code class="expression">space.vars.product\_name</code> has multiple components to it. At the Hypervisor level, resource management is handled via a combination of the following components.

### Memory Ballooning

Memory ballooning is a KVM hypervisor technique for dynamic memory management that is used to reduce the impact of memory-overcommitment on the hypervisor load. Memory ballooning allows your guest OS to dynamically evict unused pages of virtual memory, so that the KVM host can then share the unused memory with other VMs, allowing the host to overcommit memory and optimize resource use by giving memory to VMs that suddenly need it. <code class="expression">space.vars.product\_name</code> hypervisor makes use of a combination of memory ballooning and memory swapping to enable optimal utilization of host memory.

### Memory Deduplication & Consolidation

Memory deduplication and consolidation is a memory management feature used by KVM hypervisor to consolidate identical memory pages to achieve higher virtual machine density. KVM hypervisor achieves this using a linux feature called **Kernel Samepage Merging (KSM)**, that finds identical memory pages across different virtual machines and merges them into a single, shared, Copy-on-Write (CoW) page, significantly reducing overall memory usage. KSM enables <code class="expression">space.vars.product\_name</code> to improve VM memory density on a host without impacting performance.

### Resource Overcommitment / Allocation Ratios

### Memory Overcommitment (Allocation Ratio)

<code class="expression">space.vars.product\_name</code> allow you to overcommit memory (also called memory overprovisioning or memory allocation ratio) at each host level, enabling you to effectively share the host level memory resources across virtual machines. Each new hypervisor host provisioned in <code class="expression">space.vars.product\_name</code> is configured with a default memory **overcommitment ratio of 1:1.5**. This means by default each host's memory is overcommitted by 50%. You can change this configuration by navigating to the 'Cluster Hosts' view in the <code class="expression">space.vars.product\_name</code> UI, then selecting a specific host, and editing the allocation ratios.

The KVM hypervisor uses a combination of memory ballooning and memory swapping to manage memory across running virtual machines.

### CPU Overcommitment & Time-Sharing (Allocation Ratio)

<code class="expression">space.vars.product\_name</code> allows you to overcommit CPU resources (also called CPU overprovisioning or CPU allocation ratio) at each host level, enabling you to effectively share the host's CPU cores across running virtual machines.

Behind the scenes, the KVM hypervisor dynamically allocates physical CPU cores to running virtual machines, context switching between different virtual machines, to enable sharing and overcommit.

Each new hypervisor host provisioned in <code class="expression">space.vars.product\_name</code> is configured with a default CPU **overcommitment ratio of 1:16**. This means by default each host's CPU is overcommitted by 16x. We choose this default setting as most virtualized workloads tend to be memory bound and can share CPU without impacting workload performance. But when running CPU heavily workloads, you will need to change this default overcommitment ratio to suite your application's needs. You can also utilize features like CPU pinning to reduce the context switching overhead for CPU heavy workloads.

### Ephemeral Storage Overcommitment (Allocation Ratio)

<code class="expression">space.vars.product\_name</code> also allows you to overcommit the [Ephemeral Storage](/private-cloud-director/storage/ephemeral-storage#what-is-ephemeral-storage) at per host level (also called ephemeral disk overprovisioning or ephemeral disk allocation ratio). This capability makes it possible for you to over-allocate the disk storage at per hypervisor host level that is used to store root disks for VMs using ephemeral storage for their root disks. Useful when you do not expect all VMs using ephemeral storage to use all their disk space.

For example, lets say that your VM ephemeral storage path on your hypervisor is set to the default location of `/opt/data/instances` and this path is on your root partition which has 100GB of total disk space. If you set the ephemeral disk overcommitment value to 2, this means you can create virtual machines with the sum total of root disk size not exceeding 200GB. This can become a problem if all the VMs start fully utilizing their allocated root disk space.

Each new hypervisor host provisioned in <code class="expression">space.vars.product\_name</code> is configured with a default ephemeral storage **overcommitment ratio of 1:9999**. This means by default each host's ephemeral disk is overcommitted by 9999x. We set it to such a high value by default because most production environments use block storage volumes for production VM root disks and VMs that use ephemeral storage for their root disk often tend to not utilize it fully. You will need to update this default value if you plan to use and rely on ephemeral root disk storage for production virtual machines.

{% hint style="warning" %}
Each new hypervisor host is configured with a default ephemeral storage **overcommitment ratio of 1:9999**.

You must **update this default** if you plan to use and rely on ephemeral root disk storage for production virtual machines. We recommend updating it to 1 for no overcommitment, or to a reasonable overcommitment ratio based on the nature of your VM workloads and their expected use of root disk.
{% endhint %}

### View Host CPU, Memory and Storage Allocation and Utilization Info

The current allocation and utilization values for CPU, Memory and Epheremeral Strorage for each host are reported as individual columns as part of the Cluster Hosts grid view in the UI.

* For Compute:
  * **% CPU utilized** - this is the current CPU utilization for this host. This value is an aggregate across all physical CPU cores on the host and is reported in terms of total vs available MHz.
  * **% VCPUs allocated** - this represents the sum total of VCPUs that have been allocated across all virtual machines provisioned on this host. For eg, if the host has 5 VMs currently deployed on it with 4 CPUs each, you will see the VCPUs allocated value to be 20. If the host has 10 physical CPUs and has overcommitment ratio of 1:2, then you can allocate a maximum of 20 VCPUs on this host.
* For Memory:
  * **% Memory utilized** - this is the current Memory utilization for this host, reported as total vs available GiB.
  * **% Memory allocated** - this represents the sum total of RAM that have been allocated across all virtual machines provisioned on this host. For eg, if the host has 5 VMs currently deployed on it with 4 GB RAM each, you will see the Memory allocated value to be 20. If the host has 20 GB physical RAM and has overcommitment ratio of 1:1, then you can allocate a maximum of 20 GB of memory across all VMs on this host.
* For Storage:
  * **% root disk used** - this is the current root disk utilization for this host, reported as total vs available GiB.
  * **% ephemeral disk used** - this represents the sum total of ephemeral root disks that have been allocated across all virtual machines provisioned on this host that are using ephemeral root disks. For eg, if the host has 5 VMs currently deployed on it, each using 10 GB of ephemeral root disk, you will see the ephemeral disk used value to be 50GB. If the host has 250 GB physical disk and has ephemeral disk overcommitment ratio of 1:1, then you can allocate a maximum of 250 GB of ephemeral root disk across all VMs on this host.

### View and Edit Overcommitment / Allocation Ratios

You can edit CPU, memory or ephemeral disk overcommitment values (allocation ratios) at per host level by following these steps:

1. Navigate to the Cluster Hosts view in the UI,
2. Select a specific host, then click on the host name to go to the host details view. Here you will see the currently set values for allocation ratios for CPU, Memory and Ephemeral storage for this host under "Properties" section.
3. To edit the current values, go back to the Host grid view, select the host then choosing "Edit allocation ratio" as the action from the action bar.
4. Here you can see again the currently configured overcommitment values for CPU, Memory and ephemeral disk for this host.
5. You can reset these values to their system default or change specific values here.

## Understanding Host Agent and Roles

### Host Agent

The Platform9 Host agent is the first component you install for each Host. The Host agent enables you to add Hosts and configure their roles in your virtualized cluster. Based on the assigned role to each Host, the agent downloads and configures the required software, integrating with the <code class="expression">space.vars.product\_name</code> management plane.

The Host agent also provides ongoing health monitoring of the Host, including the detection of failures and errors. It helps Platform9 orchestrate upgrades when you choose to upgrade your <code class="expression">space.vars.product\_name</code> deployment to a newer version.

### Host Roles

As part of the Host authorization process, you can configure Hosts to perform specific functions by assigning them one or more roles. The following roles are supported:

**Hypervisor**

The Hypervisor role enables the Host to function as a KVM-based hypervisor in your virtualized cluster. It is recommended you assign this role to all Hosts in your cluster, unless you experience performance bottlenecks and want to avoid running VM workloads on select Hosts that have other roles such as image library or storage roles.

**Image Library**

Every cluster needs at least one Image Library, which hosts the cluster copy of virtual machine source images from which you can provision new VMs. See more on [Configuring Image Library Role here](/private-cloud-director/images-and-image-library/image-library---images#configuration-).

**Persistent (Block) Storage**

You can typically configure one or more Hosts in your cluster as a **Block Storage Node**. See more on [Block Storage Service Configuration](/private-cloud-director/storage/block-storage#block-storage-service-configuration).

**Advanced Remote Support**

For specific troubleshooting situations, Platform9 support teams may request access to gather detailed telemetry from a Host that experiences problems automatically. This mode is turned off by default. To enable Advanced Remote Support, contact Platform9 Support.

**DNS**

Enables DNS as a Service (DNSaaS), which is an optional component.

## Prerequisites

Before adding a Host, ensure that your Host meets the prerequisites and that you have configured in your Cluster Blueprint.

* Verify that your Host meets the [Pre Requisites](/private-cloud-director/getting-started/pre-requisites) for <code class="expression">space.vars.product\_name</code>.
* Ensure you have administrative access to the Host.
* Confirm that your [Virtualized Cluster Blueprint](/private-cloud-director/virtualized-clusters/virtualized-cluster-blueprint) is configured in the <code class="expression">space.vars.product\_name</code> console.
* Confirm that you have created a Cluster in the <code class="expression">space.vars.product\_name</code> console.

## Add a Host

To add a host to <code class="expression">space.vars.product\_name</code>, you need to follow the steps to install the <code class="expression">space.vars.product\_name</code> agent software on your physical host, then follow the steps to add it to <code class="expression">space.vars.product\_name</code>.

#### Step 1: Add a Host

The process for adding a host varies depending on which Platform9 product you are using. Choose the appropriate method below based on your deployment type:

* SaaS Deployment
* Self-Hosted Deployment
* Community Edition Deployment

{% tabs %}
{% tab title="SaaS Deployment" %}
For SaaS deployments, adding a Host is on the <code class="expression">space.vars.product\_name</code> console with minimal configuration requirements.

1. Navigate to **Infrastructure > Cluster Hosts** on <code class="expression">space.vars.product\_name</code> console.
2. Select **Add New Hosts**
3. Follow the on-screen instructions. You are required to run these [pcdctl](/private-cloud-director/reference/pcdctl-command-line) commands using `sudo` privileges. The command requires you to add values specific to your environment that are provided on your <code class="expression">space.vars.product\_name</code> console. It takes about 2-3 minutes to download and install the Platform9 Host agent and other necessary Platform9 software components.

You have now successfully added a Host to your SaaS deployment. Continue to **Step 2: Authorize Host and Assign Roles**
{% endtab %}

{% tab title="Self-Hosted Deployment" %}
For self-hosted deployments, you may need to manually configure DNS entries before adding the Host.

**Step 1: Add a Host to Self-Hosted Deployment**

{% hint style="info" %}
**NOTE**

Hypervisor Hosts deployed as virtual machines must have virtualization support available inside the VM. Virtual machines on ARM CPUs are currently untested.
{% endhint %}

1. Verify nested virtualization

You can choose to verify how the nested virtualization works in a VM. Check for virtualization support inside the VM by running:

{% tabs %}
{% tab title="Bash" %}

```bash
egrep "svm|vmx" /proc/cpuinfo
```

{% endtab %}
{% endtabs %}

2. Add DNS entries to each Host

An FQDN is a fully-qualified domain name for your <code class="expression">space.vars.product\_name</code> installation. You will need both infrastructure and workload region FQDNs for your self-hosted deployment.

As a root user, add DNS entries on the hypervisor Host for both the infrastructure and workload region FQDNs: `/etc/hosts` file:

{% tabs %}
{% tab title="Bash" %}

```bash
echo "<IP> <FQDN-infrastructure-region>" | tee -a /etc/hosts
echo "<IP> <FQDN-workload-region>" | tee -a /etc/hosts
```

{% endtab %}
{% endtabs %}

Replace `<IP>` with your management plane IP address and the FQDNs with your specific domain names.

Here is a sample example:

{% tabs %}
{% tab title="Bash" %}

```bash
echo "10.9.11.246 pcd.pf9.io" | tee -a /etc/hosts
echo "10.9.11.246 pcd-community.pf9.io" | tee -a /etc/hosts
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Important!**

Before adding a hypervisor Host, ensure that you have saved a Cluster Blueprint and created a Cluster on the <code class="expression">space.vars.product\_name</code> console.
{% endhint %}

* Navigate to **Infrastructure > Cluster Hosts** and the select **Add New Hosts**.
* Follow the on screen instructions. You are required to enter the administrative user password when prompted. For more details on `pcdctl` CLI see [Pcdctl Command Line](/private-cloud-director/reference/pcdctl-command-line)

You have now successfully prepared your Host for self-hosted deployment. Continue to **Step 2: Authorize Host and Assign Roles**
{% endtab %}

{% tab title="Community Edition Deployment" %}
For Community Edition deployments, the process is similar to self-hosted but uses a specific FQDN for the user interface, unless configured otherwise via a Custom Install.

**Step 1: Add a Host to Community Edition Deployment**

{% hint style="info" %}
**NOTE**

Hypervisor Hosts deployed as virtual machines must have virtualization support available inside the VM. Virtual machines on ARM CPUs are currently untested.
{% endhint %}

1. **Verify nested virtualization (for VM Hosts)**

If you want to verify that nested virtualization works in a VM, check for virtualization support inside the VM:

{% tabs %}
{% tab title="Bash" %}

```bash
egrep "svm|vmx" /proc/cpuinfo
```

{% endtab %}
{% endtabs %}

If the command returns results, your VM supports nested virtualization and can run other virtual machines.

2. **Configure DNS entries**

By default, Community Edition uses `pcd.pf9.io` for the user interface and to reference the workload region, unless you have customized it [Custom Installation](/private-cloud-director/getting-started/getting-started-with-community-edition/custom-installation).

* Workload region & UI FQDN: `pcd.pf9.io`

3. **Add DNS entries to your Host**

Log in to your Host as a root user and add DNS entries for the FQDN by running the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
echo "<IP> <FQDN>" | tee -a /etc/hosts
```

{% endtab %}
{% endtabs %}

Replace `<IP>` with your management plane IP address and the `FQDN` with your specific domain name.

Here is a sample example:

{% tabs %}
{% tab title="Bash" %}

```bash
echo "10.9.11.246 pcd.pf9.io" | tee -a /etc/hosts
```

{% endtab %}
{% endtabs %}

4. **Add the Host through the** <code class="expression">space.vars.product\_name</code> **console.**

{% hint style="info" %}
**Important!**

Before proceeding, ensure that you have saved a Cluster Blueprint and Created a cluster in the <code class="expression">space.vars.product\_name</code> console.
{% endhint %}

* Navigate to **Infrastructure > Cluster Hosts** and then select **Add New Hosts**.
* Follow the on screen instructions. You are required to enter the administrative user password when prompted. For more details on `pcdctl` CLI see [Pcdctl Command Line](/private-cloud-director/reference/pcdctl-command-line)

You have now successfully prepared your Host for Community Edition deployment. Continue to **Step 2: Authorize Host and Assign Roles**
{% endtab %}
{% endtabs %}

#### Step 2: Authorize Host and Assign Roles

A successfully configured Host is accessible on **Infrastructure > Cluster Hosts** with an `Unauthorized` status, indicating an authorization and cluster role assignment.

1. Navigate to **Infrastructure > Cluster Hosts** and select a specific Host.
2. Select **Edit Roles** to configure the appropriate roles for the Host based on your cluster architecture.
3. Configure roles based on your requirements:

* **Hypervisor Role**: Enables the Host to function as a KVM-based hypervisor. Read more on [Hypervisor Role](#hypervisor-role).
* **Networking Service Configuration:** Select appropriate Host Network Config for the Host networking requirements. Read more on [Networking Service Configuration](/private-cloud-director/virtualized-networking/networking-overview#networking-service-configuration).
* **Image Library Role:** Configures the Host to store VM images for the cluster. Read more on [Configuring Image Library Role here](/private-cloud-director/images-and-image-library/image-library---images#configuration-).
* **Block Storage Role:** Enables the Host to provide persistent storage services. Read more on [Configure a Host with Block (Persistent) Storage Node Role](/private-cloud-director/storage/volume#configure-a-host-with-block-persistent-storage-node-role).
* **Advanced Remote Support:** Enables Platform9 support to gather detailed telemetry for troubleshooting purposes. Read more on [Enabling Advanced Remote Support](#enable-advanced-remote-support).
* **DNS Checkbox**: Enables DNS as a Service (DNSaaS), which is an optional component. The DNS checkbox is for [Dns As A Service Dnsaas ](/private-cloud-director/virtualized-networking/dns-as-a-service-dnsaas). Read more on [Configuring DNS-as-a-Service](/private-cloud-director/virtualized-networking/dns-as-a-service-dnsaas#configuration).

<figure><img src="/files/2NUAtoxUwgDnCQR2Hha4" alt=""><figcaption></figcaption></figure>

4. Select **Update Role Assignment**

The <code class="expression">space.vars.product\_name</code> management plane works with the Platform9 agent installed on your Host to configure the required software. This process typically takes 3-5 minutes to complete. During this time, your Host status changes to `converging` in the <code class="expression">space.vars.product\_name</code> console.

You have successfully authorized your Host and assigned roles. Your Host is now being configured for use in your cluster.

#### Step 3: Monitor Host addition status

While your Host is in the `converging` state, you can monitor the configuration progress by examining the Host agent log files.

1. Locate the Host agent log file.

The Host agent log files are located on your Host. See the [Log Files](https://docs.platform9.com/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files#log-files) section for detailed information about log file locations. The primary Host agent log is located at `/var/log/pf9/hostagent.log` on your Host.

2. Monitor the configuration progress.

Tail the log file to monitor the status of host addition:

{% tabs %}
{% tab title="Bash" %}

```bash
tail hostagent.log
```

{% endtab %}
{% endtabs %}

3. Verify successful completion

Monitor the log output for completion indicators. Once the Host is authorized and the role assignments have taken effect, your Host status changes from `converging` to `ok` in the <code class="expression">space.vars.product\_name</code> console.

Your Host is now ready to use and can run workloads according to its assigned roles.

## Manage Host lifecycle

### Remove a Host

Entirely removing a Host from your <code class="expression">space.vars.product\_name</code> setup is a two-step process. You must first remove all roles assigned to the Host, and then, if necessary, decommission the Host to clean up any <code class="expression">space.vars.product\_name</code> related data associated with it. You must perform both steps if you plan to re-add the Host to your current or another product name setup.

#### Step 1 - Remove all roles and deauthorize a Host

Removing all roles from a Host is the first step toward entirely removing a Host from your <code class="expression">space.vars.product\_name</code> setup. Removing all roles uninstalls any specific packages and software components assigned to the Host.

**Prerequisites before removing a Host from** <code class="expression">space.vars.product\_name</code> setup.

* If the host is assigned `hypervisor` role, make sure that no VMs are running on the host using `sudo virsh list --all` The expected output is an empty list of VM UUIDs and their corresponding statuses.
* If the host is assigned a `persistent storage` role, make sure that this host is serving no storage volumes.
* If the host is assigned `image library` role, ensure that the host is serving no images in the image library. On the image library host, check if the image UUID exists in the `/var/lib/glance/images/glance` default directory, or if a custom image directory was added, check that location.

1. Navigate to the **Cluster Hosts** on the console.
2. Select the specific Host.
3. From the Actions bar dropdown, select **Remove all roles**.

**Remove all roles using CLI:**

Use the following `pcdctl` command to remove all roles from a Host:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl deauthorize-node
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**NOTE**

Removing all roles only uninstalls and removes software packages associated with all roles from the Host. It does not clean up and delete any directories or files created as part of the installation of these software components. You must decommission a Host for all <code class="expression">space.vars.product\_name</code> related data from that Host to be removed.
{% endhint %}

#### Step 2 - Decommission a Host

When you remove all roles from a Host using the command above, any <code class="expression">space.vars.product\_name</code> specific packages and software components associated with those roles are uninstalled and removed from the Host. However, any directories where the packages were installed are not cleaned up or deleted. This ensures that you still have access to the log files for those components if required for debugging.

To remove these directories and clean up any <code class="expression">space.vars.product\_name</code> related data from the Host, you need to run the decommission Host command.

**Prerequisites:**

* You must remove all roles from the Host using the <code class="expression">space.vars.product\_name</code> console or CLI before decommissioning.
* Always back up important data, such as log files and configuration files, from the Host before decommissioning.
* You must decommission a Host before you can add it again to your current or any other <code class="expression">space.vars.product\_name</code> setup. Not doing so results in problems when re-authorizing the Host in the <code class="expression">space.vars.product\_name</code> setup.

**Decommission using CLI:**

Currently, decommissioning a Host can only be performed using `pcdctl` CLI.

Use `pcdctl` to decommission a Host by running the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
$ pcdctl decommission-node

Do you wish to decommission the node? (y/n) y
Checking if any roles exist on the host
Cleaning up the node
Decommission Successful
```

{% endtab %}
{% endtabs %}

Once the command executes successfully, the Host is removed from the list of active Hosts in the <code class="expression">space.vars.product\_name</code> console. If you encounter errors during decommissioning, check the logs for details and ensure the Host is reachable.

## Host Properties

### Host ID

Each hypervisor host gets a system-assigned ID when it's created. By default, the ID value is not shown in the host grid UI, but you can view the ID information by clicking on the 'Manage Columns' button on the cluster hosts grid view, then selecting the ID field to be displayed. You can also view a host's ID on the host details view by clicking on the host name from the host grid view. You can also query it from the `pcdctl` CLI by running `pcdctl hypervisor list` command or `pcdctl hypervisor show <hypervisor-name>` command where `<hypervisor-name>` is the name of your hypervisor host.

### Host Connection Status

Host connection status, represented by the 'Connection Status' column in the Cluster Hosts grid in the UI, represents the status of connectivity between the host agent and the <code class="expression">space.vars.product\_name</code> management plane. The following are the different status values:

* **online** - The Host agent is connected to the <code class="expression">space.vars.product\_name</code> management plane and responds to heartbeats.
* **offline** - The Host agent is unable to connect to the <code class="expression">space.vars.product\_name</code> management plane due to the host being in a powered-off state or because the Host is running. However, the Platform9 host agent may be experiencing issues connecting with the management plane. For more information on debugging steps, refer to [Host Issues](https://platform9.com/docs/private-cloud-director/troubleshooting/hypervisor-or-host-issues).

### Host Role Status

Host role status, represented by the 'Role Status' column in the Cluster Hosts grid in the UI, indicates whether the host has been added to any virtualized cluster and assigned roles within the cluster. The following are the different role status values:

* **unauthorized -** The host has <code class="expression">space.vars.product\_name</code> host agent installed but has not yet been added to a cluster or assigned any specific roles.
* **applied** - The host is assigned to a cluster, has specific roles assigned to it within that cluster, and those roles have been successfully applied.
* **converging** - A new role is being applied to the host and/or the host is being authorized and added to a cluster
* **failed / error -** The host appears to be in a failed or error state. For more information, refer to [Host Issues](https://platform9.com/docs/private-cloud-director/troubleshooting/hypervisor-or-host-issues).
* **unknown** - The role status is displayed as unknown when the host connection status is offline, as it does not know if the application of any roles was in progress or if it was successful.

### Host OverCommitment / Allocation Ratios

You can update the overcommitment / allocation ratios for individual hosts by navigating to the Cluster Hosts grid view in the UI, slelecting a specific host, then selecting "Update allocation ratios" from the actions menu.

Read more about [#memory-overcommitment-allocation-ratio](#memory-overcommitment-allocation-ratio "mention") [#cpu-overcommitment-and-time-sharing-allocation-ratio](#cpu-overcommitment-and-time-sharing-allocation-ratio "mention") and [#ephemeral-storage-overcommitment-allocation-ratio](#ephemeral-storage-overcommitment-allocation-ratio "mention").

### Storage Initiator Identifiers (IQN and WWN)

Each host exposes the storage initiator identifiers that backends use to authenticate connections. Both appear as columns on the **Cluster Hosts** grid under **Infrastructure > Cluster Hosts**.

* **IQN** (iSCSI Qualified Name) – The host's iSCSI initiator name. iSCSI storage backends use this to recognize and authorize the host. Format: `iqn.YYYY-MM.<reverse-domain>:<unique-suffix>`.
* **WWN** (World Wide Name) – The host's Fibre Channel initiator port identifier. FC backends use this for zoning and host access configuration on the array side.

Use these when configuring host access on the storage array or when correlating a volume's **Management Host** (see [Volume](/private-cloud-director/storage/volume#placement-information)) with the array's view of attached hosts.

## View Volumes Mounted on a Host

Each host's details page includes a **Volumes** tab that lists the volumes currently mounted on virtual machines running on that hypervisor. Navigate to **Infrastructure > Cluster Hosts**, select a host, then open the **Volumes** tab.

The **Mounted Volumes** table shows, per volume: **Name**, **Status**, **Management Host** (the block storage host serving the volume), **Volume Backend**, **Type**, **VM** (the virtual machine the volume is attached to, linked to the VM details page), **Hypervisor Host** (matches the host you're viewing), **Device** (the guest device path, for example `/dev/vda`), **Snapshots**, **Size**, and **Bootable**.

Use this view to scope a storage problem to one hypervisor: which volumes is it currently serving, which VMs depend on them, and which backends are in play. Toggle **View Volumes in All Tenants** to widen the scope beyond the current tenant.

## Host Aggregates

A host aggregate is a **group of hosts** within your virtualized cluster that share common characteristics. Read [Host Aggregate](/private-cloud-director/virtualized-clusters/host-aggregate) for more information on how to configure them.

## Debugging Compute Service Problems

Follow [Troubleshooting And Log Files](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files) for steps to troubleshoot issues with the compute service.


# Maintenance Mode

<code class="expression">space.vars.product\_name</code> maintenance mode enables administrators to perform routine maintenance operations on their fleet of hosts in the data center in a safe and orchestrated manner. Maintenance mode allows you to safely migrate virtual machines from a host and prevent new virtual machines from being scheduled during maintenance operations.

## What Happens in Maintenance Mode?

When maintenance mode is enabled for a host:

* All running VMs on the host are automatically live-migrated to another compute host in the cluster.
  * You have the option to choose the host to which the VMs will get migrated
* The host is marked as unschedulable, ensuring no new VMs are placed on it.
* Once all VM migrations are complete, the host remains online but isolated from scheduling operations, allowing administrators to perform software updates, hardware maintenance, or diagnostics.

## Enabling Maintenance Mode

To enable maintenance mode on a host:

1. Navigate to Infrastructure → Cluster Hosts in the <code class="expression">space.vars.product\_name</code> UI.
2. Select the target host from the list (optional). If no target host is selected, maintenance mode will select an appropriate target host from the cluster.
3. Click Other → Enable Maintenance Mode.

As noted in point 2 above, you may choose to:

* Allow <code class="expression">space.vars.product\_name</code> to automatically determine the best destination for each VM based on cluster capacity and placement policies.
* Manually specify a destination host for all VM migrations.

Click Migrate VMs and Enable Maintenance Mode to initiate the process.

After host maintenance operations are completed, you can disable maintenance mode to allow the host to rejoin the scheduler.

## Disabling Maintenance Mode

Before disabling maintenance mode, ensure that any maintenance tasks on the host have been completed successfully and that the host is healthy and ready to handle new workloads.

1. Navigate to Infrastructure → Cluster Hosts.
2. Locate the host that is currently in Maintenance Mode.
   1. Hosts in maintenance mode are typically marked with a "Maintenance Mode" label or status indicator.
3. Select the host, then click on Other → Disable Maintenance Mode

The host will be marked as schedulable again and <code class="expression">space.vars.product\_name</code> will resume placing new virtual machines on this host.

## UI Observability

The Private Cloud Director UI provides real-time status indicators and detailed migration progress information throughout the maintenance mode lifecycle.

### Maintenance Mode Status on the Cluster Hosts Page

The **Infrastructure → Cluster Hosts** page reflects the current maintenance mode state for each host:

* **Scheduling column**: When maintenance mode is enabled on a host, the Scheduling column displays **Disabled** to indicate that the host is no longer accepting new VM placements.
* **Maintenance Mode label**: During the transition into maintenance mode, a status label displays **Entering Maintenance Mode** in the Scheduling column. This label is visible in the hosts table for any host that is currently undergoing the maintenance mode process.
* **Other Actions menu**: When a host is in maintenance mode, the **Enable Maintenance Mode** option in the Other Actions dropdown is disabled, and the **Disable Maintenance Mode** option becomes available. The **Enable/Disable Scheduling** option is also disabled while maintenance mode is active, since scheduling is managed by the maintenance mode process.

### Maintenance Mode Status on the Host Details Page

Clicking on a host that is in maintenance mode opens the Host details page, which displays a banner at the top of the page with the following information:

* **Migration status summary**: The banner indicates whether VM migrations are in progress or have completed successfully.
* **VM count**: The total number of VMs that were migrated out of the total that will be migrated.
* **Timestamps**: The start time of the maintenance mode operation is displayed. The completion time is displayed once all migrations are finished.
* **See Details**: A button that opens the View Migration Progress panel with detailed per-VM migration status.
* **Disable Maintenance Mode**: A button that allows you to disable maintenance mode directly from the Host details page without navigating back to the hosts list.

### View Migration Progress

Clicking **See Details** on the Host details page banner opens the View Migration Progress panel. This panel provides a detailed breakdown of the maintenance mode operation:

* **Overall Migration Status**: Displays the current state of the migration operation (for example, Completed, In Progress, or Failed).
* **Number of VMs Migrated**: Shows the count of VMs that have been successfully migrated out of the total.
* **Start Time**: The timestamp when the maintenance mode operation was initiated.
* **Completion Time**: The timestamp when all migrations finished (displayed upon completion).
* **Virtual Machines Migration table**: Lists each individual VM that was migrated, along with its migration status. This allows administrators to identify any VMs that may have failed to migrate and retry migration.

## Maintenance Mode Interoperation with Other Services

This section describes how maintenance mode interoperates with other services configured for your cluster.

### Host Aggregates

Maintenance mode will honor[ Host Aggregates](/private-cloud-director/virtualized-clusters/host-aggregate) and migrate VMs to another host from the same host aggregate.

Limitations:

* All hosts from the Host Aggregate must belong to a single cluster and not span multiple clusters.
* If the host aggregate has a single host that has gone down, maintenance mode will not find a suitable target to migrate the VMs, and maintenance mode will fail.

### DRR

A host in maintenance mode will have scheduling VMs disabled while maintenance mode is active. Consequently, DRR will exclude any hosts that are in maintenance mode in a given cluster, and only balance workloads on the remaining active hosts.

### VMs with Hard Affinity or Anti-Affinity Rules

* **Hard affinity**: Maintenance mode will fail if one or more VMs with hard affinity rules are in a powered-on state.
* **Soft affinity**: Maintenance mode will attempt to keep all VMs in the affinity group together by migrating the first VM to a randomly selected host in the cluster and then migrating the subsequent VMs to the same host until it reaches capacity.
* **Hard anti-affinity**: Maintenance mode will honor the anti-affinity policy by placing the VMs in the anti-affinity group on separate hosts. If no appropriate target host is found to migrate a VM, maintenance mode will fail.
* **Soft anti-affinity**: Maintenance mode will attempt to find target hosts that satisfy the anti-affinity policy for each VM to be migrated in a powered-on state, ensuring that all VMs part of the soft anti-affinity policy are placed on separate hosts. If maintenance mode cannot find a suitable target host, maintenance mode will migrate the VM(s) to a target host that violates the soft affinity policy, and the policy violation will be displayed on the UI.

### VM States

Maintenance mode will migrate VMs that are in an active state (powered on).

Maintenance mode will not migrate VMs with the following statuses:

* Shutdown (powered off)
* Suspended
* Error
* Unknown

Maintenance mode will not proceed if a VM exists in the following states:

* **Pending resize confirmation**: The resize operation needs to be confirmed for the VM to be eligible for migration by maintenance mode.
* **Rescued**: The VM must be un-rescued for the VM to be eligible for migration by maintenance mode.

### VMs with Special Properties

1. **Virtual TPM-enabled VMs:** Maintenance mode does not have the ability to live migrate VMs with[ Virtual TPM](/private-cloud-director/virtualized-clusters/virtual-tpm) enabled today. This ability is coming soon.
2. **Zero flavor VMs:** Maintenance mode will migrate VMs created with zero flavor in a powered-on state.
3. **VMs with hot-added CPU or memory:** Maintenance mode will migrate VMs that have hot-added CPU or memory resources.

## Troubleshooting

If VMs fail to migrate, or are left stranded or in an error state during maintenance mode, see [Troubleshoot Maintenance Mode Migration Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshoot-maintenance-mode-migrations).


# CPU Model

## What is a CPU Model

In <code class="expression">space.vars.product\_name</code>, a CPU model refers to the set of CPU features and flags exposed to a virtual machine. This is how <code class="expression">space.vars.product\_name</code> compute service manages the CPU architecture presented to a virtual machine guest operating system, allowing for flexibility and control over performance, compatibility, and migration capabilities.

## Use Cases for CPU Model

* **Performance:** Exposing specific CPU features can maximize the performance of workloads running inside the virtual machines.
* **Compatibility:** Ensuring consistent behavior across different hypervisor hosts in a virtualized cluster, regardless of the underlying host CPU, can be achieved by specifying a CPU model.
* **Migration:** CPU models play an important role in [Live Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration) and [Dynamic Resource Rebalancing Drr ](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr). Matching CPU models on source and destination hosts is crucial for successful live migration and for [Dynamic Resource Rebalancing Drr ](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr)to operate with full efficiency.

## Locating CPU Models Supported By Your Host

The libvirt KVM driver that comes built in as part of your Ubuntu installation on your hypervisor host provides a number of standard CPU model names.

These models are defined in `/usr/share/libvirt/cpu_map/*.xml`. You can inspect these files to determine which models are supported by your server.

Each file located under this directory contains information about the **feature set provided by the CPU model.**

For example, here the file `x86_SandyBridge-IBRS.xml`describes the CPU vendor, family, model and features supported by the SandyBridge-IBRS CPU Model.

{% tabs %}
{% tab title="Bash" %}

```bash
$ cat /usr/share/libvirt/cpu_map/x86_SandyBridge-IBRS.xml
<cpus>
  <model name='SandyBridge-IBRS'>
    <decode host='on' guest='on'/>
    <signature family='6' model='42'/> <!-- 0206a0 -->
    <signature family='6' model='45'/> <!-- 0206d0 -->
    <vendor name='Intel'/>
    <feature name='aes'/>
    <feature name='apic'/>
    ...
  </model>
</cpus>
```

{% endtab %}
{% endtabs %}

You can also locate the supported CPU models by running `virsh cpu-models <ARCH>` command:

{% tabs %}
{% tab title="Bash" %}

```bash
$ virsh cpu-models x86_64
...
SandyBridge
SandyBridge-IBRS
IvyBridge
IvyBridge-IBRS
Haswell-noTSX
Haswell-noTSX-IBRS
Haswell
Haswell-IBRS
```

{% endtab %}
{% endtabs %}

## Supported CPU Models

<code class="expression">space.vars.product\_name</code> currently supports following list of [Supported Cpu Models List](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model/supported-cpu-models-list).

## CPU Model Configuration Per Host

When you configure a host with 'hypervisor' role, here are the steps followed by the compute service to configure the CPU model on the host:

1. It starts by checking the [list of approved models](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model/supported-cpu-models-list), beginning with the newest models based on the CPU vendor for your host.
2. For each model, it verifies compatibility by checking if the model is marked as **usable** in the `virsh -c qemu:///system domcapabilities` output.
3. The first approved model marked as usable is selected.
4. If no approved models are usable, the system falls back to a generic model.

## CPU Model Pre-requisites for Hypervisor Hosts

* Each host that you authorize in your <code class="expression">space.vars.product\_name</code> setup with 'hypervisor' role must support at least one CPU Model that is part of the [Supported Cpu Models List](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model/supported-cpu-models-list) list.
  * Failing this requirement may result in unexpected errors such as the host going to offline state, or virtual machines failing to provision, etc.

## CPU Model Pre-requisites for Migration

Following CPU model prerequisites must be met by all hosts within a virtualized cluster, so that operations like live migration, cold migration, VM evacuation can work smoothly, and services like [Dynamic Resource Rebalancing Drr ](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr), [Virtual Machine High Availability Vm Ha ](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha)can operate efficiently.

* All hosts in the cluster must have the exact same value for the most recent CPU model supported. Read here for how [system configures CPU model per host](#cpu-model-configuration-per-host).
* This value must be part of the list of [Supported Cpu Models List](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model/supported-cpu-models-list).
* If the hosts do not all report the same most-recent supported CPU model, or if a host’s CPU model is missing from the supported models list above, then you can set the [CPU Mode and Model](/private-cloud-director/virtualized-clusters/virtualized-cluster#cpu-mode-and-model)

## Troubleshoot CPU Baseline Issues

If a host transitions to `error` or `offline` status after a host upgrade, or if live migrations begin failing with CPU compatibility errors, the cluster CPU baseline may need to be adjusted. See [Resolve CPU Baseline Mismatch After Host Upgrade](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/cpu-baseline-mismatch) for a step-by-step diagnostic and remediation guide, including how to use `cpu_model_extra_flags` for mixed-generation clusters.


# Supported CPU Models

This document provides a list of currently supported CPU Models in <code class="expression">space.vars.product\_name</code>.

## Supported Intel Models

## 2025+ Generation

```
        ClearwaterForest

        # 2024+ Generation

        SierraForest SierraForest-v1

        GraniteRapids GraniteRapids-v2 GraniteRapids-v1

        # 2021-2023 Generation

        SapphireRapids SapphireRapids-v3 SapphireRapids-v2 SapphireRapids-v1

        # 2021 Generation

        Snowridge Snowridge-v4 Snowridge-v3 Snowridge-v2 Snowridge-v1

        # 2020 Generation

        Cooperlake Cooperlake-v2 Cooperlake-v1

        # 2019-2020 Generation

        Icelake-Server Icelake-Server-noTSX Icelake-Server-v7 Icelake-Server-v6

        Icelake-Server-v5 Icelake-Server-v4 Icelake-Server-v3 Icelake-Server-v2 Icelake-Server-v1

        Icelake-Client Icelake-Client-noTSX

        # 2017-2019 Generation

        Cascadelake-Server Cascadelake-Server-noTSX Cascadelake-Server-v5

        Cascadelake-Server-v4 Cascadelake-Server-v3 Cascadelake-Server-v2 Cascadelake-Server-v1

        # 2015-2017 Generation

        Skylake-Server Skylake-Server-IBRS Skylake-Server-noTSX-IBRS

        Skylake-Server-v5 Skylake-Server-v4 Skylake-Server-v3 Skylake-Server-v2 Skylake-Server-v1

        Skylake-Client Skylake-Client-IBRS Skylake-Client-noTSX-IBRS

        Skylake-Client-v4 Skylake-Client-v3 Skylake-Client-v2 Skylake-Client-v1

        # 2014-2015 Generation

        Broadwell Broadwell-IBRS Broadwell-noTSX Broadwell-noTSX-IBRS

        Broadwell-v4 Broadwell-v3 Broadwell-v2 Broadwell-v1

        # 2013-2014 Generation

        Haswell Haswell-IBRS Haswell-noTSX Haswell-noTSX-IBRS

        Haswell-v4 Haswell-v3 Haswell-v2 Haswell-v1

        # 2012-2013 Generation

        IvyBridge IvyBridge-IBRS IvyBridge-v2 IvyBridge-v1

        # 2011-2012 Generation

        SandyBridge SandyBridge-IBRS SandyBridge-v2 SandyBridge-v1

        # 2010-2011 Generation

        Westmere Westmere-IBRS Westmere-v2 Westmere-v1

        # 2008-2010 Generation

        Nehalem Nehalem-IBRS Nehalem-v2 Nehalem-v1

        # 2007-2008 Generation

        Penryn Penryn-v1

        # 2006-2007 Generation

        Conroe Conroe-v1

        # Specialized processors

        Denverton Denverton-v3 Denverton-v2 Denverton-v1

        KnightsMill KnightsMill-v1
```

## Supported AMD Models

## 2022+ Generation

```
        EPYC-Genoa EPYC-Genoa-v1

        # 2021-2022 Generation

        EPYC-Milan EPYC-Milan-v2 EPYC-Milan-v1

        # 2019-2020 Generation

        EPYC-Rome EPYC-Rome-v4 EPYC-Rome-v3 EPYC-Rome-v2 EPYC-Rome-v1

        # 2017-2019 Generation

        EPYC EPYC-IBPB EPYC-v4 EPYC-v3 EPYC-v2 EPYC-v1

        # Legacy Opteron series (2003-2017)

        Opteron_G5 Opteron_G5-v1 

        Opteron_G4 Opteron_G4-v1 

        Opteron_G3 Opteron_G3-v1 

        Opteron_G2 Opteron_G2-v1 

        Opteron_G1 Opteron_G1-v1 

        # Chinese Zen-based processors

        Dhyana Dhyana-v2 Dhyana-v1

        # Legacy consumer processors

        phenom phenom-v1
```

## Supported Generic Models

## KVM virtualization models

```
        qemu64 qemu64-v1

        kvm64 kvm64-v1



        # RHEL/CentOS compatibility models

        cpu64-rhel6

        cpu64-rhel5
```


# Manually Configure CPU Model

Follow this document to manually configure CPU model for all hypervisor hosts in your virtualized cluster.

Occasionally you may need to manually configure the CPU model across hosts in your virtualized cluster, to ensure consistent behavior for important features such as [Dynamic Resource Rebalancing DRR ](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr), [Virtual Machine High Availability Vm Ha ](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha), [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) and overall predictable performance across your hypervisor hosts.

Follow the steps to do this:

## Step 1: Query the existing CPU model on a host

Run the following command to query the list of supported CPU models on a host

{% tabs %}
{% tab title="Bash" %}

```bash
$ virsh domcapabilities | grep "model usable='yes" | sort
```

{% endtab %}
{% endtabs %}

This command will show output similar to the following:

{% tabs %}
{% tab title="Bash" %}

```bash
<model usable='yes'>486</model>
      <model usable='yes'>Broadwell-noTSX-IBRS</model>
      <model usable='yes'>Broadwell-noTSX</model>
      <model usable='yes'>Cascadelake-Server-noTSX</model>
      <model usable='yes'>Conroe</model>
      <model usable='yes'>core2duo</model>
      <model usable='yes'>coreduo</model>
      <model usable='yes'>Haswell-noTSX-IBRS</model>
      <model usable='yes'>Haswell-noTSX</model>
      <model usable='yes'>IvyBridge-IBRS</model>
      <model usable='yes'>IvyBridge</model>
      <model usable='yes'>kvm32</model>
      <model usable='yes'>kvm64</model>
      <model usable='yes'>n270</model>
      <model usable='yes'>Nehalem-IBRS</model>
      <model usable='yes'>Nehalem</model>
      <model usable='yes'>Opteron_G1</model>
      <model usable='yes'>Penryn</model>
      <model usable='yes'>pentium2</model>
      <model usable='yes'>pentium3</model>
      <model usable='yes'>pentium</model>
      <model usable='yes'>qemu32</model>
      <model usable='yes'>SandyBridge-IBRS</model>
      <model usable='yes'>SandyBridge</model>
      <model usable='yes'>Skylake-Client-noTSX-IBRS</model>
      <model usable='yes'>Skylake-Server-noTSX-IBRS</model>
      <model usable='yes'>Westmere-IBRS</model>
      <model usable='yes'>Westmere</model>
```

{% endtab %}
{% endtabs %}

## Step 2 - Choose a Common CPU Model

Run the above command on all hypervisor hosts in your virtualized cluster to find the list of supported models. Then choose the most recent CPU Model that is supported across all hosts.

## Step 3 - Apply CPU Model Changes to All Hosts

Now that you have identified the common model, make changes to the compute service config file to apply the CPU model to all hypervisor hosts in your virtualized cluster.

Follow [CPU Mode and Model Configuration](/private-cloud-director/virtualized-clusters/nova-override#cpu-mode-and-model-configuration) and edit the CPU model value with the model you identified in step 2.

Make sure to [Restart Compute Service](/private-cloud-director/virtualized-clusters/nova-override#restart-service) on each host following the changes.


# Advanced Remote Support

Advanced Remote Support (ARS) is a secure troubleshooting mechanism that allows Platform9 support engineers to log onto your <code class="expression">space.vars.product\_name</code> hosts to analyze and resolve complex technical issues. By default, Platform9 support team members cannot interactively access your hosts. However, when you enable ARS, support engineers can securely connect through the host's existing connection to the Platform9 management plane.

Despite being based on SSH, ARS does not expose your host to SSH login from any external network. The mechanism leverages secure channels and does not require any firewall changes to your host or network infrastructure. In this guide, you will learn how to enable, configure, and disable Advanced Remote Support for your <code class="expression">space.vars.product\_acronym</code> hosts.

## Prerequisites

Before you enable Advanced Remote Support, ensure you have:

* Administrative access to the <code class="expression">space.vars.product\_acronym</code> UI
* Root or sudo access to the target host
* The SSH daemon (`sshd`) running on the target host
* The `pf9` user account created on the host (this is created automatically during <code class="expression">space.vars.product\_acronym</code> installation)

## Enable Advanced Remote Support

To enable Advanced Remote Support for a host:

1. Navigate to **Infrastructure → Cluster Hosts** in the <code class="expression">space.vars.product\_acronym</code> UI.
2. Select the checkbox next to the host where you want to enable remote support.
3. Click **Edit Roles**.
4. Select the **Advanced Remote Support** checkbox.
5. Click **Update Role Assignment**.

**Expected outcome:** The host now allows Platform9 support engineers to establish secure SSH connections through the management plane.

## Verify SSH Daemon Configuration

To ensure the SSH daemon is properly configured:

1. Connect to your host using your standard SSH method.
2. Verify that the `sshd` service is running by executing the appropriate command for your operating system.
3. Confirm that the SSH daemon configuration allows key-based authentication.

Consult your Linux operating system's documentation for specific instructions on verifying and configuring the SSH daemon.

## Configure Sudo Access (Optional, Recommended)

When ARS is enabled, Platform9 support engineers log into the host using the `pf9` user account. By default, this account has restricted privileges. To allow support engineers to run diagnostic commands with elevated privileges, you need to configure sudo access for the `pf9` user.

### Requirements for Sudo Access

To grant sudo access:

* sudo must be enabled for the 'pf9' user
* sudo must allow the 'pf9' user to authenticate without a password (ARS uses one-time SSH keys, so the 'pf9' user does not have a password by default)

### Configure Sudo on Debian and Ubuntu Systems

To configure sudo access on Debian-based systems:

1. Edit the sudo rules by running the visudo command.
2. Add or verify the following line to allow wheel group members to authenticate without a password:

{% tabs %}
{% tab title="Bash" %}

```bash
pf9 ALL=(ALL) NOPASSWD: ALL
```

{% endtab %}
{% endtabs %}

3. Save and exit the editor.

**Expected outcome:** The 'pf9' user can now run commands with sudo privileges without entering a password.

For other Linux distributions, consult your operating system's documentation for specific instructions on configuring passwordless sudo access.

## Coordinate with Platform9 Support

After enabling Advanced Remote Support, coordinate with your Platform9 support representative to arrange access:

1. Identify the specific host that requires troubleshooting by sharing:
   * The contents of the host's `/etc/pf9/host_id.conf` file, or
   * The host's hostname
2. Agree on a time window when the support engineer can access the host.
3. Confirm that ARS is enabled and properly configured on the target host.

## Disable Advanced Remote Support

When troubleshooting is complete, you should disable Advanced Remote Support to restore standard access controls.

To disable Advanced Remote Support:

1. Navigate to **Infrastructure → Cluster Hosts** in the <code class="expression">space.vars.product\_name</code> UI.
2. Select the checkbox next to the host where you want to disable remote support.
3. Click **Edit Roles**.
4. Deselect the **Advanced Remote Support** checkbox.
5. Click **Update Role Assignment**.

**Expected outcome:** Platform9 support engineers can no longer establish SSH connections to the host through the management plane.


# Host Aggregate

A host aggregate is a **group of hosts** within your virtualized cluster that share common characteristics. Host aggregates are a mechanism to create specialized groupings of hosts that have specific properties, with the purpose of enabling provisioning of VMs that require those properties on those hosts.

You can create a host aggregate by navigating to the <code class="expression">space.vars.product\_name</code> UI and selecting 'Infrastructure' -> 'Host Aggregates' from the left nav bar. When you create a new host aggregate, you get to specify a name for the aggregate and a set of key-value pairs that describe the specific properties of the hosts that will belong to this host aggregate.

For example, you may wish to create a group of hosts within your virtualized cluster that have Oracle licensing enabled on them. You can do this by creating an aggregate called "Oracle Hosts" and within that aggregate, creating a key-value pair with values "Oracle-License" and "True". You can then add the hosts that you know have Oracle license installed on them to this aggregate.

This will enable you to provision VMs that require Oracle licensing to place on these hosts. You will do this by first creating a flavor that references this host aggregate and then using that flavor to create your VM. Read about [Flavors and Host Aggregates](/private-cloud-director/virtualized-clusters/virtualmachine#flavors-and-host-aggregates) for how to do this.


# Virtual Machines

A virtual machine, or a VM, is a software-based representation of a physical computer. A VM allows you to run an operating system and applications using the resources of a host machine or a hypervisor, acting like a separate, isolated computer with its own virtual CPU, memory, storage, and network capabilities, all managed by the hypervisor.

You can create a virtual machine by first [populating an image in the image library](/private-cloud-director/images-and-image-library/image-library---images), then creating one or more VM flavors, and then creating a new VM.

## VM Flavors

Once you have the images you would like to use, it's time to look at the resource configuration for the VMs to be deployed. <code class="expression">space.vars.product\_name</code> uses T-shirt-sized configurations of resource allocations, called Flavors, to allow you to specify the resource allocation for VMs. Read more about [VM Flavors](/private-cloud-director/virtualized-clusters/virtualmachine/vm-flavors) here.

## VM Affinity and Anti-affinity Rules <a href="#server-groups" id="server-groups"></a>

<code class="expression">space.vars.product\_name</code> supports the creation of VM affinity and anti-affinity groups using a mechanism called Server Groups. Read more about [VM Affinity Anti-Affinity Rules](/private-cloud-director/virtualized-clusters/virtualmachine/vm-affinity-anti-affinity-rules) here.

## Create a VM

By now, you have your desired VM image, your preferred Flavor, and you have previously setup at least one Network that you can use to deploy your VM. So navigate to Virtual Machines in the Navigation pane and start 'Deploy Virtual Machine'.

We will describe some of the VM creation options here.

### VM source

{% hint style="warning" %}
**Warning**

This is an important selection that affects where your virtual machine disk is stored, and whether the VM will be using [Ephemeral Storage](/private-cloud-director/storage/ephemeral-storage) or [Block Storage](/private-cloud-director/storage/block-storage)
{% endhint %}

#### Boot VM from Image

Use this option to provision a VM that uses [Ephemeral Storage](/private-cloud-director/storage/ephemeral-storage) for its root disk. This storage option is typically used for non-production VMs where the disk need not be preserved across host failures or VM deletion. The VM root disk will be created using local storage on the hypervisor on which it is provisioned and populated with the contents of the source image. If the hypervisor were to go down, it would not be possible to access or recover the VM disk till the hypervisor comes back online. If the host's local disk experiences corruption or other issues, the VM may be impacted and may be unrecoverable. If the VM were to be deleted, the VM's root disk would be deleted along with it and would not be recoverable afterward.

#### Boot VM from New Volume

Use this option if you would like to create a VM that uses [Block Storage](/private-cloud-director/storage/block-storage) volume for its root disk. A VM created in this manner will persist across host failures and VM termination, unless you choose to delete the volume when the VM terminates. When you choose this option, a new persistent volume is created on the block storage you configured for your virtualized cluster, and its contents are populated from the source image.

#### Boot VM from Existing Volume

This option is like the previous one, except that you choose to boot from an existing volume rather than provisioning a new volume from a base image.

#### Install OS from ISO Wizard

You can now create Windows and Linux VMs directly from installation ISOs through the **Deploy New VM** wizard in the UI. This replaces the previous CLI-only workflow documented in the [Windows](/private-cloud-director/tutorials/create-windows-vm-from-iso/creating-a-windows-virtual-machine--vm--from-an-iso-image4io) and [Ubuntu](/private-cloud-director/tutorials/deploy-a-virtual-machine-using-an-ubuntu-iso-image) tutorials.

**To install an OS from an ISO:**

1. In the **Deploy New VM** wizard, select **Install from ISO** from the **Boot VM From** dropdown.
2. Select an installation ISO from your image library and set the **Root Volume Size**. A new volume of the specified size will be created for the OS installation.
3. (Optional) For Windows installs, enable **Drivers ISO** and select the VirtIO drivers ISO for optimal performance.
4. Complete the remaining steps (optionally attach additional volumes, select a flavor, configure networking, customize VM) and click **Deploy Virtual Machine**.
5. Connect to the VM console and complete the OS installation. The VM is configured with a boot order of boot volume first, then ISO, so once the OS installer finishes and the VM reboots, it will automatically boot from the installed disk.

**Notes:**

* Only zero-disk flavors are supported, since the root disk is provided by the volume created in step 2. The VM disk size will equal the volume size.
* Bulk VM creation is not supported when installing from an ISO.
* The installation ISO can remain attached to the VM after installation. It will not interfere with boot because of the boot order, and it remains in the image library. You can detach it manually if you prefer.

#### Guest OS hostname

By default, the value you enter in the **Name** field is used both as the VM's display name in Private Cloud Director and as the guest OS hostname. To set a hostname that differs from the VM name, use the **Guest OS Hostname (Optional)** field in the **Basic VM Info** section. This is useful when you want a descriptive VM name for tracking and reporting, paired with a shorter, standards-compliant hostname inside the OS.

If you leave the **Guest OS Hostname (Optional)** empty, the **VM** **Name** is used as the guest OS hostname.

When you deploy more than one VM in a single operation, Private Cloud Director appends a numeric suffix to the hostname so that each VM has a unique name, using the same pattern as VM names.

If you set a value in **Guest OS Hostname (Optional)** and also include a cloud-init script that sets the hostname, the cloud-init script takes precedence. Use only one method to set the hostname to avoid unexpected behavior.

### Choose a Specific Subnet

When the network you select has more than one subnet, the **Networking** step of the **Deploy New VM** wizard displays a **Subnet** column for each network row. The selector defaults to **Automatic**; expanding it lists each available subnet as `<CIDR> (<subnet-name>)` — for example, `192.0.0.0/5 (test)`. Select one to place the VM's network interface on that specific subnet. If you leave it on **Automatic**, the Networking Service chooses the subnet for you.

Choosing a specific subnet does not limit how many VMs you can create in a single operation. If, however, you switch a network's IP mode from **Automatic** to **Private IP** and enter a specific address, the wizard's **Customize VM** step restricts you to a single VM and shows:

> You can create only 1 VM when a specific port is selected from a network.

For dual-stack or other advanced port configurations, pre-create a port on the subnet you want and attach the port during the **Networking** step. See [Virtual Network](/private-cloud-director/virtualized-networking/networks-and-ports) and [Add Multiple Dual Stack Ports to a VM](/private-cloud-director/tutorials/create-dual-stack-port).

### SSH Key

Using SSH Keys is the recommended and secure method to access your VMs. To set up your SSH Key, navigate to 'Networks and Security' in the Navigation Pane, and import your SSH Keys.

Once you have the key imported, you can select that key when creating a new VM.

### Select Server Group

Optionally select an [affinity or anti-affinity group](#vm-affinity-and-anti-affinity-rules) for this VM to be part of. This will impact the VM's placement on a host in your virtualized cluster.

### Using cloud-init

Using cloud-init allows you to customize the virtual machine as it boots, such as provisioning additional software or custom scripts.

### Assign Security Groups

Security groups allow you to limit port access to Virtual Machines. If you don't already have a Security Group configured, you can cover this later.

### Specify Metadata

Metadata attributes allow you to specify additional key-value pairs, which can be used to query and organize your virtual machines.

{% hint style="success" %}
**Success**

With this information specified, you should see the VM provision within a few minutes. You can then access the VM over the network or via its console.
{% endhint %}

## VM Properties

### VM UUID

Each newly created virtual machine is assigned a unique UUID. By default, this field is not displayed in the virtual machines grid view in the UI. You can change this by clicking the 'Manage Columns' button above the virtual machines grid view and selecting the UUID field. You can also select an individual VM to open the VM details view and see the ID property listed there. Alternatively, you can get VM ID using the pcdctl CLI by running the `pcdctl server list` command or the `pcdctl server show` command and supplying the name of the VM.

## Migrate Existing VMs from VMware

If you have existing workloads that you would like to migrate onto your <code class="expression">space.vars.product\_name</code> cluster, [Project vJailbreak](https://github.com/platform9/vjailbreak) can help.

### Download the vJailbreak Appliance

Head over to Project vJailbreak to download the appliance image. The appliance is packaged as an OVA that you can deploy into your <code class="expression">space.vars.product\_name</code> environment. Note that vJailbreak should have access to your storage network to efficiently migrate VM data over during the data copy phase.

### Register your Source VMware Environment and Your New Virtualized Cluster

As your vJailbreak appliance loads, you'll see instructions for accessing its user interface, where you can specify your source VMware environment information, as well as the information for your new cluster.

### Select VMs and Schedule Migrations

With this information specified, you can now initiate migrations from the vJailbreak user interface.


# VM Flavors

<code class="expression">space.vars.product\_name</code> uses t-shirt sized configurations of resource allocations, called flavors, to allow you to specify the resource allocation for VMs. This feature enables administrators to templatize and standardize on a few different VM configurations that developers can use, rather than having a configuration sprawl.

<code class="expression">space.vars.product\_name</code> ships with a set of flavors out of box, but administrators can create new flavors and make them available to end users.

## Flavor Properties

Flavors capture the following configuration information for a VM:

1. vCPUs - number of CPU cores
2. RAM - amount of memory
3. Disk - size of the VM ephemeral disk, **only used when** creating a VM from an image or a [VM Ephemeral Disk Snapshot](/private-cloud-director/virtualized-clusters/virtualmachine/virtual-machine-snapshot#type-1---vm-ephemeral-disk-snapshot). Ignored when creating a VM using a volume for root disk.
4. Metadata - used to define:
   1. What [Host Aggregate](/private-cloud-director/virtualized-clusters/host-aggregate) the VM should get placed on, if any, by specifying metadata (key-value pairs) that correspond to a specific [Host Aggregate](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster#host-aggregates). When creating a VM using this flavor, the VM will be placed on hosts in your virtualized cluster that are part the selected host aggregate (that match the specified metadata attributes).

## Create Flavors

Navigate to the Flavors menu in the UI under Virtual Machines, and create or edit flavors to suit your desired CPU, memory and disk configuration.

<figure><img src="/files/AmLThjZWSmlIZolgv1iR" alt=""><figcaption></figcaption></figure>

## Zero CPU or Memory Flavors

<code class="expression">space.vars.product\_name</code> allows you to create flavorless VMs where you specify CPU and memory information dynamically at VM creation time, instead of having to use a pre-define value via a flavor. This is done by using special flavors that have both CPU and memory set to zero. Read more about creating [flavorless VMs here](/private-cloud-director/virtualized-clusters/virtualmachine/flavorless-vms-with-hot-plug).

## Zero Disk Flavors

You can create a VM flavor with disk size set to zero.

Using such a flavor is a requirement when creating a VM that will use a block storage volume for it's root disk.

When used to create a VM that will use ephemeral root disk, the root disk size for the VM will be set to be equal to the size of the image.


# VM Affinity Anti-Affinity Rules

This document describes how to create affinity and anti-affinity rules for your virtual machines, the placement behavior each rule produces, and how those rules interact with other <code class="expression">space.vars.product\_name</code> operations such as Dynamic Resource Rebalancing (DRR), VM High Availability (VM HA), maintenance mode, and migration.

## Overview

Server groups are the mechanism <code class="expression">space.vars.product\_name</code> uses to apply affinity and anti-affinity rules. A server group is a logical group of virtual machines that shares a single placement policy. Membership in a server group controls how the member VMs are placed across the hypervisor hosts in your virtualized cluster.

<code class="expression">space.vars.product\_name</code> supports four policies. Two are strict ("hard") and two are best effort ("soft").

| Policy             | Rule strength | Placement goal                                 | If the goal cannot be met            |
| ------------------ | ------------- | ---------------------------------------------- | ------------------------------------ |
| Affinity           | Strict (hard) | All VMs in the group run on the same host      | The operation fails                  |
| Soft affinity      | Best effort   | Keep all VMs in the group on the same host     | VMs may be placed on different hosts |
| Anti-affinity      | Strict (hard) | No two VMs in the group run on the same host   | The operation fails                  |
| Soft anti-affinity | Best effort   | Spread VMs in the group across different hosts | Two or more VMs may share a host     |

* **Affinity (hard).** Restricts all VMs in the group to the same hypervisor host. If no single host can hold every member, the operation fails.
* **Soft affinity.** Attempts to keep all VMs in the group on the same host, but allows them to run on different hosts when that is not possible.
* **Anti-affinity (hard).** Restricts the VMs in the group to separate hosts, so that no two members run on the same host. If there are not enough eligible hosts, the operation fails.
* **Soft anti-affinity.** Attempts to keep the VMs in the group on separate hosts, but allows two or more of them to share a host when that is not possible.

A server group has exactly one policy. You cannot combine affinity and anti-affinity, or hard and soft, within the same group.

## Prerequisites

* A virtualized cluster with one or more hosts that have the hypervisor role.
* Enough host capacity for the policy you choose. A hard affinity group needs a single host with enough capacity for every member. A hard anti-affinity group needs at least as many eligible hosts as it has VMs. Anti-affinity and soft anti-affinity are only meaningful when the cluster has more than one hypervisor host.
* The server group must exist before you create the VMs that will belong to it.
* The appropriate role. Administrators and self-service users can create and delete server groups. Read-only users cannot.

## Create a Server Group

To create a server group from the UI:

1. Navigate to **Virtual Machines** > **Server Groups** and click the **Create Server Group** button.
2. In the **Name** field, enter a descriptive name for the server group.
3. From the **Policy** dropdown, select one of **Affinity**, **Soft Affinity**, **Anti-Affinity**, or **Soft Anti-Affinity**. See [Overview](#overview) for what each policy does.
4. Click the **Create Server Group** button.

You can now add a VM to this server group as part of the VM creation wizard.

## Add a VM to a Server Group

A VM's server group is selected when the VM is created, in the VM creation wizard.

You cannot add an existing VM to a server group, and you cannot move a VM from one server group to another, after the VM has been created. To change membership, recreate the VM in the desired group. You can create the new VM from an image of the original VM to preserve its disk contents.

## How Affinity Rules Interact With <code class="expression">space.vars.product\_name</code> Operations

After a server group is created and its VMs are running, the group's policy continues to govern placement during automated and administrative operations. The table below summarizes the behavior for powered-on VMs.

| Policy                   | DRR                                                                                                                                                                                     | VM HA                                                                                                                                                                                                                                                                                       | Maintenance mode                                                                                                                                                                                     | Live and cold migration                                                                                                                                                                   |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Affinity** (hard)      | DRR does not rebalance members of a hard affinity group. The Compute Service cannot find a valid target for a single member without breaking the rule, so the VM is skipped.            | On host failure, VM HA looks for a single host with enough capacity for the entire group and restarts all members there together. If no such host exists, HA fails for the group and the UI reports the failure.                                                                            | Entering maintenance mode fails if any powered-on member of a hard affinity group is on the host. The UI reports the policy violation.                                                               | Migration fails. The platform cannot move every member together in a single operation, so a single member cannot be relocated without breaking the rule. See [Limitations](#limitations). |
| **Soft affinity**        | DRR may rebalance members. A member with no explicit migration priority is treated as low priority for rebalancing; a member with a set migration priority is honored at that priority. | VM HA tries to restart all members together on a single host with enough capacity. If none is available, it restarts the members individually, which may place them on different hosts, and the UI reports an affinity policy violation.                                                    | Best effort. <code class="expression">space.vars.product\_name</code> keeps members together where capacity allows and places the remainder on other hosts. The group may end up split across hosts. | Allowed. The UI warns that the move may violate the soft affinity policy but does not block it.                                                                                           |
| **Anti-affinity** (hard) | When choosing a migration target, DRR excludes any host that already runs another member of the group.                                                                                  | On host failure, VM HA restarts each VM on a host that is not running another member VM. If no compliant host is available, HA fails for that VM and the UI reports the failure. A member on an offline host is shown as disconnected.                                                      | Entering maintenance mode fails if no target host preserves the anti-affinity rule.                                                                                                                  | Migration fails when the selected target host already runs another member of the group. The UI filters out non-compliant hosts.                                                           |
| **Soft anti-affinity**   | Same rebalancing behavior as soft affinity. Members are honored at their set migration priority, or treated as low priority when none is set.                                           | VM HA tries to restart each member on a host that is not running another member. If none is available, it may restart on a host that already runs a member and reports an anti-affinity policy violation. If it cannot restart the VM at all, HA fails and the VM is shown as disconnected. | Best effort. Members may be co-located on a host when capacity requires it.                                                                                                                          | Allowed even when the target host runs another member. The UI warns of the membership but does not block the move.                                                                        |

{% hint style="info" %}
**Hot add of CPU or memory does not affect affinity rules.** A hot add operation does not move a VM to a different host, so a hot-added VM keeps the same placement and never violates an affinity or anti-affinity rule, regardless of the policy type.
{% endhint %}

## Limitations

* <code class="expression">space.vars.product\_name</code> cannot move all members of a server group together in a single, atomic operation. As a result:
  * Live or cold migration of a VM in a **hard affinity** group fails. Moving one member VM on its own would break the rule that all members share a host, and the platform cannot relocate the entire group in one step.
  * Live or cold migration of a VM in a **hard anti-affinity** group fails when the selected target host already runs another member of the group.
  * Soft affinity and soft anti-affinity groups are not affected, because their rules are best effort and the platform is allowed to break them when necessary.
* **Group membership is set at VM creation only.** You select a VM's server group when you create the VM. You cannot add an existing VM to a group, or move a VM between groups, afterward.
* **One policy per group.** A server group enforces a single policy. Affinity and anti-affinity, or hard and soft, cannot be mixed in the same group.
* **Capacity and host count govern strict policies.** A hard affinity group needs one host with enough capacity for every member, and a hard anti-affinity group needs at least as many eligible hosts as it has members. When these conditions are not met, VM creation, an HA restart, or a maintenance-mode evacuation can fail because no valid host is available.

## VMware Equivalents

If you are coming from VMware, <code class="expression">space.vars.product\_name</code> affinity rules map to vSphere DRS affinity rules as follows.

<table data-header-hidden="false" data-header-sticky><thead><tr><th>Private Cloud Director policy</th><th>Closest VMware DRS rule</th></tr></thead><tbody><tr><td>Affinity</td><td>"Keep Virtual Machines Together", or a VM-Host "Must run on hosts in group" rule</td></tr><tr><td>Soft affinity</td><td>A preferential "should run" affinity rule</td></tr><tr><td>Anti-affinity</td><td>"Separate Virtual Machines", or a VM-Host "Must not run on hosts in group" rule</td></tr><tr><td>Soft anti-affinity</td><td>A preferential "should not run" separation rule</td></tr></tbody></table>

As in vSphere, a hard rule blocks an HA restart that would violate it, while a soft rule lets HA restore availability first, even if that temporarily breaks the rule, and rebalances toward compliance afterward.


# Virtual Machine Actions

### Rebuild VM

The rebuild operation allows for the re-creation of virtual machine from a new or existing image while preserving the VM's:

* UUID
* Storage Volumes
* IP addresses (private or public)
* Network ports

If the VM utilizes ephemeral storage (local disks), rebuilding will result in the loss of data on those disks. However, if the VM uses block storage volumes or shared storage for its root disk and other data, those volumes are reattached to the rebuilt virtual machine, and data should be preserved.

### Rescue / Unrescue VM

Rescue mode is a powerful operation that allows you to access and repair a non-bootable virtual machine. When activated, this action boots your virtual machine into a temporary environment using the image that was used to create the vm, with full root access to the file system. This mode is useful for:

* Troubleshooting and fixing configuration file issues
* Recovering or copying data to a remote location
* Gaining emergency access similar to single-user mode or safe mode with networking

You can rescue a virtual machine using a new or different image from the one used to create it. When you select this option in the UI, the UI shows a list of images from the image catalog, so you can choose a different image to boot from.

To enable an image to be used for rescuing VMs, add the following metadata keys to the image:

* **`hw_rescue_device`** – The type of device to attach the rescue image as (`cdrom`, `disk`, or `floppy`)
* **`hw_rescue_bus`** – The bus the rescue device should use (`scsi`, `virtio`, `ide`, or `usb`)

For a typical VM qcow2 image to be used as rescue image, the values you would use would be `hw_rescue_device: disk` and `hw_rescue_bus: scsi.`

Once your maintenance or recovery tasks are complete, you can return the VM to normal operation by unrescuing the VM. To do that using the UI, select the VM in the VM grid view, then choose the `unrescue` action from the power actions drop down menu.

### Stop

Performs a graceful shutdown of the virtual machine by sending an ACPI shutdown signal to the guest operating system. The VM will attempt to shut down cleanly, allowing running processes to terminate properly and data to be flushed to disk. If the guest OS does not respond to the ACPI signal, the VM may not shut down. In such cases, use Hard Reboot instead.

### Reboot

Performs a graceful restart of the virtual machine by sending an ACPI reboot signal to the guest operating system. This allows the OS to restart cleanly, preserving data integrity.

### Hard Reboot

Forces an immediate restart of the virtual machine without waiting for the guest OS to shut down cleanly. This is equivalent to pressing the physical reset button on a server and immediately powers the VM back on. This action may result in data loss or corruption as running applications do not have time to shut down properly. Use only when a normal reboot fails or the VM is unresponsive.

### Suspend

Saves the current running state of the virtual machine to disk, including memory contents and CPU state, then stops the VM. When resumed, the VM will continue exactly where it left off, with all applications and data in the same state.

### Pause

Freezes the virtual machine execution by halting the virtual CPU without writing the state to disk. The VM's memory remains allocated in RAM, but no CPU cycles are consumed. This is a temporary state intended for short pauses.

### Rename

Changes the display name of the virtual machine. This updates only the VM's label in the <code class="expression">space.vars.product\_name</code> interface and does not affect the guest OS hostname or network identity.

### Change Owner

Transfers ownership of the virtual machine to a different user account. The new owner will have full control over the VM and its associated resources. Please ensure that the new owner has available quota to take ownership of the VM before transferring ownership.

### Edit Metadata

Allows you to add, modify, or remove custom key-value metadata tags associated with the virtual machine.

### Resize

Allows you to choose a new flavor to resize the VM to. Resizing a VM results in the VM being rebuilt (powered off, then recreated using the source image on the same or different host) using the new flavor. Please note that the disk size of the new flavor must be larger than or equal to the current flavor.

### Clone

Creates a complete copy of the virtual machine, including its disks and configuration. The new VM is independent of the original and can be modified without affecting the source VM.

### Migration Priority

Sets the priority level for VM migrations triggered by Dynamic Resource Rebalancing (DRR). Higher priority migrations are processed first during automated resource rebalancing or operations. You can specify migration priority as Low, Normal, High, or Excluded. If a migration priority is not set for a VM, it defaults to Normal priority. Selecting Excluded excludes the VM from any DRR migrations.

### Mount ISO

Attaches an ISO from the image library to a VM in a powered-off state. This action presents the mounted ISO as a read-only optical drive to the guest OS.

### Unmount ISO

Detaches an ISO that was previously attached to the VM. This action requires the VM to be powered off before the ISO can be detached.


# Virtual Machine Migration

Discover VM migration in PCD, a vital process for moving virtual machines between hosts to ensure system availability and optimize resource utilization. Learn about use cases, types, and commands for

VM migration in <code class="expression">space.vars.product\_name</code> refers to the process of moving a running or stopped virtual machine (VM) from one compute host (hypervisor) to another within the same <code class="expression">space.vars.product\_name</code> environment. This is a fundamental capability for maintaining system availability, performing infrastructure maintenance, and balancing workloads in data centers.

## Common Use Cases

Following are the common scenarios in which migration is initiated in <code class="expression">space.vars.product\_name</code>.

* **Maintenance mode:** When you invoke [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) on host, VMs are live migrated to other compatible hosts in the virtualized cluster.
* **Load Balancing:** <code class="expression">space.vars.product\_name</code> [Dynamic Resource Rebalancing (DRR)](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr) uses live migration to distribute VMs more evenly across hosts in a virtualized cluster to optimize resource utilization and prevent individual hosts from becoming overloaded.
* **Host Failure (VM Evacuation):** <code class="expression">space.vars.product\_name</code> [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) service uses VM evacuation to recover VMs from a failing or failed host.

## Types of VM Migration

<code class="expression">space.vars.product\_name</code> supports following types of VM migrations:

* Cold migration
* Live migration
* VM evacuation

Let's dive into the specifics of each.

## Cold Migration

Cold migration is a method of moving a virtual machine from one host to another, where the VM is `shut down`` ``(powered off)`or `suspended` during the process. The VM disks and configuration files are transferred to the destination host as part of the cold migration process.

**Downtime:** Cold migration essentially requires VM downtime, as the VM must be powered off or suspended before it can be cold migrated.

### Cold Migration Prerequisites

* VM should be powered off or suspended.
* Cold migration **uses the management network** for copying over virtual machine disk information. Make sure that the management network interface has **sufficient network bandwidth** **and capacity** to support **disk copy operation**.
* All hosts within the virtualized cluster must meet [CPU Model Pre-requisites](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model#cpu-model-pre-requisites).
* Source and destination hosts must run compatible operating systems. Cold migration is supported between hosts running the same Ubuntu version, **and from Ubuntu hosts to Platform9 OS powered by Rocky Linux by CIQ hosts** (the supported path for transitioning to Platform9 OS). For more information, refer to the [Operating System Compatibility Requirements](#operating-system-compatibility-requirements).

#### Using CLI

1. Authenticate [pcdctl CLI](/private-cloud-director/reference/pcdctl-command-line)
2. Run the command below to perform the cold migration. The command below will cold migrate the VM to a suitable host within the virtualized cluster.

{% tabs %}
{% tab title="Bash" %}

```bash
$ pcdctl server migrate <VM_UUID>
```

{% endtab %}
{% endtabs %}

3. To migrate VM to a specific host use the below command:

{% tabs %}
{% tab title="BASH" %}

```bash
$ pcdctl server migrate --host <HOST_UUID> <VM_UUID>
```

{% endtab %}
{% endtabs %}

#### Using UI

* Log in to the UI and select Virtual Machine menu on the left side nav bar.
* Select the desired VM then click on the "other" option on the action bar and select the `migrate` action.
* Select the target hypervisor host and hit the "Migrate VM" button.

## Live Migration

Live migration is a method of moving a virtual machine from one host to another, where the VM remains **running** throughout the migration process.

The live migration process copies a virtual machine's memory from the source to the destination host while the VM is running. Any memory pages that get modified or "dirtied" on the VM during this time are copied over again. Finally, the VM enters a brief pause period during which its remaining memory and CPU state are copied over to the destination host, and finally, the VM is resumed on the destination host.

**Downtime:** The VM will experience a brief downtime (typically milliseconds to seconds) during the pause period before resuming on the destination host.

Live migration can be classified further by the way it treats virtual machine storage:

* **Live Migration for a VM using Ephemeral Shared Storage:** The virtual machine has an ephemeral root disk that is located on [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage).
  * Live migration enables the testing and validation of whether the source and target hosts are utilizing the same underlying shared storage.
  * When validated, live migration is performed without copying over VM disk.
* **Live Migration of a VM using Ephemeral Local Storage:** The Virtual machine has an ephemeral root disk that is located on [Ephemeral Local Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-local-storage).
  * Live migration will copy over the entire virtual machine root disk from the source host to the destination host.
* **Live Migration of a VM using Volumes only:** The Virtual machine is using [block storage volumes](/private-cloud-director/storage/volume), rather than ephemeral disk.
  * Live migration will migrate the VM without performing a storage copy. The same volumes will continue to be attached to the VM once it's migrated to the target host.

### Live Migration Prerequisites

* All hosts within the virtualized cluster must run the same major and minor version of Ubuntu.
* Live migration uses the **management network to copy virtual machine memory** and, optionally, disk information by default. Make sure that the management network interface has **sufficient network bandwidth** **and capacity** to support **VM memory migration traffic**. A separate network interface can be configured for live migration traffic to isolate it from the management network.
* When using live migration in a production setup, specifically for features like [Dynamic Resource Rebalancing (DRR)](/private-cloud-director/virtualized-clusters/virtualized-cluster/dynamic-resource-rebalancing-drr), it is recommended to configure all virtual machines to use **shared storage only** for the best performance of the migration operation.
* When using [Ephemeral Local Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-local-storage), live migration requires **sufficient network bandwidth and capacity** to **support virtual machine disk copy**.
  * This is not required when using [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) at the virtualized cluster level or using [Block Storage Volumes](/private-cloud-director/storage/volume) exclusively for all virtual machines.
* All hosts within the virtualized cluster must meet [CPU Model Pre-requisites](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model#cpu-model-pre-requisites) for live migration to work. Read more about [CPU Models](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model).
* Source and destination hosts must run the same operating system version. For more information, refer to the [Operating System Compatibility Requirements](#operating-system-compatibility-requirements).

{% hint style="warning" %}
**WARNING**

Live migration between hosts with different Ubuntu versions (22.04 and 24.04) is not supported and will result in migration failures. Ensure all hosts in your virtualized cluster run the same operating system version.
{% endhint %}

#### Live Migrate Using CLI

1. Authenticate [pcdctl CLI](/private-cloud-director/reference/pcdctl-command-line)
2. Run the command below to perform the live migration.

{% hint style="info" %}
**NOTE**

Here target host UUID is mandatory.
{% endhint %}

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl server migrate --live-migration <VMUUID> --host <HOST_UUID>
```

{% endtab %}
{% endtabs %}

#### Live Migrate Using UI

1. Log in to the <code class="expression">space.vars.product\_acronym</code> UI and select **Virtual Machine**.
2. On Virtual Machine search for a specific VM and select the VM.
3. Click **other** tab and then select the **Migrate**.

This will list down the eligible target hypervisors for Live migration.

4. Select the target hypervisor and click **Migrate VM**.

## VM Evacuation

VM Evacuation is specifically designed for disaster recovery. If a host fails or becomes unresponsive, VM evacuation attempts to restart a VM that was running on the failed host onto a healthy host.

### VM Evacuation Prerequisites

* VM evacuation can only be performed for VMs using [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) or [Block Storage Volumes](/private-cloud-director/storage/volume). If the VM is using [Ephemeral Local Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-local-storage), evacuation can not work as the virtual machine disk can not be copied over from the source host to a target host.
* All hosts within the virtualized cluster must meet [CPU Model Pre-requisites](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model#cpu-model-pre-requisites) for VM evacuation to work. Read more about [CPU Models](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/cpu-model).
* Source and destination hosts must run the same operating system version. For more information, refer to the [Operating System Compatibility Requirements](#operating-system-compatibility-requirements).

{% hint style="info" %}
**NOTE**

The evacuation operation will be triggered automatically for VMs on hosts that experience an outage in clusters with VM High Availability enabled. You can perform a VM evacuation operation manually via the command line.
{% endhint %}

#### Using CLI

1. Authenticate [pcdctl CLI](/private-cloud-director/reference/pcdctl-command-line)
2. Run the command below to perform the live migration.

{% hint style="info" %}
**NOTE**

Here target host UUID is mandatory.
{% endhint %}

{% tabs %}
{% tab title="Bash" %}

```bash
$ pcdctl server evacuate --host <FAILED_HOST UUID>
```

{% endtab %}
{% endtabs %}

### Operating System Compatibility Requirements

{% hint style="info" %}
**NOTE**

Live migration and VM evacuation require source and destination hosts to run the same operating system version and KVM version. **Cold migration** additionally supports a one-way Ubuntu-to-Platform9-OS path; see [Cross-OS Cold Migration: Ubuntu Hosts to Platform9 OS Hosts](#cross-os-cold-migration-ubuntu-hosts-to-platform9-os-hosts) below.
{% endhint %}

* Hosts running Ubuntu 22.04 can only migrate VMs to other Ubuntu 22.04 hosts.
* Hosts running Ubuntu 24.04 can only migrate VMs to other Ubuntu 24.04 hosts.
* Cross-version migrations (Ubuntu 22.04 ↔ Ubuntu 24.04) are not supported.

**Known Limitation:** While initial migration from Ubuntu 22.04 to Ubuntu 24.04 may appear to succeed, subsequent migrations back to Ubuntu 22.04 may fail. This is due to incompatibilities between KVM versions on different operating systems.

#### Cross-OS Cold Migration: Ubuntu Hosts to Platform9 OS Hosts

Cold migration is supported from hosts running Ubuntu to hosts running **Platform9 OS powered by Rocky Linux by CIQ**. This is the supported path for transitioning your virtualized cluster from Ubuntu to Platform9 OS.

* Supported: Ubuntu 22.04 → Rocky Linux 10.2, Ubuntu 24.04 → Rocky Linux 10.2.
* The migration is **one-way**: cold migrating VMs back from Platform9 OS hosts to Ubuntu hosts is not supported.
* **Only cold migration is supported across operating systems.** Live migration and VM evacuation are still not supported. Power VMs off (or suspend them) before initiating the migration.
* All other Cold Migration Prerequisites still apply (CPU model compatibility, management network bandwidth, etc.).

#### Recommended Approach for OS Upgrades

When upgrading your cluster from Ubuntu 22.04 to Ubuntu 24.04 consider the following:

* Plan for one-way migration only (22.04 → 24.04).
* Complete all hosts upgrades before resuming normal migration operations.
* Do not attempt to migrate VMs back to older OS versions.

**When Migrating from Ubuntu Hosts to Platform9 OS Hosts**

When migrating from Ubuntu hosts to hosts running **Platform9 OS powered by Rocky Linux by CIQ**, consider the following:

* Cold-migrate VMs from Ubuntu source hosts to Platform9 OS destination hosts. Live migration and VM evacuation across OSes are not supported. Power VMs off (or suspend them) before initiating the migration.
* The migration is **one-way**. Do not attempt to cold-migrate VMs back to Ubuntu hosts after they have been moved to Platform9 OS hosts.
* Complete the migration of all VMs from a given Ubuntu host before decommissioning or repurposing it for Platform9 OS.

## Migration Limitations

### Cross-Cluster Live Migration

Cross-cluster live migration of virtual machines is not currently supported. This feature will be introduced with a future release of Private Cloud Director.

### Live Migration of GPU-Enabled VMs

You can live migrate virtual machines configured with [VGPUs](/private-cloud-director/gpu/gpu-support-pcd/setup-vgpu). VMs configured with GPU passthrough can not be live migrated today.

### Live Migration of Legacy Boot-from-Volume Hotplug VMs

VMs that use a hot-add-capable flavor and boot from a persistent volume, and that were created before the <code class="expression">space.vars.product\_acronym</code> 2025.10 release, may fail to live migrate until a one-time corrective step is applied. See [Resolve Live Migration Failures for Legacy Boot-from-Volume Hotplug VMs](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/hotplug-boot-from-volume-live-migration-failure).


# Flavorless VM Support

Although using flavors to create virtual machines is a fundamental feature of <code class="expression">space.vars.product\_name</code>, many organizations—especially those moving away from traditional virtualization vendors—may prefer a simpler process that doesn’t require specifying a flavor when creating a VM.

## Zero-Size Flavor

To support this scenario, Private Cloud Director allows creation of VMs using a **special zero-size flavor**, where both the **number of vCPUs and the RAM are set to zero**. Private Cloud Director ships out of the box with one such flavor, but you can also create new flavors where the number of vCPUs and the RAM are set to zero.

When you choose such a flavor when creating a new VM, you can then specify the amount of CPU and Memory that the VM should get provisioned with.

Using the zero-size flavor is a requirement in order to use CPU or memory hot-add feature

## Create a Zero-Size Flavor

1. In the Private Cloud Director UI, navigate to Virtual Machines ▷ Flavors and click Create Flavor.
2. **Enable Hot-plug**: toggle **on**.
3. Set a Disk size.
4. *(Optional)* Add Metadata if you want Private Cloud Director to place VMs on specific host aggregates.
5. Check the Make Public checkbox if you would like to make this flavor available to all tenants.
6. Click Create Flavor.

## Deploy a VM with a Zero-Size Flavor

1. Go to Virtual Machines ▷ Virtual Machines and click Deploy New VM.
2. During Flavor selection: enable the **Hot-add Compatible** toggle to filter the list.
3. Choose the zero-size flavor you created above.
4. In the Hotplug RAM (MB) and Hotplug vCPUs fields, enter the initial resources you would like to create this VM with. Maximum values are automatically configured by the system.
5. Finish the remaining steps and deploy the VM.

Private Cloud Director will create the VM with the requested CPU and Memory resources.

## Hot-add CPU or Memory to a Running VM

Read [VM Hot-Add CPU Or Memory](https://platform9.com/docs/private-cloud-director/2025.8/private-cloud-director/cpu-memory-hot-add-support) about how to hot add resources to a running VM.

## Use Flavorless VM with CLI

To automate creation of flavorless VM using CLI, API or terraform provider, use the zero-size flavor when creating the VM and use the following values for the metadata fields.

{% tabs %}
{% tab title="Bash" %}

```bash
{
    "HOTPLUG_MEMORY": "<specify memory in MB for the VM>",
    "HOTPLUG_CPU": "<specify number of VCPUs for the VM>"
}
```

{% endtab %}
{% endtabs %}

## Limitations & Best Practices

* VM disk size **cannot** be changed via the hot-add feature.
* On Windows Server 2022 guests, the memory value reported by the guest OS after a hot-add may not reflect the VM's current allocation. See [Limitations](/private-cloud-director/virtualized-clusters/virtualmachine/vm-hot-add-cpu-or-memory#limitations).


# VM Hot Add CPU Or Memory

<code class="expression">space.vars.product\_name</code> supports hot-add of CPU and / or memory resources to a running virtual machine. Using this feature, you can increase CPU cores or memory allocated to a running VM, without having to power-cycle it. This is valuable for production VMs that need more resources to meet workload demands but that can not be powered off.

### Pre-requisites <a href="#pre-requisites" id="pre-requisites"></a>

* Read more here about [Zero-Size Flavor](/private-cloud-director/virtualized-clusters/virtualmachine/flavorless-vms-with-hot-plug#zero-size-flavor). Using **zero-size flavor** when creating your VM is a requirement for hot-add. Before you can hot-add CPU or memory resources to a VM, the VM must have been created using zero-size flavor.

### Hot-add CPU or Memory to a Running VM <a href="#hot-add-cpu-or-memory-to-a-running-vm" id="hot-add-cpu-or-memory-to-a-running-vm"></a>

1. Select a powered-on VM that was created with a zero-size flavor.
2. On the VM grid view, navigate to **▷** **Other** **Actions ▷ Hotplug**.
3. Enter the new vCPU and/or memory value for the VM. The maximum allowable values are automatically set based on the VM configuration.
4. Click Hotplug VM.

### Limitations <a href="#limitations" id="limitations"></a>

* Hot add can only be used to increase the current CPU or memory allocation of a running VM. Reducing currently allocated CPU or memory is supported only while the VM is stopped.
* You can not use the hot-add feature to resize the disk of a running VM today.
* On VMs running Windows Server 2022, the guest OS's System Information (`systeminfo`) may display the flavor's maximum configured hotplug memory value instead of the VM's current memory allocation after a memory hot-add operation.
* Hot-add-capable VMs that boot from a persistent volume and were created before the <code class="expression">space.vars.product\_acronym</code> 2025.10 release may fail to live migrate until a one-time corrective hotplug is applied. See [Resolve Live Migration Failures for Legacy Boot-from-Volume Hotplug VMs](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/hotplug-boot-from-volume-live-migration-failure).


# Virtual Machine Snapshot

A virtual machine snapshot is an image of a virtual machine's disk at a specific point in time. It's a read-only copy that can be used to create a new virtual machine or restore a VM to a previous state. Essentially, it's a backup mechanism and a templating tool rolled into one.

You can create **two types** of VM snapshots in <code class="expression">space.vars.product\_name</code>.

## Type 1 - VM Ephemeral Disk Snapshot

This type of snapshot is created when you take a snapshot of a virtual machine that is using an ephemeral disk as it's root disk. A new image gets created in the Image Library that captures the point in time state of the VM's disk, including its data and configuration, at the moment the snapshot is taken.

**Snapshot File Type** **& Location**

* Ephemeral disk snapshots are always of type `qcow2`
* They are always stored in the Image Library Service as a single `qcow2` file.

Note that if a VM is created with ephemeral root disk, and then mounts one or more block storage volumes as data volumes, taking a snapshot of such a VM will only create a snapshot of it's root disk. The data volumes will not be snapshotted or tracked as part of the VM snapshot. To fully recreate such VM using a snapshot in the future, you need to first snapshot the VM and then separately take a volume snapshot of the data volumes. Then during VM restore or recreation time, you can create a new VM first using the VM snapshot, and then mount the volumes data volumes on this VM using the volume snapshots you created above.

{% hint style="danger" %}
**Important**

Taking a snapshot of a VM with ephemeral root disk will only snapshot the root disk. Any volumes attached to the VM will NOT be included in the snapshot.
{% endhint %}

### Live Snapshots

A snapshot is considered live when taken against a running virtual machine with no downtime.

Such snapshot is simply disk-only snapshot and as such is guaranteed to be crash consistent but *not application consistent.*

When you issue the command to create a snapshot of a running VM, the <code class="expression">space.vars.product\_name</code> compute service informs the hypervisor to freeze the VM to allow the creation of a “delta” file before resuming the execution of the VM. This is done to prevent the VM from writing directly to its disk while the disk is copied. VM continues to write to the delta file while the VM disk is being copied. When the copy is done, the VM is frozen again to allow the “delta” to be merged with the VM's disk, and the execution is then resumed with the disk fully merged.

Inconsistencies can appear on the first freeze if the VM is not aware that the hypervisor is taking a snapshot, because the applications and the kernel running on the instance are not told to flush their buffers.

For applications requiring strict data consistency, it is recommended to either shut down the VM or use application-level mechanisms to ensure data is flushed to disk before taking the snapshot.

{% hint style="warning" %}
**Warning**

Live VM snapshots are disk-only snapshots and may not capture the in-memory state of applications. For applications requiring strict data consistency, it is recommended to either shut down the VM or use application-level mechanisms to ensure data is flushed to disk before taking the snapshot.
{% endhint %}

## Type 2 - VM Volume Snapshot

This type of snapshot is created when you snapshot a virtual machine that is using a volume for it's root disk. When you snapshot such a virtual machine:

**Snapshot File Type and Location**

* The VM's root disk volume is snapshotted as a bootable volume snapshot.
* If the VM has any other volumes mounted, **those volumes are also snapshotted**.
* The Image Library service creates a **metadata entry** in the Image Library database that tracks a reference to the root volume snapshot and the data volume snapshots, along with other required information.
* **No other file gets created or stored** in the Image Library for this snapshot.

When a new VM gets created using this snapshot, the Image Library service will use the metadata stored in the database behind the scenes to identify the root volume snapshot and the data volume snapshots if any. This information will then be supplied to the Compute Service which will then use those snapshots to create a new VM.

## Snapshot Creation

You can take a snapshot of virtual machines in running or powered off state. In the UI, select the snapshot action for the VM to take the snapshot.

## Snapshot Properties

All VM snapshots will have the `image_type=snapshot` property set. You can view this property in the properties column.

## Delete a VM Snapshot

To delete a VM snapshot, navigate to **Images & VM Snapshots** > **VM Snapshots**, select one or more snapshots, and choose the delete action. The wizard opens a **Delete Snapshot(s)** dialog listing the snapshots being removed.

When you delete a VM Volume Snapshot, the dialog also offers a **Delete all the associated volume snapshots** checkbox. The checkbox is selected by default and removes the root and data volume snapshots that the VM snapshot references in the same operation:

> The associated volume snapshot(s) for the selected VM snapshot(s) will be deleted as part of this operation. Uncheck the checkbox if you wish to retain them.

The dialog lists the volume snapshots that will be deleted. Uncheck the box to keep those child volume snapshots after the parent VM snapshot is removed; in that case, only the VM snapshot's metadata entry is deleted.


# Virtual Machine Leasing

## VM Leasing

VM leasing lets you enforce automatic *power-off* or *deletion* of virtual machines (VMs) after a fixed amount of time. Leases are defined **once per tenant** and inherited by every existing and future VM that belongs to that tenant. Administrators and self-service users can also override the lease for an individual VM when needed.

### Lease-End Actions

| Action        | Result                                                        | Typical Use Case                                   |
| ------------- | ------------------------------------------------------------- | -------------------------------------------------- |
| **Power Off** | VM is gracefully shut down. Disks and metadata are retained.  | Development or lab workloads you may revive later. |
| **Delete**    | VM (and its attached ephemeral disks) is permanently removed. | Disposable environments or cost-control scenarios. |

## Tenant-Level Lease Policy

You can set the tenant-level lease policy by navigating to **Settings** **>** **Tenants & Users** **>** **Tenants** (on left navigation pane) **>** choose a **Tenant** **>** click the **Manage Lease Policy** button on the table header.

The tenant lease policy management screens allow you to manage the following lease configuration:

**Enable Lease Policy**: Turns leasing on or off for the tenant.

**Lease Duration**: The total time in *Days, Hours, Minutes* that each VM may run before the lease expires. The timer starts at VM creation.

**Action Upon Lease Expiration:** Defines the action triggered on lease expiration.

* **Power Off VM (default**): the VM is shut down.
* **Delete VM:** the VM is deleted.

{% hint style="info" %}
**Info**

Caution If you shorten the lease and existing VMs have already exceeded the new duration, the chosen lease-end action executes immediately.
{% endhint %}

## Per-VM Lease Customization

You can modify a lease at the VM level by navigating to **Virtual Machines** on the left-hand navigation pane **>** **Virtual Machines** **>** select a **virtual machine** **>** **Other** dropdown in the table header **>** **Customize Lease**

You can perform the following actions on the VM lease from the customization screen.

* **Specify action upon lease expiration:** Defines the action triggered on lease expiration for the selected VM. **This selection overrides the action set at the tenant level.**
  * **Power Off VM (default**): the VM is shut down.
  * **Delete VM:** the VM is deleted.
* **End sooner**: Pick a new date/time earlier than the current expiration to force an earlier shutdown or deletion.
* **Extend unexpired VMs**: Extend the lease by *up to one additional lease period* as defined in the tenant policy (e.g., +5 days if the policy is 5 days). You can repeat extensions indefinitely.
* **Extend expired VMs**: Even after a VM’s lease has expired, you can grant an extension (subject to the same maximum).

{% hint style="info" %}
**Info**

After a VM is powered-off by the lease engine, any manual power-on is only momentary. PCD will power the VM off again almost immediately.
{% endhint %}

## Permissions Matrix

The following table outlines the VM leasing permissions that the Admin, Self-Service, and Read-Only roles have.

| Role                  | Set Tenant Policy | Edit Lease (VM-level) | Create VM |
| --------------------- | ----------------- | --------------------- | --------- |
| **Admin**             | ✔                 | ✔                     | ✔         |
| **Self-service User** | ✖                 | ✔                     | ✔         |
| **Read-only User**    | ✖                 | ✖                     | ✖         |

## Troubleshooting

| Symptom                                      | Cause                                      | Resolution                                             |
| -------------------------------------------- | ------------------------------------------ | ------------------------------------------------------ |
| VM powers off immediately after you start it | Lease already expired                      | Extend the lease or disable the tenant-level policy.   |
| Cannot extend lease beyond a certain date    | Max extension equals one full lease period | Increase the lease duration at the tenant level first. |


# Advanced Scheduling Options

This document describes advanced virtual machine scheduling strategies that you can use for specific enterprise use cases.

## NUMA Support

NUMA awareness allows the operating system of a virtual machine to intelligently schedule the workloads that it runs and minimize cross-node memory bandwidth.

In order to configure NUMA nodes for a virtual machine, you need to specify `hw:numa_nodes` property as part of a virtual machine flavor, then use that flavor to create the virtual machine.

For example, to restrict a virtual machine's vCPUs to a single host NUMA node, set `hw:numa_nodes=1`as part of the VM's flavor.

You can update an existing flavor and add this as new metadata - note that it will only apply to new VMs created using the flavor after the update. To update an existing flavor using `pcdctl` CLI, run the following command (replace with the name of the flavor you wish to update):

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl flavor set <flavor-name> --property hw:numa_nodes=1
```

{% endtab %}
{% endtabs %}

Some workloads have very demanding requirements for memory access latency or bandwidth that exceed the memory bandwidth available from a single NUMA node. For such workloads, it is beneficial to spread the virtual machine across multiple host NUMA nodes, even if the VM's RAM/vCPUs could theoretically fit on a single NUMA node. To force a virtual machine's vCPUs to spread across two host NUMA nodes, set `hw:numa_nodes=2`as part of the VM's flavor.

The allocation of a virtual machine's vCPUs and memory from different host NUMA nodes can be configured. This allows for asymmetric allocation of vCPUs and memory, which can be important for some workloads. You can configure the allocation of VM vCPUs and memory across each VM NUMA node using the `hw:numa_cpus.{num}` and `hw:numa_mem.{num}` metadata properties as part of the VM flavor. For example, to spread the 6 vCPUs and 6 GB of memory of a VM across two NUMA nodes and create an asymmetric 1:2 vCPU and memory mapping between the two nodes, run the following command using `pcdctl` CLI (replace with the name of the flavor you wish to update):

{% tabs %}
{% tab title="Bash" %}

```bash
$ pcdctl flavor set <flavor-name> --property hw:numa_nodes=2

$ pcdctl flavor set <flavor-name> \
  --property hw:numa_cpus.0=0,1 \
  --property hw:numa_mem.0=2048

$ pcdctl flavor set <flavor-name> \
  --property hw:numa_cpus.1=2,3,4,5 \
  --property hw:numa_mem.1=4096
```

{% endtab %}
{% endtabs %}

{% hint style="warning" %}
**Important**

The `{num}` parameter is an index of *VM* NUMA nodes and may not correspond to *host* NUMA nodes. For example, on a platform with two NUMA nodes, the scheduler may opt to place VM NUMA node 0, as referenced in `hw:numa_mem.0` on host NUMA node 1 and vice versa. Similarly, the CPUs bitmask specified in the value for `hw:numa_cpus.{num}` refer to *VM* vCPUs and may not correspond to *host* CPUs. As such, this feature cannot be used to constrain VMs to specific host CPUs or NUMA nodes.
{% endhint %}

{% hint style="danger" %}
**Important**

If the combined values of `hw:numa_cpus.{num}` or `hw:numa_mem.{num}` are greater than the available number of CPUs or memory respectively, this may result in VM provisioning failure.
{% endhint %}

## CPU Pinning

By default, virtual machine vCPU processes are not assigned to any particular host CPU. This allows for features like overcommitting of CPUs. In heavily contended systems, this provides optimal system performance, however that may come at the expense of performance and latency for individual VMs.

Some virtualized workloads require real-time or near real-time behavior, which is not possible with the latency introduced by this default CPU scheduling policy. For such VMs, it is beneficial to control which host CPUs are bound to a VM's vCPUs. This process is known as CPU pinning. No other VMs can use the CPUs of a pinned VM, thus preventing resource contention between VMs.

You can configure a VM to use CPU pinning by specifying the [`hw:cpu_policy`](https://docs.openstack.org/nova/latest/configuration/extra-specs.html#hw:cpu_policy) metadata property as part of the VM's flavor. There are three policies: `dedicated`, `mixed` and `shared` (the default). The `dedicated` CPU policy is used to specify that all CPUs of a VM should use pinned CPUs. To configure a flavor to use the `dedicated` CPU policy, run:

{% tabs %}
{% tab title="Bash" %}

```bash
$ pcdctl flavor set <flavor-name> --property hw:cpu_policy=dedicated
```

{% endtab %}
{% endtabs %}


# Attaching Virtual Persistent Memory To Guests

The virtual persistent memory (vPMEM) feature in <code class="expression">space.vars.product\_name</code> enables administrators to configure vPMEMs for virtual machines using physical persistent memory (PMEM) that can provide virtual devices.

Note that the vPMEM feature is not enabled by default in your <code class="expression">space.vars.product\_name</code>. Contact your <code class="expression">space.vars.product\_name</code> support liaison if you would like this feature enabled.

### Pre-requisites

* Persistent Memory Hardware
* Following modules loaded at the linux kernel level for hypervisor hosts:

  `dax_pmem`, `nd_pmem`, `device_dax`, `nd_btt`

### Configure PMEM Namespaces

Your first step is to configure your PMEM namespaces as vPMEM backends using the `ndctl` utility.

```bash
$ sudo ndctl create-namespace -s 30G -m devdax -M mem -n ns3
{
  "dev":"namespace1.0",
  "mode":"devdax",
  "map":"mem",
  "size":"30.00 GiB (32.21 GB)",
  "uuid":"937e9269-512b-4f65-9ac6-b74b61075c11",
  "raw_uuid":"17760832-a062-4aef-9d3b-95ea32038066",
  "daxregion":{
    "id":1,
    "size":"30.00 GiB (32.21 GB)",
    "align":2097152,
    "devices":[
    {
      "chardev":"dax1.0",
      "size":"30.00 GiB (32.21 GB)"
    }
    ]
  },
  "name":"ns3",
  "numa_node":1
}
```

Now you can list the available PMEM namespaces on the host:

<pre class="language-bash"><code class="lang-bash"><strong>$ ndctl list -X
</strong>[
  {
    ...
    "size":6440353792,
    ...
    "name":"ns0",
    ...
  },
  {
    ...
    "size":6440353792,
    ...
    "name":"ns1",
    ...
  },
  {
    ...
    "size":6440353792,
    ...
    "name":"ns2",
    ...
  },
  {
    ...
    "size":32210157568,
    ...
    "name":"ns3",
    ...
  }
]
</code></pre>

### Configure Memory Resource Classes

Now that you have the namespaces created, the next step is to specify which PMEM namespaces should be available to virtual machines. You can do this by defining one or more PMEM resource classes in the compute service advance configuration [as specified here](/private-cloud-director/virtualized-clusters/nova-override#configure-pmem-namespace).

Make sure to restart the compute service on each host after doing the above configuration.

### Configure a flavor

Now that you have the required PMEM namespaces and classes configured, you can specify a comma-separated list of the `$LABEL`s that correspond to the resource classes you defined in the step above, to the flavor’s `hw:pmem` property. Multiple instances of the same label are permitted.

```bash
pcdctl flavor set --property hw:pmem='6GB' my_flavor
pcdctl flavor set --property hw:pmem='6GB,LARGE' my_flavor_large
pcdctl flavor set --property hw:pmem='6GB,6GB' m1.medium
```

Based on the above examples, a `pcdctl server create` request with `my_flavor_large` will spawn a new virtual machine with two vPMEMs. One, corresponding to the `LARGE` label, will be `ns3`; the other, corresponding to the `6G` label, will be arbitrarily chosen from `ns0`, `ns1`, or `ns2`.

\ <br>


# Virtual TPM

This guide outlines the implementation and configuration requirements for Virtual Trusted Platform Module (vTPM) v2.0 support in <code class="expression">space.vars.product\_name</code>.

## What is Virtual Trusted Platform Module (vTPM)

A Trusted Platform Module (TPM) is a specialized hardware chip on your computer's motherboard that is designed to enhance your computer's security by securely storing cryptographic keys that are used for encryption and decryption.

vTPM v2.0 is a software-based representation of a traditional TPM 2.0 chip. It carries out the same hardware-based security functions as a physical Trusted Platform Module, such as attestation, key and random number generation, but without the physical TPM chip being required.

<code class="expression">space.vars.product\_name</code>'s vTPM solution leverages open source Barbican service for encryption management. <code class="expression">space.vars.product\_name</code>'s Virtual TPM service enables TPM support by default on <code class="expression">space.vars.product\_name</code> hypervisor hosts.

## TPM Version and Models Supported

The Virtual TPM configuration is controlled through **metadata that can be applied at the virtual machine image level**.

<code class="expression">space.vars.product\_name</code> currently only supports TPM version 2.0. <code class="expression">space.vars.product\_acronym</code> supports two models for vTPM

* `tpm-tis`: This option emulates a TPM device based on the TPM Interface Specification, which is the standard for TPM version 1.2.
* `tpm-crb`: This option emulates a TPM device based on the TPM 2.0 CRB (Chip Reference Board) specification.

## Image Preparation and Configuration

When you add TPM metadata to an image, any VM created using the image will automatically enable vTPM with the specified configuration. The metadata parameters control:

* The TPM model type (`tpm-tis` or `tpm-crb`)

You can apply these configurations by adding metadata to the image as below:

#### Image-level Properties

Following is the TPM metadata that you need to associate with a virtual machine image in order to enable vTPM for the VMs created with the image.

{% tabs %}
{% tab title="YAML" %}

```yaml
hw_tpm_version = 2.0
hw_tpm_model = tpm-crb
```

{% endtab %}
{% endtabs %}

For example, you might start with a standard Windows image without TPM support and later add TPM 2.0 support by updating the image metadata. Any new VMs created from this image will have TPM 2.0 enabled, while existing VMs remain unchanged.

Similarly, if you create a VM using an image with these tags, do not remove these metadata keys while you have active vTPM VMs as it may lead to unexpected failures.

#### Flavor-level Properties

To enable the support of vTPM live migration, ensure the flavor has the metadata

{% tabs %}
{% tab title="YAML" %}

```yaml
hw:tpm_secret_security = deployment
```

{% endtab %}
{% endtabs %}

Please note that this can only be set on the flavor metadata, not the image metadata.

#### Shared swtpm State Directory

The directory `/var/lib/libvirt/swtpm` must be shared across all compute hosts. Use NFS with the following mount options:

```
vers=3,local_lock=all
```

#### Consistent swtpm UID Across Hosts

Ensure that the swtpm user ID matches across all hosts. Platform9 pins this to UID 64130 starting 2026.4 release. If your UIDs are not consistent across hosts, please fix them so they are consistent. While doing this, you may also have to change the ownership of `/var/lib/swtpm-localca` to the swtpm user again.

## VM Deployment and Verification

1. Create a VM with vTPM support through the <code class="expression">space.vars.product\_name</code> UI
2. Make sure that the VM reaches "Active" state
3. Perform TPM verification:

#### TPM Verification for Windows VMs

{% tabs %}
{% tab title="YAML" %}

```yaml
1. Press Windows key + R
2. Execute tpm.msc
```

{% endtab %}
{% endtabs %}

#### TPM Verification For Linux VMs:

{% tabs %}
{% tab title="YAML" %}

```yaml
ls /dev | grep tpm    # Should show TPM device
```

{% endtab %}
{% endtabs %}

#### General TPM Verification:

{% tabs %}
{% tab title="YAML" %}

```yaml
# List running VMs
virsh list

# Verify TPM configuration
virsh dumpxml <VM_ID>
```

{% endtab %}
{% endtabs %}

Expected TPM configuration in XML:

{% tabs %}
{% tab title="YAML" %}

```yaml
<tpm model='tpm-crb'>
  <backend type='emulator' version='2.0' persistent_state='yes'>
    <encryption secret='<secret>'/>
  </backend>
  <alias name='tpm0'/>
</tpm>
```

{% endtab %}
{% endtabs %}

## Secret Management Verification

Run the following command to make sure that the secrets got created successfully:

{% tabs %}
{% tab title="YAML" %}

```yaml
openstack secret list
```

{% endtab %}
{% endtabs %}

Each VM with TPM should have a corresponding secret entry.

## vTPM State File

VTPM state files are located on directory `/var/lib/libvirt/swtpm` by default on the hypervisor host.

## Live Migration of vTPM-Enabled VMs

### Prerequisites for vTPM Live Migration

Live-migrating a VM with vTPM enabled requires additional setup beyond the standard [live migration prerequisites](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration-prerequisites). Because the vTPM emulator (`swtpm`) maintains state files for each VM, those files must be accessible on both the source and destination host during and after the migration.

All four conditions below must be satisfied before attempting live migration of a vTPM-enabled VM:

#### 1. Flavor Must Include the TPM Secret Security Property

The VM's flavor must have the following extra spec set:

```
hw:tpm_secret_security = deployment
```

Without this property, the Compute Service will not attempt to transfer the vTPM secret during migration. The VM may migrate but the vTPM device will fail to start on the destination host.

This property can only be set on the flavor, not the image:

```bash
pcdctl flavor set <FLAVOR_NAME> --property hw:tpm_secret_security=deployment
```

{% hint style="warning" %}
This flavor property must be set before the VM is created. Updating the flavor after a VM has been created does not retroactively enable vTPM migration support for that VM.
{% endhint %}

#### 2. Shared swtpm State Directory (NFS)

The swtpm state directory `/var/lib/libvirt/swtpm` must be mounted from the same NFS share on all compute hosts in the cluster. If the state directory is local to each host, the destination host will not have access to the VM's TPM state after migration.

Mount options must include:

```
vers=3,local_lock=all
```

Verify the mount is consistent across hosts:

```bash
# Run on each host — output should be identical
mount | grep swtpm
```

#### 3. Consistent swtpm UID Across All Hosts

The `swtpm` user must have the same numeric UID on every compute host. <code class="expression">space.vars.product\_name</code> pins this to **UID 64130** starting with the 2026.4 release. If hosts were added at different times or from different base images, UIDs may differ, causing permission errors when the destination host tries to access the shared state directory.

Check the swtpm UID on each host:

```bash
id swtpm
```

If the UID is not 64130, update it:

```bash
sudo usermod -u 64130 swtpm
sudo chown -R swtpm /var/lib/swtpm-localca
sudo chown -R swtpm /var/lib/libvirt/swtpm
```

Restart `libvirtd` on the host after changing the UID:

```bash
sudo systemctl restart libvirtd
```

#### 4. VM HA Shared Storage Requirements (If VM HA Is Enabled)

If your cluster has [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) enabled and the cluster contains vTPM VMs:

* The VM's ephemeral storage directory (the virtual machine storage path from the Cluster Blueprint) must be on shared storage mounted on all hypervisor hosts.
* The vTPM state directory (`/var/lib/libvirt/swtpm`) must also be on shared storage.
* Both directories must be owned by the `pf9` user and `pf9group` group.

### Diagnose vTPM Live Migration Failures

#### Symptom: Migration Fails with "swtpm" or "TPM" in the Error

If a live migration fails and the Compute Service log or `virsh` output references `swtpm`, `tpm`, or a permissions error, work through the following checks.

**Check 1: Verify swtpm UID consistency**

```bash
# On source host
id swtpm

# On destination host
id swtpm
```

If the UIDs differ, fix them using the steps in the "Consistent swtpm UID" section above.

**Check 2: Verify the shared NFS mount is healthy**

```bash
# On both source and destination hosts
ls -la /var/lib/libvirt/swtpm/
```

If the directory is empty on the destination host while the source has state files, the NFS mount is not working correctly. Check NFS mount status:

```bash
mount | grep swtpm
df -h /var/lib/libvirt/swtpm
```

Re-mount the NFS share if it has become stale:

```bash
sudo umount /var/lib/libvirt/swtpm
sudo mount /var/lib/libvirt/swtpm
```

**Check 3: Verify the flavor property**

Check that the VM's flavor includes `hw:tpm_secret_security = deployment`:

```bash
pcdctl flavor show <FLAVOR_NAME>
```

Look for `hw:tpm_secret_security` in the `properties` section. If it is absent, the flavor needs to be updated and a new VM created with the updated flavor — the property cannot be applied retroactively.

**Check 4: Inspect the libvirt and swtpm logs**

On the destination host, check for errors from `swtpm` after a failed migration attempt:

```bash
sudo journalctl -u swtpm@<VM_UUID> --since "10 minutes ago"
```

Also check the libvirt migration log on the source host:

```bash
sudo grep -i "swtpm\|tpm\|migration" /var/log/libvirt/libvirtd.log | tail -50
```

#### Symptom: VM Migrates Successfully but vTPM Is Unavailable After Migration

If the VM reaches **ACTIVE** on the destination host but the guest OS reports that the TPM device is missing or inaccessible:

1. Verify that the swtpm state file for the VM exists on the shared NFS path:

   ```bash
   ls /var/lib/libvirt/swtpm/<VM_UUID>/
   ```
2. Verify that the `swtpm` user on the destination host owns the state file:

   ```bash
   stat /var/lib/libvirt/swtpm/<VM_UUID>/
   ```

   The owner should be `swtpm` (UID 64130). If not, fix ownership:

   ```bash
   sudo chown -R swtpm:swtpm /var/lib/libvirt/swtpm/<VM_UUID>/
   ```
3. Perform a hard reboot of the VM to allow the guest to re-detect the TPM device:

   ```bash
   pcdctl server reboot --hard <VM_UUID>
   ```

#### Symptom: Maintenance Mode Does Not Migrate vTPM VMs

Maintenance Mode does not support live migration of vTPM-enabled VMs in releases prior to 2026.4. If you are running an earlier release and need to move a vTPM VM off a host for maintenance:

1. Shut down the VM (note: this causes downtime).
2. Perform a cold migration to the desired host.
3. Restart the VM on the destination host.

See [Virtual Machine Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration) for cold migration instructions.

## Related Pages

* [Virtual Machine Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration)
* [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha)
* [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode)
* [VM Flavors](/private-cloud-director/virtualized-clusters/virtualmachine/vm-flavors)


# Compute Service Advanced Configuration

You may occasionally need to adjust the default Compute Service configuration to enable or disable specific settings. This document describes the steps to do that.

## What is nova\_override.conf

`nova_override.conf` is the configuration file used by <code class="expression">space.vars.product\_name</code> compute service to store all the default configuration values.

You can use this override file to:

* Adjust logging levels or enabling more verbose logging by enabling debug mode. (Mostly commonly used case).
* Customize compute driver options.
* Modify quota defaults for specific tenants.

The `nova_override.conf` file is located in the following directory.

{% tabs %}
{% tab title="Bash" %}

```bash
/opt/pf9/etc/nova/conf.d/nova_override.conf
```

{% endtab %}
{% endtabs %}

## What to know before

* Updating nova\_override.conf file is an **advance operation and should only be performed by Administrators** with complete understanding of the property being edited and the possible impact. We highly recommend that you do this **only under guidance from the Platform9 support** or solution architect teams.
* Any changes made to the config file persist across upgrades to <code class="expression">space.vars.product\_name</code>, ensuring that your customized settings remain intact across upgrades.
* Any change to the config file requires service restart for the changes to take effect.
* The `nova_override` file **must be edited on** **each hypervisor host** to ensure that the changes to apply consistently across all hosts.

{% hint style="danger" %}
**Important**

Updating nova\_override.conf file is an **advance operation** and must only be performed by Administrators. We highly recommend **doing this only under Platform9 support guidance**.
{% endhint %}

## Restart Service

After editing the `nova_override.conf` file, you must restart the compute service for the changes to take effect. Run the following command on the host:

{% tabs %}
{% tab title="Bash" %}

```bash
systemctl restart pf9-ostackhost
```

{% endtab %}
{% endtabs %}

## Common Settings

For a complete list of Compute Service configuration options, refer to [nova.conf](https://docs.openstack.org/ocata/config-reference/compute/config-options.html).

Below we have some common scenarios and configurations options you may wish to edit.

### Enable Debug Logging

This is the most commonly used option. You can turn this on when troubleshooting issues with VM creation or VM volume attachment. Note that enabling this will generate highly verbose logs and may exhaust disk space on the host. We recommend disabling it in production after troubleshooting.

{% tabs %}
{% tab title="Bash" %}

```bash
[DEFAULT]
debug = true  # Enable detailed logging (useful for troubleshooting)
log_dir = /var/log/nova  # Log file storage location
```

{% endtab %}
{% endtabs %}

### Scheduler Configuration

This parameter determines the number of times the virtual machine scheduler will try to provision a virtual machine instance before failing. Only adjust this under specific guidance from Platform9 support team.

{% tabs %}
{% tab title="Bash" %}

```bash
[scheduler]
max_attempts = 5  # Number of retries for scheduling instances
```

{% endtab %}
{% endtabs %}

### Resource Over-Provisioning

`cpu_allocation_ratio` and `ram_allocation_ratio` define the CPU and memory over commitment ratios at the hypervisor host level. The defaults are 1:5 for RAM and 1:16 for CPU over commitment.

{% tabs %}
{% tab title="Bash" %}

```bash
[DEFAULT]
cpu_allocation_ratio = 16.0  # Overcommit CPU resources (Default: 16)
ram_allocation_ratio = 1.5  # Overcommit RAM resources (Default: 1.5)
```

{% endtab %}
{% endtabs %}

### Reserving Resources for System Overhead

You can reserve CPU and Memory resources at each hypervisor level that will not be allocated to virtual machines that get provisioned on that hypervisor. This is crucial for production environments to ensure that the hypervisor remains functional even when the virtual machines provisioned on it may be under heavy load. \_\_`reserved_host_memory_mb` represents memory in MB to be reserved for system overhead and `reserved_host_memory_mb` represents host cpus to be reserved. The default values for these are 512 for memory and 0 for CPU.

{% tabs %}
{% tab title="Bash" %}

```bash
reserved_host_memory_mb = 512  # host memory to reserve for system overhead 
reserved_host_cpus = 0 # host cpus to reserve for system overhead
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

Kernel-level overcommit checks during VM creation are disabled by setting `/proc/sys/vm/overcommit_memory` to `1`. This enables dynamic VM creation and hotplug operations. After host upgrades, the file is automatically reset to `1` even if it was manually modified.
{% endhint %}

## Advance Settings

{% hint style="warning" %}
**Warning**

These settings impact VM operations and performance. Modify only if necessary and test changes in a non-production environment.
{% endhint %}

To restrict the Compute Service so that a cluster's hosts fetch images only from that cluster's Image Library host, see [Restrict the Image Library Service to a Specific Cluster](/private-cloud-director/images-and-image-library/restrict-image-library-to-cluster).

### Compute Driver Configuration

{% tabs %}
{% tab title="Bash" %}

```bash
[DEFAULT]
compute_driver = nova.virt.libvirt.LibvirtDriver
```

{% endtab %}
{% endtabs %}

Do not change unless you are configuring a different virtualization technology (e.g., VMware, Xen, KVM).

### Networking Configuration

{% tabs %}
{% tab title="Bash" %}

```bash
[neutron]
service_plugins = router,metering
```

{% endtab %}
{% endtabs %}

Changing networking parameters without proper testing may break VM connectivity.

### Disk & Storage Settings

{% tabs %}
{% tab title="Bash" %}

```bash
[libvirt]
images_type = qcow2  # Default disk image format
live_migration_flag = VIR_MIGRATE_LIVE
```

{% endtab %}
{% endtabs %}

Modify only if you are implementing specific storage optimizations.

### CPU Mode and Model Configuration

{% tabs %}
{% tab title="Bash" %}

```bash
[libvirt]
# Specifies the CPU mode for virtual machine instances.
cpu_mode = custom
# Defines the custom CPU model(s) to be used with the hypervisor.
cpu_models = Broadwell-noTSX-IBRS
```

{% endtab %}
{% endtabs %}

This configuration is typically used when you need consistent CPU features across different hypervisors.

### Enable Storage Multipath

1. Make sure that service `multipathd` is running on all hosts.
2. Use the following config option to enable storage multipath support.

{% tabs %}
{% tab title="Bash" %}

```bash
volume_use_multipath = True
```

{% endtab %}
{% endtabs %}

### Configure PMEM Namespace

Use the following parameters to configure what Persistent Memory namespaces should be available for virtual machines to use when utilizing the vPMEM feature.

**Pre-requisites:**

* The PMEM namespaces must be precreated for this configuration to work.

Syntax to follow:

```bash
"$LABEL:$NSNAME[|$NSNAME][,$LABEL:$NSNAME[|$NSNAME]]"
```

* `$NSNAME` is the name of the PMEM namespace.
* `$LABEL` represents one specific resource class of PMEM memory that you wish to create and make available to your virtual machines, this is then used to generate the resource class name as `CUSTOM_PMEM_NAMESPACE_$LABEL`.

The configuration syntax allows the admin to associate one or more namespace `$NSNAME`s with an arbitrary `$LABEL` that can subsequently be used in a flavor to request one of those namespaces. It is recommended, but not required, for namespaces under a single `$LABEL` to be the same size.

The example below creates two resource classes of PMEM memory, 6GB and LARGE. It maps the class 6GB to namespaces ns0, ns1, ns2 and class LARGE to namespace ns3.

{% tabs %}
{% tab title="Bash" %}

```bash
[libvirt]
# pmem_namespaces=$LABEL:$NSNAME[|$NSNAME][,$LABEL:$NSNAME[|$NSNAME]]
pmem_namespaces = 6GB:ns0|ns1|ns2,LARGE:ns3
```

{% endtab %}
{% endtabs %}


# Troubleshooting And Log Files

## Troubleshooting Compute Service Problems

If your <code class="expression">space.vars.product\_name</code> Service Health dashboard indicates that the Compute Service is <mark style="color:red;">`unhealthy`</mark>, it may be because a large percentage of your hosts with the Hypervisor role assigned are either offline or the Compute service is unresponsive on those hosts. Refer to the [#log-files](#log-files "mention")to debug the issue further.

## Important Directories

`/var` – All logs go under `/var/log/pf9.` The only exception is the `pcdctl` log files which go under `/var/log`

`/opt` – All the packages for services installed by <code class="expression">space.vars.product\_name</code> go under `/opt/pf9`

`/var/opt/pf9` - Subdirectories for the Platform9 host agent and networking service go under here with temp files or state files.

## Log Files

Essential log files for debugging:

1. **Log files for all services** - Each host stores all its log files for the various components running on it at `/var/log/pf9`**.** Here you will find logs for compute, image library, storage, networking, and other services, depending on the roles assigned to that host. See the documentation for each service for more information about its log files.
2. **Compute service log** - The log file for the compute service is located `/var/log/pf9/ostackhost.log` on all hosts with the hypervisor role assigned. Useful for debugging issues with virtual machine creation or updates.
3. **Host agent log** - The log file for the Platform9 host agent that is installed on each host is located at `/var/log/pf9/hostagent.log`. This is helpful for debugging issues with host agent install or connectivity with the management plane.
4. **Communication agent log** - Located at `/var/log/pf9/comms/comms.log`. Log file for the Platform9 communications agent, which is responsible for ensuring the health and uptime of the host agent. Helpful for debugging issues regarding connectivity.
5. **Libvirt logs** - Located at `/var/log/libvirt/qemu/<vm-id>` where `<vm-id>`should be the UUID of the VM, and at `/var/log/libvirt/libvirtd.log`. Libvirt logs help with debugging any resource allocation or other issues with virtual machine instances.

## Troubleshooting Steps

1. Verify that the host has outbound network connectivity to the internet and the management plane controller:
   * `$ curl -s https://<FQDN>`
   * `$ ping www.google.com`
   * `$ telnet`[`www.google.com`](http://www.google.com/)`443`
2. If your environment has a proxy server, make sure that you have configured the Platform9 host agent installed on the host to route traffic via the proxy:
   * `$ sudo bash <path to host agent installer> --proxy=<proxy server>:<proxy port>`
3. Verify that the hostname on the host matches the hostname shown in the GUI.
4. The UUID given in the `host_id` file should match the host UUID shown on the GUI. `$ sudo cat /etc/pf9/host_id.conf`
5. If NTP is enabled, make sure the NTP servers are configured correctly across all the nodes within the configuration file `/etc/systemd/timesyncd.conf.d/conf.d`. You can verify whether the host's time is in sync using `$ sudo timedatectl status`.
6. Check that the Platform9 packages are installed on the host and the version which is shown on the GUI. `$ sudo apt list | grep -i pf9*`
7. Confirm that the host agent is running on your host:
   * `$ sudo systemctl status pf9-hostagent`
   * `$ sudo systemctl status pf9-ostackhost`
   * `$ sudo systemctl status pf9-comms`
   * `$ sudo systemctl status pf9-sidekick`
8. Check the below service logs for any “*`errors/timeouts/warnings`”,* during the time of issue. verify if any recent network changes might have impacted the connectivity.
   * `$ sudo cat /var/log/pf9/hostagent.log`
   * `$ sudo cat /var/log/pf9/ostackhost.log`
   * `$ sudo cat /var/log/pf9/comms/comms.log`
   * `$ sudo cat /var/log/pf9/sidekick/sidekick.log`
9. If these steps prove insufficient to resolve the issue, kindly reach out to the [Platform9 Support team](https://support.platform9.com/) for additional assistance.

### Most Common causes <a href="#most-common-causes" id="most-common-causes"></a>

* Connectivity to the management plane controller is broken or unreachable.
* Dependent services are down.
* Packages are corrupted or not installed.
* NTP is not synchronized.
* Insufficient disk space available.
* Ensure that the host that you are trying to add meets the [pre-requisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/pre-requisites#host-hardware-and-configuration).

## Operational Recovery Runbooks

The following runbooks cover the most common operational recovery scenarios for the Compute Service:

* [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state) — diagnose and restore VMs that land in ERROR after a host reboot or maintenance event, including hard reboot, evacuation, and rebuild options.
* [Recover libvirt and Compute Service Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-libvirt-and-compute-service) — diagnose and restart `libvirtd` and the Platform9 Compute Service when a hypervisor host becomes unresponsive.
* [Diagnose VM Scheduling Failures ("No Valid Host Was Found")](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/diagnose-vm-scheduling-failures) — step-by-step checklist for resolving VM placement failures caused by resource exhaustion, disabled hosts, host aggregate mismatches, or image property constraints.
* [Recover from Messaging Layer Failures Affecting VM Creation](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-messaging-layer-failures) — identify and recover from messaging layer problems that cause VM creates to hang or fail silently at scale.
* [Diagnose a Host Agent Stuck in Converging State](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/host-agent-stuck-converging) — locate the cause when a host remains in `converging` status after role assignment, including service start failures and log-rotation permission errors.
* [Diagnose Role Assignment Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/role-assignment-failures) — resolve HTTP 500 errors on role assignment and "Interface None is missing an IP" errors caused by incomplete host network configuration.
* [Troubleshoot Maintenance Mode Migration Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshoot-maintenance-mode-migrations) — diagnose and recover when VMs fail to migrate or are left stranded or in an error state during maintenance mode.
* [Resolve CPU Baseline Mismatch After Host Upgrade](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/cpu-baseline-mismatch) — correct cluster CPU baseline issues after a host upgrade, including how to use `cpu_model_extra_flags` for mixed-generation clusters.
* [Troubleshoot VM HA](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshoot-vm-ha) — step-by-step diagnostics for VM HA not triggering, Consul health prerequisites, shared and FC storage validation, enable/disable failures, and post-upgrade re-validation.
* [Resolve Live Migration Failures for Legacy Boot-from-Volume Hotplug VMs](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/hotplug-boot-from-volume-live-migration-failure) — one-time corrective step for boot-from-volume, hot-add-capable VMs created before the 2025.10 release that still fail live migration.

[HOST ISSUES - Previous<br>](https://platform9.com/kb/pcd-ts/host/host-issues)


# Troubleshoot Advanced Remote Support Issues

## Problem

* Platform9 Support Team is unable to access the host using ARS even if ARS is enabled.
* How to troubleshoot ARS (Advanced Remote Support) issues for hosts so that Platform9 Support Team can access the hosts anytime for log analysis and initial troubleshooting?

## Steps to Troubleshoot

{% stepper %}
{% step %}

#### Reset the pf9 user password (if expired or locked)

Perform these steps as root or a sudo user.

{% code title="Reset pf9 password" %}

```bash
sudo passwd pf9
```

{% endcode %}

If the account is expired, force a password change on next login:

{% code title="Force password change" %}

```bash
sudo chage -d 0 pf9
```

{% endcode %}
{% endstep %}

{% step %}

#### Ensure Platform9 services are running

Check the status of the core Platform9 services on the host:

{% code title="Check service status" %}

```bash
systemctl status pf9-hostagent
systemctl status pf9-comms
systemctl status pf9-sidekick
```

{% endcode %}

If any service is failing, restart it:

{% code title="Restart a service" %}

```bash
systemctl restart <service-name>
```

{% endcode %}
{% endstep %}

{% step %}

### Self-Hosted Installation

Following instructions apply to the [Self Hosted](/private-cloud-director/getting-started/self-hosted#overview) version of Private Cloud Director. Ignore these if using the SaaS hosted version.

#### Grant Temporary Access to Kubernetes Cluster Worker Nodes

Platform9 does not have direct access to the on-premise management plane Kubernetes cluster worker nodes. To allow Platform9 Support to access these nodes temporarily:

* Create a user named `pf9-support` on the worker nodes and assign a secure password.
* Share that password securely with the Platform9 Support Team (for example, via a Support Ticket or other secure communication channel).
  {% endstep %}
  {% endstepper %}

{% hint style="info" %}
If these steps do not resolve the issue, kindly reach out to the Platform9 Support Team for additional assistance: <https://support.platform9.com/hc/en-us/requests/new?ticket\\_form\\_id=360000924873>
{% endhint %}


# Troubleshooting Offline Or Failed Hosts

When you see a hypervisor host in a <code class="expression">space.vars.product\_name</code> environment that's **offline, failed, or in an error state**, here's how to debug the issue:

## Most common causes

* Connectivity to the <code class="expression">space.vars.product\_name</code> management plane controller is broken or unreachable.
* Dependent services are down.
* Packages are corrupted or not installed.
* NTP is not synchronized.
* Insufficient disk space available.
* Ensure that the host you are trying to add meets the [pre-requisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/pre-requisites#host-hardware-and-configuration).

## Steps To Troubleshoot

{% stepper %}
{% step %}

#### Verify outbound network connectivity

Ensure the host can reach the internet and the <code class="expression">space.vars.product\_name</code> management plane controller:

```bash
$ curl -s https://<FQDN>
$ ping www.google.com
$ telnet www.google.com 443
```

{% endstep %}

{% step %}

#### Proxy configuration (if applicable)

If your environment uses a proxy, configure the Platform9 host agent on the host to route traffic via the proxy:

```bash
$ sudo bash <path to host agent installer> --proxy=<proxy server>:<proxy port>
```

{% endstep %}

{% step %}

#### Verify hostname consistency

Make sure the hostname on the host matches the hostname shown in the GUI.
{% endstep %}

{% step %}

#### Verify host UUID

Confirm the UUID in the `host_id` file matches the host UUID shown in the GUI:

```bash
$ sudo cat /etc/pf9/host_id.conf
```

{% endstep %}

{% step %}

#### Verify NTP configuration

If NTP is enabled, ensure NTP servers are configured correctly across nodes in:

```
/etc/systemd/timesyncd.conf.d/conf.d
```

Verify whether the host's time is in sync:

```bash
$ sudo timedatectl status
```

{% endstep %}

{% step %}

#### Verify Platform9 packages

Check that Platform9 packages are installed and match the version shown in the <code class="expression">space.vars.product\_name</code> GUI:

```bash
$ sudo apt list | grep -i pf9*
```

{% endstep %}

{% step %}

#### Confirm host agent services are running

Check the status of the Platform9 services:

```bash
$ sudo systemctl status pf9-hostagent
$ sudo systemctl status pf9-ostackhost
$ sudo systemctl status pf9-comms
$ sudo systemctl status pf9-sidekick
```

{% endstep %}

{% step %}

#### Inspect service logs

Check logs for errors, timeouts, or warnings around the time of the issue and verify whether recent network changes could have impacted connectivity:

```bash
$ sudo cat /var/log/pf9/hostagent.log
$ sudo cat /var/log/pf9/ostackhost.log
$ sudo cat /var/log/pf9/comms/comms.log
$ sudo cat /var/log/pf9/sidekick/sidekick.log
```

{% endstep %}

{% step %}

#### Contact Platform9 Support

If these steps don't resolve the issue, reach out to the [Platform9 Support team](https://support.platform9.com/) for additional assistance.
{% endstep %}
{% endstepper %}


# Troubleshooting Host Onboarding Issues

This document describes how to identify and resolve issues that occur when onboarding a host using the `pcdctl prep-node` command.

## Most Common Causes

* Residual configuration from a prior failed node preparation can cause role assignment or installation to fail.
* The user running the prep-node command may not have proper `sudo` privileges, causing command failures.
* Packages are corrupted or not installed properly; `dpkg` or `apt` lock files preventing package installation or updates during prep-node execution. Review `/var/log/dpkg.log` to verify whether packages are partially installed or misconfigured.
* Incorrect proxy configurations provided to `prep-node` can block the download of required packages or scripts.
* `Firewalld` or other firewall rules may block required ports, preventing communication with the management plane.
* NTP is not synchronized, which may cause authentication or communication failures.
* Connectivity to the <code class="expression">space.vars.product\_name</code> management plane controller is broken or unreachable.

## Steps To Troubleshoot

The `pcdctl prep-node` without a prior configuration prompts interactively for account information that can be directly retrieved from the **GUI -> Infrastructure -> Cluster Hosts -> Add a New Host**. The GUI is pre-filled with the Account URL, Username, Region, and Tenant for the logged-in user.

All the configuration details like the Platform9 Account URL, Username, Password, Region, and Tenant are persisted in a local `config.json` file under `/pf9/db/`.

Logs for `pcdctl` command execution are stored in the `/pf9/logs/pcdctl-<DATE>.log` file.

{% stepper %}
{% step %}

#### Review pcdctl logs

Start by reviewing the pcdctl log file to trace the exact error.

{% tabs %}
{% tab title="pcdctl logs" %}

```bash
"msg":"Received a call to fetch keystone authentication for fqdn: https://[FQDN] and user: [USER] and tenant: [TENANT], mfa_token: <br>"}
"msg":"Error calling keystone API:Post \"https://[FQDN]/keystone/v3/auth/tokens?nocatalog\": dial tcp: lookup example1.pcd.platform9.co on 127.0.0.53:53: no such host<br>"}
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Verify input values

Ensure to provide all the details correctly, without typos and extra space. The installer authenticates with the Keystone API using these values. Any incorrect entry can cause authentication failure or DNS resolution errors.
{% endstep %}

{% step %}

#### Verify network connectivity

Verify that the host has outbound network connectivity to the internet and the <code class="expression">space.vars.product\_name</code> management plane controller:

* `$ curl -s https://<FQDN>`
* `$ ping www.google.com`
* `$ telnet`[`www.google.com`](http://www.google.com/)`443`
  {% endstep %}

{% step %}

#### Gather verbose logs

Execute `pcdctl prep-node` with the `--verbose` flag to gather detailed logs of the host preparation process, including each command executed, checks performed, and any warnings or errors encountered.
{% endstep %}

{% step %}

#### Validate prerequisites and host checks

Ensure the [primary prerequisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/pre-requisites#hypervisor-host-prerequisites) are met. Review the checks below for additional validation.

* Verify the host is running a supported Ubuntu version. Currently, Platform9 supports Ubuntu 22.04 and 24.04 for Private Cloud Director host onboarding.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo cat /etc/os-release | grep -E '^NAME=|^VERSION_ID='
```

{% endtab %}
{% endtabs %}

* Confirm the host has sufficient CPU cores and memory. Minimum 8 CPU cores and 16 GB RAM are recommended for host onboarding.

{% tabs %}
{% tab title="Bash" %}

```bash
# CPU cores
$ sudo grep -c ^processor /proc/cpuinfo

# Total memory
$ sudo free -h | grep Mem:
```

{% endtab %}
{% endtabs %}

* Verify the root partition (`/`) has adequate free space. Minimum 250 GB of free disk space is required.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo df -h /
```

{% endtab %}
{% endtabs %}

* Ensure no other package manager (e.g., apt or dpkg) is running in the background. If either command returns a running process, wait for it to finish or terminate it before continuing.

{% tabs %}
{% tab title="Bash" %}

```bash
# Review if any dpkg, apt process is held
$ sudo lsof /var/lib/dpkg/lock
$ sudo lsof /var/lib/apt/lists/lock

# Review package manager logs
$ sudo cat /var/log/dpkg.log 
$ sudo cat /var/log/apt/history.log
```

{% endtab %}
{% endtabs %}

* Confirm the root or current user has passwordless sudo privileges; the user must have unrestricted sudo privileges for all operations.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo -l
    (ALL : ALL) ALL
    (ALL) NOPASSWD: ALL
```

{% endtab %}
{% endtabs %}

* Check the status of `firewalld` to ensure it does not block Platform9 service communication. Platform9 recommends stopping and disabling `firewalld` on these hosts using `sudo systemctl stop firewalld` and `sudo systemctl disable firewalld`.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo systemctl is-active firewalld
inactive
```

{% endtab %}
{% endtabs %}

* Ensure NTP is enabled for accurate time synchronization across hosts. Verify if `systemd-timesyncd` is already running on the host, as it provides basic time synchronization.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo systemctl status systemd-timesyncd
$ sudo timedatectl status
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Proxy configuration (if applicable)

For [hosts using a proxy server for outbound connectivity](https://platform9.com/docs/private-cloud-director/2025.8/private-cloud-director/pre-requisites#connectivity-via-https-proxy-server), ensure that the `/etc/environment` file has required variables configured. Also configure the package manager (apt) to properly fetch required packages through the proxy server; refer to the linked documentation.
{% endstep %}

{% step %}

#### Verify hostagent installation and logs

As the last step, the hostagent package is downloaded and installed on the host. Verify the `pf9-hostagent.service` status and `/var/log/pf9/hostagent.log` file to track the progress.

{% tabs %}
{% tab title="Bash" %}

```bash
$ sudo systemctl status pf9-hostagent
$ sudo cat /var/log/pf9/hostagent.log
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Post-onboarding: authorization and role assignment

After onboarding the host to the <code class="expression">space.vars.product\_name</code>, it can be [Authorized & Assigned Roles](https://platform9.com/docs/private-cloud-director/private-cloud-director/add-hosts-Virtualized-Cluster#step-2-authorize-host-and-assign-roles). This involves downloading and installation of service specific packages and service initialization, which can be monitored through `/var/log/pf9/hostagent.log`. The `.deb` packages are downloaded inside `/var/cache/pf9apps` directory.
{% endstep %}

{% step %}

#### Contact support

If these steps prove insufficient to resolve the issue, reach out to the [Platform9 Support team](https://support.platform9.com/) for additional assistance.
{% endstep %}
{% endstepper %}

## Re-Onboarding Recovery <a href="#re-onboarding-recovery" id="re-onboarding-recovery"></a>

## Overview

When a host fails to onboard or re-onboard, the cause is often stale agent state left over from a previous attempt. This can include a broken libvirtd configuration, a mismatched host identity file, leftover host configuration, or partially applied role state. Cleaning up that state before retrying gives the onboarding process a clean starting point.

In this section, you will identify and remove stale host state, resolve broken libvirtd conditions, and re-onboard the host cleanly.

{% hint style="warning" %}
Before removing any host state, confirm that no VMs are running on the host. Use `virsh list --all` to verify. If VMs are present, migrate or shut them down first.
{% endhint %}

## Decommission the Host Before Re-Onboarding

If the host was previously onboarded and is reachable, use `pcdctl` to decommission it cleanly before removing its state. This avoids leaving stale compute service records in the management plane database.

```bash
pcdctl decommission-node
```

Do NOT use the `-r` / `--skip-installed-role-check` flag. If roles are still applied, deauthorize them from the <code class="expression">space.vars.product\_name</code> UI first (**Infrastructure > Cluster Hosts**, select the host, click **Edit Roles**, remove all roles), then run `pcdctl decommission-node`.

See [Hypervisor Role Deauthorization and Reauthorization](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/hypervisor-role-deauthorization-and-reauthorization) for the full decommission guidelines.

## Remove Stale Host Identity and Configuration

After decommissioning, or if the host cannot be decommissioned cleanly, remove the identity and configuration files that persist host state across re-onboarding attempts:

```bash
# Remove host identity file (forces a new host UUID on next onboarding)
sudo rm -f /etc/pf9/host_id.conf

# Remove persisted pcdctl config (account URL, credentials, tenant)
sudo rm -f /pf9/db/config.json

# Remove downloaded package cache
sudo rm -rf /var/cache/pf9apps/*

# Remove Platform9 state directory
sudo rm -rf /opt/pf9/data/state/*
```

{% hint style="info" %}
Removing `/etc/pf9/host_id.conf` causes the management plane to register the host as a new host on the next `pcdctl prep-node` run. The old host entry in the UI can be deleted from **Infrastructure > Cluster Hosts** after re-onboarding succeeds.
{% endhint %}

## Resolve Broken libvirtd State

A failed or interrupted role application can leave libvirtd in a broken state that prevents the Hypervisor role from applying cleanly on re-onboarding. Symptoms include:

* `pf9-ostackhost` service failing to start after role assignment.
* Errors in `/var/log/pf9/hostagent.log` mentioning `libvirtd` socket or connection refused.
* `virsh list` returning connection errors.

To check and reset libvirtd:

```bash
# Check libvirtd status
sudo systemctl status libvirtd

# If libvirtd is failed or inactive, attempt a restart
sudo systemctl restart libvirtd

# Verify the libvirtd socket is present
ls -la /var/run/libvirt/libvirt-sock
```

If libvirtd still fails to start, check its log for the specific error:

```bash
sudo journalctl -u libvirtd --since "1 hour ago" | tail -50
```

Common causes of a broken libvirtd state after a failed role application:

* Stale QEMU hook scripts left in `/etc/libvirt/hooks/` from a prior Platform9 installation.
* Corrupted `/etc/libvirt/libvirtd.conf` settings written by a failed role application.

If the libvirtd configuration is corrupted, restore the defaults and restart:

```bash
sudo mv /etc/libvirt/libvirtd.conf /etc/libvirt/libvirtd.conf.bak
sudo apt-get install --reinstall libvirt-daemon-system
sudo systemctl restart libvirtd
```

After libvirtd is running cleanly, re-attempt role assignment from the <code class="expression">space.vars.product\_name</code> UI.

## Re-Onboard the Host

Once stale state is cleared and libvirtd is healthy, re-run the onboarding command:

```bash
pcdctl prep-node
```

Provide the account URL, username, region, and tenant (project) when prompted, or pass them as flags. After `prep-node` completes, authorize the host and assign roles from **Infrastructure > Cluster Hosts** in the UI.

Monitor onboarding progress in the hostagent log:

```bash
tail -f /var/log/pf9/hostagent.log
```

Look for `converge successful` or `role application complete` messages. If the host gets stuck in `converging` state, see [Diagnose a Host Agent Stuck in Converging State](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/host-agent-stuck-converging).

If re-onboarding fails again after clearing stale state, contact the [Platform9 Support team](https://support.platform9.com/) with the output of `/var/log/pf9/hostagent.log` and the `pcdctl` log from `/pf9/logs/`.


# Hypervisor Role Deauthorization and Reauthorization

In `space.vars.product_acronym`, when a hypervisor role is applied, it auto configures the compute and network services on the hosts depending on the settings defined in the Cluster Blueprint.

This guide enables you to debug issues that you may run into when authorizing or deauthorizing new compute service hosts.

## A) Improper clean up of state files

Files such as `/opt/pf9/data/state/compute_id` define the unique compute id for a hypervisor host. When the same host is re-authorised, a new `compute_id` can fail to create a resource provider record in the placement API.

What to do if this happens:

* Deauthorize the hypervisor role from the host cleanly.
* Decommission the host.
* Re-authorise the host.

## B) Forceful decommission leaving stale compute services in database

Forceful decommission of the host using `pcdctl` CLI without the role check (`pcdctl decommission-node -r`) can leave compute service state in the databases, causing conflicts when the same host is added back with the same hostname.\
Note: Option `-r/--skip-installed-role-check` is no longer supported.

What to do if this happens:

* As a best practice, forceful decommission is not recommended. However if it was used and there are stale compute services left in the Database, delete them with:

```bash
# openstack compute service delete <uuid>
```

* Restart the pf9-ostackhost service:

```bash
# systemctl restart pf9-ostackhost
```

## C) Forceful decommission leaving stale instances on the host

Forceful decommission using `pcdctl` without the role check (`pcdctl decommission-node -r`) can leave a non-empty instance list in `virsh list` output, causing the role removal to fail due to existing instances on the hypervisor.

What to do if this happens:

* As a best practice, forceful decommission is not recommended. However if there are stale instances left on the host, delete them using the commands below.

List all VMs:

```bash
# virsh list --all
```

Destroy and undefine those VMs (use the ID from the previous command):

```bash
# virsh destroy <id>
# virsh undefine <id>
```

* Restart the pf9-ostackhost service:

```bash
# systemctl restart pf9-ostackhost
```

## Host Authorization And Deauthorization Guidelines

{% stepper %}
{% step %}

#### Option 1 — Host has only hypervisor role

Suggested process to de-auth/auth:

* Deauthorize the hypervisor role from the host using the `space.vars.product_name` console cleanly.
* Decommission the host using `pcdctl decommission-node` (DO NOT use the `-r` option to skip role checks).
* Onboard the host using the instructions in the UI to add a new host.
* Re-authorise the host with the hypervisor role using the `space.vars.product_name` console.
  {% endstep %}

{% step %}

#### Option 2 — Host has multiple roles and you want to retain other roles

Suggested process to de-auth/auth while retaining existing host configuration and other roles (Image Library, Persistent Storage, etc.)

As a part of the networking configuration when the hypervisor role is applied OpenVSwitch bridges are created based on the host config defined in the blueprint. These bridges are not deleted when the hypervisor role is removed from a host. This requires manual cleanup of the bridges on the host for clean deauthorization. Skipping this step can result in connectivity loss if the host is rebooted.

Steps for cleanup:

* Deauthorize the hypervisor role from the host using the `space.vars.product_name` console cleanly.
* Delete all the leftover OVS bridges and reapply the original netplan:

  ```bash
   sudo ovs-vsctl list-br | xargs -r -I {} sudo ovs-vsctl del-br {};
   netplan apply;
  ```
* In case you want to reauthorize the host with the hypervisor role reapply the role from the `space.vars.product_name` console.
  {% endstep %}
  {% endstepper %}

{% hint style="warning" %}
Forceful decommission using `pcdctl decommission-node -r` is not recommended because it can leave stale state (compute services or instances) that cause conflicts when re-adding the host.
{% endhint %}


# Failed to Deploy Virtual Machine

This guide provides step-by-step instructions for troubleshooting and resolving issues when creating a virtual machine (VM) fails in <code class="expression">space.vars.product\_name</code>.

## Various VM deployment Methods

* Launch an instance from an Image to quickly deploy a pre-configured environment.
* Launch an instance from a New Volume to create a fresh setup with dedicated storage.
* Launch an instance from an Existing Volume to utilize previously used storage for seamless continuation.
* Launch an instance from a VM snapshot to capture current state and restore it precisely as it was.
* Launch an instance from a Volume Snapshot to ensure data integrity by reverting to a specific point in time.

## Most common causes

* Insufficient resources (CPU, RAM, storage).
* Incorrect network configurations or security groups.
* Unavailable or corrupted images.
* Issues with the scheduler or compute nodes.
* Permission or quota restrictions.
* Virtualised cluster mismatches.

## Deep Dive

The <code class="expression">space.vars.product\_name</code> VM creation process is similar for all VM deployment methods mentioned earlier, and the workflow is orchestrated primarily by the **Compute service** (Nova). This flow involves a series of steps with critical validations at each stage to ensure the request is valid, resources are available, and the VM is provisioned correctly.

{% hint style="info" %}
Below logs can be only reviewed in Self Hosted Private Cloud Director. For SAAS model, kindly contact the Platform9 Support Team
{% endhint %}

{% stepper %}
{% step %}

#### User Request & API Validation

This is the initial stage where the user's request is received and authenticated.

* **User Request:** A user submits a request to create a VM (also called an instance) via the OpenStack CLI, <code class="expression">space.vars.product\_name</code> dashboard, or direct API call. Key parameters are specified, including the **image**, **flavor**, **network**, **security group**, and **key pair**.
* **Keystone Authentication:** The request is sent to the `nova-api-osapi` Pod, which immediately validates the user's authentication token with the **Identity service** (keystone). This ensures the user is who they claim to be. The output below shows the initial VM creation request was successfully received by the Nova API and was accepted with a 202 status code.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
$ kubectl logs deployment/nova-api-osapi -n <WORKLOAD_REGION> | grep "POST /v2.1"
INFO nova.osapi_compute.wsgi.server [None [REQ_ID] [USER_ID] [TENANT_ID] - - default default] [IP] "POST /v2.1/[tenant_id]/servers HTTP/1.1" status: 202 len: [.] time: [.]
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
Here a unique `REQ_ID` will be generated, which will be further used for tracking the request in other component logs.
{% endhint %}

* **Authorization & Quota Checks:** The Nova API performs two key validations:
  * **Authorization:** It verifies that the user has the necessary permissions to create a VM within the specified project.
  * **Quota Check:** It confirms the project has enough available resources (vCPUs, RAM, instances, etc.) to fulfil the request based on the chosen flavor.
* **Initial Database Entry:** The database name is `nova`. The `nova-conductor` service is the only service that writes to the database. The other Compute services access the database through the `nova-conductor` service. If all checks pass, `nova-conductor` creates a database record for the new VM and sets its status to **BUILDING(None)**.
  {% endstep %}

{% step %}

#### Scheduling & Resource Selection

After the initial validation, the request is sent to the Nova Scheduler, which decides where to place the VM.

* **Message Queue:** The Nova API sends a message to the **Nova Scheduler** via a message queue (RabbitMQ), containing all the VM's requirements.
* The **Nova scheduler** queries the **Placement API** to find a suitable **resource provider (compute node)** that has enough resources based on host filters and host weighing.
* **Host Filtering:** The `nova-scheduler` begins by filtering out unsuitable hosts. This process checks for:
  * **Resource availability:** It ensures the host has sufficient free RAM, disk space, and vCPUs.
  * **Compatibility:** It verifies the host is compatible with the image properties and any specific requirements.
  * **Availability Zones:** It confirms the host is in the requested availability zone.
  * **Image Metadata:** It checks the image metadata if there is a specific metadata filter for the image. E.g. Images with metadata SRIOV, vTPM, etc.
  * Many more other filters.

{% hint style="info" %}
Other filters:

Details on Nova filters are available on [Scheduler filters](https://docs.openstack.org/nova/rocky/user/filter-scheduler.html).
{% endhint %}

* **Host Weighing:** The remaining hosts are then ranked based on a weighting system. This can be configured to prioritise hosts with the least load or those that have been least recently used to ensure balanced resource distribution.

{% hint style="warning" %}
At this stage, if the scheduler doesn’t find any suitable host to deploy an instance, it gives a **“No Valid Host Found”** error.
{% endhint %}

* **Placement Reservation:** The `nova-scheduler` service queries `placement API` to fetch eligible compute nodes. Once a host is selected, the scheduler **makes a provisional allocation** by creating a **"claim"** via: `PUT /allocations/[VM_UUID]`. Placement API `PUT` requests will have VM allocation ID logs that look like the following:

{% tabs %}
{% tab title="Sample Logs" %}

```dart
$ kubectl logs deployment/placement-api -n <WORKLOAD_REGION> | grep <req_id>
INFO placement.requestlog [[REQ_ID] [REQ_ID] [USER_ID] [TENANT_ID] - - default default] [IP] "PUT /allocations/[VM_UUID]" status: 204 len: 0 microversion: 1.36
```

{% endtab %}
{% endtabs %}

The `nova-scheduler` pod logs can be reviewed against the request ID captured from `nova-api-osapi` the pod. In the snippet below, VM requests will verify a suitable host for VM deployment.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
$ kubectl logs deployment/nova-scheduler -n <WORKLOAD_REGION> | grep <req_id>
WARNING nova.scheduler.filters.aggregate_image_properties_isolation [None [REQ_ID] [USER_ID] [TENANT_ID] - - default default] Host '[HOST_UUID]' has a metadata key 'availability_zone' that is not present in the image metadata.
```

{% endtab %}
{% endtabs %}

* The **Nova-scheduler** sends the database update request with the host information to **Nova-conductor** which further updates the database and sets VM status to **BUILDING (Scheduling)**. Then the request is passed to the **Nova-compute** service.
  {% endstep %}

{% step %}

#### Compute & Final Service-Level Validation

The **Nova-compute** service on the selected host performs the final provisioning steps.

* **Resource Allocation:** The **Nova Compute** service receives the scheduling decision and begins allocating resources. It interacts with:
  * **Glance:** It requests the VM image. **Validation occurs here** as `glance-api` pod can perform a signature check to ensure the image's integrity. If an image is not available, then it errors out. Below is an example of an `GET` image request.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO eventlet.wsgi.server [None [REQ_ID] [USER_ID] [TENANT_ID] - - default default] 127.0.0.1 - - [...] "GET /v2/images/[IMAGE_UUID] HTTP/1.0" 200 1132 0.032435
```

{% endtab %}
{% endtabs %}

* **Neutron:** It requests network resources, and `neutron-server` pod **validates** that the specified network and security groups exist and are accessible to the user. It then allocates a virtual network interface and an IP address. Below example shows the IP and network interface port information. **Nova-conductor** further updates the database and sets VM status to **BUILDING(Networking)**.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
$ kubectl logs deployment/neutron-server -n <WORKLOAD_REGION>
INFO neutron.wsgi [[REQ_ID] [REQ_ID] [USER_ID] [TENANT_ID] - - default default] 127.0.0.1 "GET /v2.0/floatingips?fixed_ip_address=[VM_IP_Address]&port_id=[VM_Interface_Port_ID] HTTP/1.1" status: 200  len: [.] time: [.]
```

{% endtab %}
{% endtabs %}

* **Cinder (if applicable):** If a persistent boot volume is requested, **Cinder validates** that the volume is available and attaches it to the VM. **Nova-conductor** further updates the database and sets VM status to **BUILDING(Block\_Device\_Mapping)**.
* **Hypervisor Instruction:** Once all resources are confirmed, `nova-compute` instructs the `pf9-ostackhost` service on the hypervisor (Libvirtd KVM) to create the VM using the image, flavor, and other parameters. The VM then boots. The `pf9-ostackhost` logs look like the below, which outline details like claim successful, device path, network information, time elapsed to spawn an instance, etc.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO nova.compute.claims [[REQ_ID] [USERNAME] service] [instance: [VM_UUID]] Claim successful on node [SELECTED_NODE_NAME]
..
INFO os_vif [[REQ_ID] [USERNAME] service] Successfully plugged vif VIFOpenVSwitch(active=False,address=[MAC_ADDRESS],bridge_name='br-int',has_traffic_filtering=True,id=[INTERFACE_ID],network=Network([NETWORK_ID]),plugin='ovs',port_profile=VIFPortProfileOpenVSwitch,preserve_on_delete=False,vif_name='[VM_TAP_INTERFACE]')
..
INFO nova.compute.manager [[REQ_ID] [USERNAME] service] [instance: [VM_UUID]] Took 3.74 seconds to spawn the instance on the hypervisor.
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### VM Configuration & Finalization

The final step involves configuring the guest OS and updating the status.

* **Cloud-init:** As the VM boots, **Cloud-init** runs with `169.254.169.254` IP address and retrieves metadata from Nova. The cloud-init logs are available within the VM. It performs validations on this metadata before:
  * Injecting the SSH key.
  * Configuring networking and the hostname.
  * Executing any custom user data scripts.
* **Status Update:** The `nova-compute` service updates the VM's status in the database to **ACTIVE**, indicating a successful creation. The VM is now ready for the user to access.
  {% endstep %}
  {% endstepper %}

## Procedure

{% hint style="info" %}
OpenStack CLI references virtual machines as 'server'.
{% endhint %}

{% stepper %}
{% step %}

#### Get the VM status

Use the PCD UI or CLI to check the error message. Look for `status` and `fault` fields to understand the issue.

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack server show <VM_UUID>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Validate Compute Service Status

Get the Compute Service state and ensure it is `up` and status is `enabled`.

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack compute service list
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Trace the VM Events

Retrieve the Request ID i.e. `REQ_ID` from the server event list details, which uniquely identifies the request. This `REQ_ID` is displayed in the first column of the server events list command and helps track request failures.

{% hint style="warning" %}
This `REQ_ID` is crucial for troubleshooting the VM creation issues.
{% endhint %}

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack server event list <VM_UUID>
$ openstack server event show <VM_UUID> <REQ_ID>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Review the Pods and their logs on the Management plane

{% hint style="info" %}
Step 4 is applicable only for Self Hosted Private Cloud Director
{% endhint %}

Management plane has Pods like `Nova-api-osapi`, `Nova-scheduler` and `Nova-conductor`. Review all these pods:

* Check if they are in "CrashLoopBackOff/OOMkilled/Pending/Error/Init" state.
* Verify if all containers in the pods are Running.
* See the events section in pod describe output.
* Review pods logs using `REQ_ID` or `VM_UUID` for relevant details.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl get pods -o wide -n <WORKLOAD_REGION> | grep -i "nova"

$ kubectl describe -n <WORKLOAD_REGION> <NOVA_API_OSAPI_POD>
$ kubectl describe -n <WORKLOAD_REGION> <NOVA_SCHEDULER_POD>
$ kubectl describe -n <WORKLOAD_REGION> <NOVA_CONDUCTOR_POD>

$ kubectl logs -n <WORKLOAD_REGION> <NOVA_API_OSAPI_POD>
$ kubectl logs -n <WORKLOAD_REGION> <NOVA_SCHEDULER_POD>
$ kubectl logs -n <WORKLOAD_REGION> <NOVA_CONDUCTOR_POD>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Validate the Image and Flavor

Check if the image is available and not corrupted. Ensure the resources requested in the flavor are available on the underlying hosts.

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack image show <IMAGE_ID>
$ openstack flavor show <FLAVOR_ID>
$ openstack hypervisor stats show
$ openstack hypervisor show <HYPERVISOR_NAME>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Validate the service status on the affected VM's underlying hypervisor

Validate if services listed below are running on the Underlying Hypervisor:

{% tabs %}
{% tab title="Command" %}

```bash
$ sudo systemctl status pf9-hostagent
$ sudo systemctl status pf9-ostackhost 
$ sudo systemctl status pf9-cindervolume-base
$ sudo systemctl status pf9-neutron-ovn-metadata-agent
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Check the logs on the affected VM's hypervisor

* Compute Node: `Ostackhost` logs are responsible for provisioning the compute resource required by the VM. Review the latest logs and search for `REQ_ID` or `VM_UUID`.

{% tabs %}
{% tab title="Command" %}

```bash
$ less /var/log/pf9/ostackhost.log
```

{% endtab %}
{% endtabs %}

* Cinder Storage Node: `cindervolume-base` logs are responsible for provisioning the storage resources required by the VM. Review the latest logs and search for `REQ_ID` or `VM_UUID`.

{% tabs %}
{% tab title="Command" %}

```bash
$ less /var/log/pf9/cindervolume-base.log
```

{% endtab %}
{% endtabs %}

* Network Node: `pf9-neutron-ovn-metadata-agent` logs are responsible for provisioning the connectivity and networking resources required by the VM. Review the latest logs and search for `REQ_ID` or `VM_UUID`.

{% tabs %}
{% tab title="Command" %}

```bash
$ less /var/log/pf9/pf9-neutron-ovn-metadata-agent.log
```

{% endtab %}
{% endtabs %}
{% endstep %}
{% endstepper %}

If these steps prove insufficient to resolve the issue, kindly reach out to the [Platform9 Support Team](https://support.platform9.com/) for additional assistance.


# VM Boot Stuck With Booting From Hard Disk Console Message

## Problem

VM Boot Stuck with "Booting from Hard Disk" Console Message with VM status `ACTIVE` on the console.

## Cause

The VM Image is built to set boot in UEFI mode, not legacy BIOS which is default.

## Diagnostics

{% stepper %}
{% step %}

#### Check VM status

Verify the VM is running and VM Status shows active on the PCD GUI.
{% endstep %}

{% step %}

#### Connect QCOW2 image using qemu-nbd

Run:

```bash
sudo qemu-nbd --connect=/dev/nbd0 <IMAGE_NAME>.qcow2
```

If the above command fails with the error:

qemu-nbd: Failed to open /dev/nbd0: No such file or directory

then load the nbd kernel module:

```bash
sudo modprobe nbd max_part=8
```

Then retry:

```bash
sudo qemu-nbd --connect=/dev/nbd0 <IMAGE_NAME>.qcow2
```

{% endstep %}

{% step %}

#### Inspect disk partition types

Check if any disk has type "EFI" (which indicates the qcow2 image is set to boot in UEFI mode):

```bash
sudo fdisk -l /dev/nbd0
```

{% endstep %}

{% step %}

#### Disconnect the NBD

When finished, disconnect:

```bash
sudo qemu-nbd --disconnect /dev/nbd0
```

{% endstep %}
{% endstepper %}

## Resolution

{% hint style="info" %}
These steps are applicable only when the Diagnostics step that inspects partitions shows any disk with type "EFI".
{% endhint %}

{% stepper %}
{% step %}

#### Update image properties in PCD GUI

Using the PCD GUI edit the image properties (VM → Images → Select Image → Edit properties) and add the following values:

```bash
hw_firmware_type=uefi
hw_machine_type=q35
```

{% endstep %}

{% step %}

#### Rebuild the VM

Rebuild the VM using the updated image properties.
{% endstep %}
{% endstepper %}

## Validation

1. Check if the VM boots successfully.
2. If these steps prove insufficient to resolve the issue, reach out to the [Platform9 Support Team](https://support.platform9.com/hc/en-us) for additional assistance.


# Windows VM Fails To Boot

## Problem

A Windows virtual machine migrated from VMware to PCD fails to boot successfully. During startup, the VM stops at the “Choose your keyboard layout” screen and does not proceed further into Windows.

## Cause

The issue occurs because Windows boot files or BCD (Boot Configuration Data) become corrupted or misaligned during the migration process.

When Windows cannot locate valid boot information, it automatically enters recovery mode and prompts for keyboard layout selection.

Possible contributing factors include:

* Missing or outdated VirtIO drivers during migration
* Incomplete or interrupted conversion process
* Boot partition corruption due to disk mapping or UUID mismatch
* File system inconsistencies on the migrated disk

## Diagnostics

{% stepper %}
{% step %}

#### Confirm boot state

Verify that the VM halts at the “Choose your keyboard layout” screen during boot.
{% endstep %}

{% step %}

#### Inspect disk from recovery media

Use WinPE or a Windows recovery ISO to access the disk and check for file system integrity issues.
{% endstep %}
{% endstepper %}

## Resolution

Running a simple filesystem check from a WinPE ISO or Windows recovery media can resolve the issue.

{% code title="Run chkdsk" %}

```powershell
chkdsk C: /f
```

{% endcode %}

Running chkdsk /f fixes most common NTFS or MBR inconsistencies that trigger WinRE.

## Additional Information

The chkdsk command runs the Check Disk utility to scan the specified drive (C: in this case) for disk or filesystem errors, and the /f flag instructs it to automatically fix any issues found. If your Windows installation uses a different drive letter, replace C: with the appropriate one for your system volume.


# Recover VMs in ERROR State After Host Reboot or Patching

## Overview

After a hypervisor host is rebooted or patched, virtual machines that were running on that host may land in **ERROR** state instead of returning to **ACTIVE**. This guide explains why that happens, how to safely diagnose the cause, and how to bring VMs back into service — including when to attempt a hard reboot, when to use evacuation, and when to escalate.

In this guide, you will restore ERROR-state VMs to a running state following a host reboot or maintenance event.

## Why VMs Land in ERROR After a Host Reboot

When a host is rebooted outside of [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode), running VMs are not migrated first. The Compute Service treats an unexpected host disappearance as a failure. Several conditions lead to VMs landing in ERROR:

* The host took too long to come back and the Compute Service timed out the pending state.
* The VM's state on disk was inconsistent at the time of the reboot (for example, a write was in progress).
* The Compute Service on the host did not restart cleanly after the reboot, so it reported VMs as failed during reconciliation.
* [VM HA](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha) attempted to evacuate the VM to another host but the evacuation failed (for example, the destination host had insufficient resources or the storage was unavailable at that moment), leaving the VM in ERROR.

{% hint style="warning" %}
**Data safety caution:** Before resetting a VM's state, confirm the host is back online and healthy. Performing a hard reboot or rebuild on a VM whose underlying disk is still inaccessible may corrupt the guest OS. Verify storage health before proceeding.
{% endhint %}

## Prerequisites

* You can reach the <code class="expression">space.vars.product\_name</code> UI or have `pcdctl` access.
* The hypervisor host is back online and shows as **Active** in the UI under **Infrastructure > Cluster Hosts**.
* Storage (ephemeral shared storage or block storage volumes) is accessible.

## Diagnose the ERROR State

### Step 1: Identify Affected VMs

List all VMs currently in ERROR state. Use the <code class="expression">space.vars.product\_name</code> UI or the CLI:

```bash
pcdctl server list --status ERROR
```

For each ERROR-state VM, retrieve its details and fault message:

```bash
pcdctl server show <VM_UUID>
```

Look for the `fault` field in the output. Common fault messages and their meanings:

| Fault message                                             | Likely cause                                                           |
| --------------------------------------------------------- | ---------------------------------------------------------------------- |
| `No valid host was found`                                 | Evacuation found no suitable destination host                          |
| `Build of instance ... aborted: Instance failed to spawn` | Compute Service on the host could not start the VM                     |
| `Connection to libvirt failed`                            | `libvirtd` was not running when the Compute Service tried to reconcile |
| `Exceeded maximum number of retries`                      | Repeated provisioning attempts all failed                              |

### Step 2: Check the Host Is Healthy

Before touching the VM, confirm the host is fully recovered:

```bash
pcdctl compute service list
```

The host's `nova-compute` entry should show **state: up** and **status: enabled**. If it shows **state: down**, the Compute Service on that host has not recovered yet. Address the host first — see [Troubleshoot libvirt and Compute Service Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-libvirt-and-compute-service) and [Troubleshooting Offline or Failed Hosts](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-offline-or-failed-hosts).

### Step 3: Review the Compute Service Log on the Host

Log in to the hypervisor and inspect the Compute Service log for errors around the time of the reboot:

```bash
sudo grep -i "error\|exception\|fail" /var/log/pf9/ostackhost.log | tail -100
```

Look for libvirt errors, storage attachment failures, or reconciliation failures referencing the VM's UUID.

## Recovery Procedures

Choose the appropriate procedure based on the fault and your storage configuration.

### Procedure A: Hard Reboot (VM Disk on Shared Storage or Block Volume)

A hard reboot instructs the Compute Service to reset and restart the VM in place on the same host. Use this when:

* The host is healthy and the Compute Service is up.
* The VM's root disk is on ephemeral shared storage or a block storage volume (not ephemeral local storage).
* The fault indicates a transient failure (libvirt timeout, failed reconciliation) rather than a missing resource.

1. From the <code class="expression">space.vars.product\_name</code> UI, navigate to **Virtual Machines**, select the VM, and choose **Hard Reboot** from the **Actions** menu.

   Alternatively, from the CLI:

   ```bash
   pcdctl server reboot --hard <VM_UUID>
   ```
2. Wait for the VM to transition to **ACTIVE**. This typically takes one to three minutes.
3. If the VM returns to ERROR after the hard reboot, proceed to Procedure B or C.

### Procedure B: Reset VM State and Retry (When the VM Is Stuck Transitioning)

If the VM is stuck in a transitioning task state (for example, `rebooting`, `rebuilding`, or `powering-off`) and does not complete, you can reset the state to **ERROR** explicitly and then attempt recovery:

{% hint style="info" %}
**Self-Hosted deployments only**

The `reset-state` command requires access to the region management plane. In SaaS deployments, contact Platform9 Support to perform this operation on your behalf.
{% endhint %}

```bash
pcdctl server reset-state <VM_UUID>
```

After the state is reset to ERROR, attempt a hard reboot (Procedure A).

### Procedure C: Evacuate to Another Host

Use evacuation when:

* The host that originally ran the VM is still offline or unhealthy.
* The VM's root disk is on a block storage volume or ephemeral shared storage (evacuation requires the disk to be accessible from a new host).
* A hard reboot has failed and the host cannot be recovered quickly.

{% hint style="warning" %}
**Evacuation is not possible for VMs using ephemeral local storage.** If the VM was using local (non-shared) ephemeral storage and the host's disk is unavailable, the VM data may be unrecoverable. Contact Platform9 Support for guidance.
{% endhint %}

1. Confirm the VM's storage type. In the <code class="expression">space.vars.product\_name</code> UI, look at the VM's volume attachments. If the root disk is a named volume, evacuation will work. If the root disk shows as ephemeral and the VM's cluster does not use shared storage, evacuation will not work.
2. Run the evacuation command, targeting a specific healthy host:

   ```bash
   pcdctl server evacuate --host <DESTINATION_HOST_UUID> <VM_UUID>
   ```

   Or, to evacuate all ERROR-state VMs from a specific host (for example, after a partial failure):

   ```bash
   pcdctl server evacuate --host <DESTINATION_HOST_UUID> --on-shared-storage
   ```

   See [Virtual Machine Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration) for full evacuation prerequisites.
3. Monitor the VM status until it reaches **ACTIVE**.

### Procedure D: Rebuild from Recovery (Last Resort)

If the VM cannot be hard rebooted or evacuated and the disk is accessible, a rebuild re-creates the VM on its existing volume using the original image. This replaces the VM's in-memory and CPU state but preserves attached volumes.

```bash
pcdctl server rebuild <VM_UUID> --image <IMAGE_UUID>
```

{% hint style="warning" %}
Rebuild rewrites the root disk from the named image. Any data written to the ephemeral root disk after VM creation will be lost. Only use rebuild if the root disk is a block volume and you have confirmed its data integrity, or if the VM is stateless and re-initialization from the base image is acceptable.
{% endhint %}

## After Recovery: Verify VM Health

After the VM returns to ACTIVE:

1. Confirm the VM is reachable on the network (SSH or ICMP ping).
2. Check guest OS logs for application errors that may have resulted from the abrupt shutdown.
3. If the host was rebooted for patching and the Compute Service is still disabled on that host, re-enable it:

   ```bash
   pcdctl compute service enable nova-compute <HOST_FQDN>
   ```
4. Review whether [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) should be used for future planned maintenance to avoid ERROR-state VMs.

## Prevent ERROR States During Planned Maintenance

The most reliable way to avoid ERROR-state VMs during host reboots is to use Maintenance Mode:

1. Enable Maintenance Mode on the host before rebooting. The Compute Service will live-migrate all running VMs to other hosts first.
2. Reboot or patch the host.
3. Verify the host is healthy after it comes back online.
4. Disable Maintenance Mode to allow new VMs to schedule on the host again.

See [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) for the full procedure.

## Related Pages

* [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode)
* [Virtual Machine Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration)
* [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha)
* [Recover libvirt and Compute Service Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-libvirt-and-compute-service)
* [Diagnose VM Scheduling Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/diagnose-vm-scheduling-failures)


# Recover libvirt and Compute Service Failures

## Overview

<code class="expression">space.vars.product\_name</code> relies on `libvirtd` (the libvirt daemon) to manage virtual machine lifecycle on each hypervisor host. If `libvirtd` becomes unresponsive or fails, the Platform9 Compute Service (`pf9-ostackhost`) cannot communicate with the hypervisor, causing hosts to appear offline and VMs to become unreachable.

This guide covers how to diagnose `libvirtd` failures, when and how to safely restart the affected services, and how to confirm that the host and its VMs have fully recovered.

In this guide, you will restore a hypervisor host that has a failing or unresponsive `libvirtd` or Compute Service to a healthy, operational state.

## Prerequisites

* SSH access to the affected hypervisor host.
* Access to the <code class="expression">space.vars.product\_name</code> UI or `pcdctl` CLI to monitor host and VM state.
* Sufficient free resources on other cluster hosts if VM migration is needed during recovery.

## Symptoms

You may be experiencing a `libvirtd` or Compute Service failure if:

* The <code class="expression">space.vars.product\_name</code> Service Health dashboard shows the host as **Offline** or the Compute Service as **Unhealthy**.
* Running `virsh list` on the host hangs or returns an error such as `error: failed to connect to the hypervisor` or `error: unable to connect to server at 'qemu:///system'`.
* VMs show as **ERROR** or **Unknown** in the UI, even though the host itself is reachable over SSH.
* The Compute Service log (`/var/log/pf9/ostackhost.log`) contains repeated lines like `libvirt connection refused`, `Connection reset by peer`, or `Timeout waiting for response`.

## Diagnose the Failure

### Step 1: Check libvirtd Status

Log in to the hypervisor host and check whether `libvirtd` is running:

```bash
sudo systemctl status libvirtd
```

* **Active (running):** The daemon is up. The issue may be a socket or permissions problem rather than a crash. Check the libvirt log (Step 3).
* **Failed or inactive:** `libvirtd` has crashed or was stopped. Proceed to the restart procedure.
* **Activating (start):** The daemon is stuck starting. Check for a hung process (Step 2).

### Step 2: Check for Hung Processes

If `libvirtd` appears to be running but `virsh list` hangs, a qemu process or socket may be blocking the daemon:

```bash
# Check if virsh can connect at all (run with a short timeout)
timeout 10 virsh list 2>&1 || echo "virsh timed out or failed"

# Check for zombie or D-state (uninterruptible) qemu processes
ps aux | grep -E "(qemu|libvirt)" | grep -v grep
```

A large number of `D` (uninterruptible sleep) processes under `qemu-system-x86_64` usually indicates a storage I/O hang. Resolve the underlying storage issue before restarting `libvirtd`, or the daemon will hang again.

### Step 3: Review libvirt Logs

Examine the libvirt daemon log for errors leading up to the failure:

```bash
sudo tail -200 /var/log/libvirt/libvirtd.log | grep -i "error\|crit\|fail"
```

For per-VM issues, check the QEMU log for the affected VM:

```bash
sudo tail -100 /var/log/libvirt/qemu/<VM_UUID>.log
```

### Step 4: Check the Compute Service Status

```bash
sudo systemctl status pf9-ostackhost
```

Also check the Compute Service log:

```bash
sudo grep -i "error\|exception\|libvirt" /var/log/pf9/ostackhost.log | tail -100
```

## Restart Procedure

{% hint style="warning" %}
Restarting `libvirtd` will cause a brief interruption to running VMs on the host while the connection is re-established. Running VMs themselves are not terminated by a `libvirtd` restart, but they will be temporarily unresponsive. If your workloads cannot tolerate any interruption, use [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode) to migrate VMs off the host first.
{% endhint %}

Restart services in the following order. Wait for each service to reach `active (running)` before starting the next.

### Step 1: Restart libvirtd

```bash
sudo systemctl restart libvirtd
sudo systemctl status libvirtd
```

After the restart, confirm that `virsh` can connect:

```bash
virsh list
```

This should return without hanging and display any VMs currently defined on the host (running or shut off).

### Step 2: Restart the Platform9 Compute Service

Once `libvirtd` is healthy, restart the Platform9 Compute Service:

```bash
sudo systemctl restart pf9-ostackhost
sudo systemctl status pf9-ostackhost
```

### Step 3: Restart the Platform9 Host Agent (If the Host Is Still Showing Offline)

If the host is still shown as **Offline** in the <code class="expression">space.vars.product\_name</code> UI after the Compute Service restart, restart the host agent as well:

```bash
sudo systemctl restart pf9-hostagent
sudo systemctl status pf9-hostagent
```

The host agent is responsible for reporting host health to the management plane. After restarting it, wait two to three minutes and then check the UI.

## Validate Recovery

### Confirm Host Is Online

In the <code class="expression">space.vars.product\_name</code> UI, navigate to **Infrastructure > Cluster Hosts** and verify the host status returns to **Active**.

Using the CLI:

```bash
pcdctl compute service list
```

The `nova-compute` entry for the host should show **state: up** and **status: enabled**.

### Confirm VMs Have Reconciled

After the Compute Service restarts, it re-queries `libvirtd` to reconcile the state of all VMs on the host. Most VMs will transition from **ERROR** or **Unknown** back to **ACTIVE** automatically within two to three minutes.

```bash
pcdctl server list --host <HOST_UUID>
```

Any VMs that remain in **ERROR** after five minutes may require additional recovery steps. See [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state).

### Confirm virsh Reflects Running VMs

```bash
virsh list --all
```

VMs that are **ACTIVE** in <code class="expression">space.vars.product\_name</code> should appear as `running` here. If a VM is **ACTIVE** in the UI but `shut off` in `virsh list`, restart the VM using a hard reboot:

```bash
pcdctl server reboot --hard <VM_UUID>
```

## Additional Checks

### Verify All Platform9 Services Are Running

The Platform9 stack requires multiple services to be healthy on each hypervisor:

```bash
sudo systemctl status pf9-hostagent
sudo systemctl status pf9-ostackhost
sudo systemctl status pf9-comms
sudo systemctl status pf9-sidekick
```

If any service is in a `failed` state, restart it individually:

```bash
sudo systemctl restart <service-name>
```

### Check for Disk Space Issues

A full disk is a common cause of `libvirtd` failures and Compute Service crashes:

```bash
df -h /var/log
df -h /var/lib/libvirt
```

If disk usage is above 90%, free space before restarting services.

### Check for Socket File Issues

If `libvirtd` is running but `virsh` cannot connect, the UNIX socket may be in a bad state:

```bash
ls -la /var/run/libvirt/libvirt-sock
```

Restarting `libvirtd` normally recreates the socket. If the socket file persists after a restart and `virsh` still cannot connect, remove it manually and restart:

```bash
sudo rm /var/run/libvirt/libvirt-sock
sudo systemctl restart libvirtd
```

## When to Contact Support

Escalate to Platform9 Support if:

* `libvirtd` crashes immediately after each restart (check for recurring errors in `/var/log/libvirt/libvirtd.log`).
* VMs do not reconcile to **ACTIVE** after the host and Compute Service are healthy.
* You see storage I/O errors in the libvirt or QEMU logs that indicate underlying storage hardware issues.
* Multiple hosts in the cluster are affected simultaneously.

Contact [Platform9 Support](https://support.platform9.com/) and provide the libvirt log, the Compute Service log, and the output of `pcdctl compute service list`.

## Related Pages

* [Troubleshooting Offline or Failed Hosts](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-offline-or-failed-hosts)
* [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state)
* [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode)
* [Compute Service Advanced Configuration](/private-cloud-director/virtualized-clusters/nova-override)


# Diagnose VM Scheduling Failures

## Overview

When you attempt to create or resize a VM and the Compute Service cannot find a suitable host, the VM lands in **ERROR** state with the fault message **"No valid host was found"**. This is one of the most common VM creation failure modes in <code class="expression">space.vars.product\_name</code>.

This guide walks you through diagnosing and resolving scheduling failures by checking the factors the Compute Service evaluates when placing a VM: available resources (vCPU, RAM, disk), host aggregate membership, host state, and the Placement service.

In this guide, you will identify why a VM is failing to schedule and take the corrective action needed to get it placed.

## Prerequisites

* Access to the <code class="expression">space.vars.product\_name</code> UI or `pcdctl` CLI.
* For Self-Hosted deployments: access to `kubectl` against the region namespace (for pod log inspection).

## Understand the Placement Decision

Before checking individual factors, it helps to know what the Compute Service evaluates when scheduling a VM:

1. **Resource availability** — does any host have enough free vCPU, RAM, and disk to satisfy the flavor?
2. **Host state** — is the `nova-compute` service enabled and up on the host?
3. **Host aggregate membership** — if the flavor specifies aggregate metadata, does the host belong to a matching aggregate?
4. **Image properties** — does the image require specific host capabilities (for example, vTPM support)?

If any filter eliminates all hosts, scheduling fails with "No valid host was found."

## Step 1: Confirm the Fault Message

Retrieve the failed VM's fault detail:

```bash
pcdctl server show <VM_UUID>
```

Look for the `fault` field. A scheduling failure typically reads:

```
No valid host was found. There are not enough hosts available.
```

or

```
No valid host was found. Exceeded max scheduling attempts...
```

Also retrieve the server events to find the request ID:

```bash
pcdctl server event list <VM_UUID>
pcdctl server event show <VM_UUID> <REQUEST_ID>
```

The event detail often contains more specific information about which filter rejected all hosts.

## Step 2: Check Compute Service State on All Hosts

A disabled or down `nova-compute` service on a host makes that host ineligible for new VMs.

```bash
pcdctl compute service list
```

Review the output. For each host in your cluster, check:

* **state**: must be `up`. A `down` state means the Compute Service is not reporting to the management plane.
* **status**: must be `enabled`. A `disabled` status means the host was manually excluded from scheduling (for example, left disabled after a maintenance window).

**To re-enable a disabled host:**

```bash
pcdctl compute service enable nova-compute <HOST_FQDN>
```

If a host shows `state: down`, the Compute Service on that host is not running. See [Recover libvirt and Compute Service Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-libvirt-and-compute-service).

## Step 3: Check Available Resources Against the Flavor

Determine what the VM's flavor requires:

```bash
pcdctl flavor show <FLAVOR_NAME_OR_ID>
```

Note the `vcpus`, `ram` (in MB), and `disk` (in GB) values.

Then check the aggregate resource availability across hosts:

```bash
pcdctl hypervisor stats show
```

This shows total and free vCPUs, RAM, and disk across all enabled hypervisors in the region. If free resources are below what the flavor requires, no host can accept the VM.

For per-host breakdown:

```bash
pcdctl hypervisor list --long
```

Compare the `Free RAM MB`, `Free Disk GB`, and `Running VMs` columns against the flavor requirements. A host is only eligible if it has enough free RAM and disk, and enough vCPU headroom.

**If resources are fully consumed:** either add hosts to the cluster, resize the flavor, or shut down unused VMs to free resources.

## Step 4: Check Host Aggregate Membership

If the flavor includes host aggregate metadata (such as `aggregate_instance_extra_specs:<key>=<value>`), the VM can only schedule on hosts that belong to a matching aggregate.

Check the flavor's extra specs:

```bash
pcdctl flavor show <FLAVOR_NAME_OR_ID>
```

Look for `properties` entries starting with `aggregate_instance_extra_specs:`.

Then verify that at least one healthy host belongs to the required aggregate:

```bash
pcdctl aggregate list
pcdctl aggregate show <AGGREGATE_NAME>
```

The `hosts` field lists which hosts are members. If no healthy host is a member of the aggregate the flavor targets, scheduling fails.

**To add a host to an aggregate:**

Navigate to **Infrastructure > Host Aggregates** in the <code class="expression">space.vars.product\_name</code> UI and add the host. Alternatively, use the CLI:

```bash
pcdctl aggregate add host <AGGREGATE_NAME> <HOST_FQDN>
```

See [Host Aggregate](/private-cloud-director/virtualized-clusters/host-aggregate) for more information on configuring aggregates and flavor-level targeting.

## Step 5: Check Image Properties

Certain image properties require specific host capabilities and restrict which hosts the VM can schedule on. Common examples:

| Image property               | Requirement                                            |
| ---------------------------- | ------------------------------------------------------ |
| `hw_tpm_version = 2.0`       | Host must have vTPM (`swtpm`) installed and configured |
| `hw:numa_nodes`              | Host must have multiple NUMA nodes                     |
| Aggregate isolation metadata | Host must belong to a specific aggregate               |

Check the image's properties:

```bash
pcdctl image show <IMAGE_ID>
```

If the image requires a capability that only some hosts have (for example, vTPM), confirm that those hosts are healthy, enabled, and have sufficient resources.

## Step 6: Inspect Placement Service Logs (Self-Hosted Deployments Only)

{% hint style="info" %}
**Self-Hosted deployments only**

The following steps require `kubectl` access to the region namespace in the management plane. In SaaS deployments, contact Platform9 Support to retrieve Placement and Compute Service logs for your region.
{% endhint %}

When the previous steps do not reveal an obvious cause, inspect the Placement service and Compute Service logs in the management plane to see exactly which filters rejected all hosts.

#### Check the Compute Service (nova-scheduler) Logs

Retrieve the request ID from the VM events (Step 1), then search for it in the nova-scheduler logs:

```bash
kubectl logs deployment/nova-scheduler -n <WORKLOAD_REGION> | grep <REQUEST_ID>
```

Look for lines containing `FilterScheduler` or the names of individual filters such as `RamFilter`, `DiskFilter`, `ComputeFilter`, or `AggregateInstanceExtraSpecsFilter`. These lines indicate which filter rejected which hosts.

#### Check the Placement API Logs

```bash
kubectl logs deployment/placement-api -n <WORKLOAD_REGION> | grep <REQUEST_ID>
```

A `GET /resource_providers` call with an empty result set (zero providers returned) means the Placement service found no resource provider (host) with enough inventory to satisfy the VM's resource requirements. This confirms a resource exhaustion problem rather than a filter mismatch.

#### Check nova-conductor for Quota or Policy Rejections

```bash
kubectl logs deployment/nova-conductor -n <WORKLOAD_REGION> | grep <REQUEST_ID>
```

Errors in `nova-conductor` may indicate tenant quota limits rather than host resource limits. Check the tenant's compute quotas:

```bash
pcdctl quota show --project <PROJECT_ID>
```

Compare the `instances`, `cores`, and `ram` used versus allowed values.

## Common Scenarios and Fixes

| Symptom                                           | Likely cause                                 | Fix                                                      |
| ------------------------------------------------- | -------------------------------------------- | -------------------------------------------------------- |
| All hosts show `state: down`                      | Compute Service not running on any host      | Restart `pf9-ostackhost` on affected hosts               |
| Some hosts `status: disabled`                     | Left disabled after maintenance              | Run `pcdctl compute service enable nova-compute <HOST>`  |
| Resources show 0 free RAM/disk                    | Cluster fully consumed                       | Add hosts, or free resources by stopping unused VMs      |
| Flavor has aggregate metadata but no host matches | Aggregate misconfiguration or host not added | Add a healthy host to the required aggregate             |
| Image requires vTPM but no vTPM hosts available   | No eligible hosts                            | Verify vTPM-capable hosts are in the cluster and healthy |
| Quota exceeded                                    | Tenant has hit its resource quota            | Raise the quota or delete unused resources               |

## Related Pages

* [VM Flavors](/private-cloud-director/virtualized-clusters/virtualmachine/vm-flavors)
* [Host Aggregate](/private-cloud-director/virtualized-clusters/host-aggregate)
* [Advanced Scheduling Options](/private-cloud-director/virtualized-clusters/virtualmachine/advance-vm-scheduling-options)
* [Failed to Deploy Virtual Machine](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/failed-to-deploy-virtual-machine)
* [Recover libvirt and Compute Service Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-libvirt-and-compute-service)
* [Virtual TPM](/private-cloud-director/virtualized-clusters/virtual-tpm)
* [Tenant Quotas, User Quotas and VM Leases](/private-cloud-director/identity-and-multi-tenancy/tenant-quotas-user-quotas-and-vm-leases)


# Recover from Messaging Layer Failures

## Overview

<code class="expression">space.vars.product\_name</code>'s Compute Service components communicate with each other through a messaging layer. When the messaging layer becomes unhealthy, VM create requests appear to hang or fail silently — the VM stays in **BUILD** state indefinitely, or a burst of VM creates causes a large number of failures at scale.

This guide explains how to recognize a messaging layer problem, what symptoms distinguish it from other failure modes, and how to recover. Management-plane remediation steps (such as restarting messaging layer pods) apply only to Self-Hosted deployments; SaaS customers should contact Platform9 Support.

In this guide, you will identify a messaging layer failure and restore normal VM creation behavior.

## Prerequisites

* Access to the <code class="expression">space.vars.product\_name</code> UI or `pcdctl` CLI.
* For Self-Hosted deployments: `kubectl` access to the region namespace.

## Recognize a Messaging Layer Failure

Messaging layer failures produce a characteristic pattern that distinguishes them from resource exhaustion or host failures:

| Symptom                                      | Messaging layer failure         | Resource exhaustion             | Host failure                       |
| -------------------------------------------- | ------------------------------- | ------------------------------- | ---------------------------------- |
| VMs stuck in BUILD indefinitely              | Yes                             | No — fails quickly              | Partial — depends on timing        |
| Error message                                | None, or generic timeout        | "No valid host was found"       | "Failed to spawn" or libvirt error |
| Affects all new VM creates simultaneously    | Yes (all fail at the same time) | Yes (all fail at the same time) | No (only VMs on the affected host) |
| Existing running VMs affected                | No                              | No                              | Yes (VMs on failed host)           |
| `pcdctl compute service list` shows hosts up | Yes                             | Yes                             | No (affected host shows down)      |

If you see VMs stuck in **BUILD** for more than ten minutes with no fault message, and `pcdctl compute service list` shows all hosts as `state: up`, a messaging layer issue is the most likely cause.

## Diagnose the Failure

### Step 1: Confirm VMs Are Stuck in BUILD

```bash
pcdctl server list --status BUILD
```

If this returns a large number of VMs, or if VMs have been in BUILD for an unusually long time (more than ten minutes for a standard VM), proceed to the next step.

### Step 2: Check for Timeout Errors in the Compute Service Log

On an affected hypervisor host, inspect the Compute Service log:

```bash
sudo grep -i "timeout\|connection refused\|AMQP\|rabbit\|MessagingTimeout" \
    /var/log/pf9/ostackhost.log | tail -50
```

Errors referencing `MessagingTimeout`, `AMQP`, or `rabbit` confirm that the Compute Service cannot reach the messaging layer.

Also check the `nova-conductor` or `nova-scheduler` pod logs for messaging errors (Self-Hosted only):

{% hint style="info" %}
**Self-Hosted deployments only**

The following steps require `kubectl` access to the region namespace. In SaaS deployments, contact Platform9 Support to inspect management-plane component health and perform any pod-level remediation.
{% endhint %}

```bash
kubectl logs deployment/nova-conductor -n <WORKLOAD_REGION> | \
    grep -i "timeout\|AMQP\|rabbit\|MessagingTimeout" | tail -50

kubectl logs deployment/nova-scheduler -n <WORKLOAD_REGION> | \
    grep -i "timeout\|AMQP\|rabbit\|MessagingTimeout" | tail -50
```

### Step 3: Check Messaging Layer Pod Health (Self-Hosted Only)

Check whether the messaging layer pods in the region namespace are healthy:

```bash
kubectl get pods -n <WORKLOAD_REGION> | grep -i rabbit
```

Look for pods in `CrashLoopBackOff`, `OOMKilled`, `Pending`, or `Error` state. A `CrashLoopBackOff` messaging layer pod is a strong indicator of the root cause.

Describe the pod to see recent events:

```bash
kubectl describe pod <RABBITMQ_POD_NAME> -n <WORKLOAD_REGION>
```

Check the pod logs:

```bash
kubectl logs <RABBITMQ_POD_NAME> -n <WORKLOAD_REGION> --tail=100
```

Look for memory pressure events, disk quota errors, or authentication failures.

## Recovery Procedure

### For Self-Hosted Deployments

If the messaging layer pod is in a crash state, attempt a pod restart:

```bash
kubectl rollout restart deployment/<RABBITMQ_DEPLOYMENT_NAME> -n <WORKLOAD_REGION>
```

Wait for the pod to return to `Running` state:

```bash
kubectl rollout status deployment/<RABBITMQ_DEPLOYMENT_NAME> -n <WORKLOAD_REGION>
```

After the messaging layer pod is healthy, the Compute Service components reconnect automatically. Allow two to three minutes for reconnection, then verify that new VM creates succeed.

If restarting the pod does not resolve the issue, or if the pod crashes again immediately, escalate to Platform9 Support. The underlying cause may be resource exhaustion (memory or disk), a misconfiguration, or a persistent authentication issue that requires deeper investigation.

### For SaaS Deployments

You do not have access to manage messaging layer infrastructure in a SaaS deployment. Contact [Platform9 Support](https://support.platform9.com/) and provide:

* The time range when VM creates began failing.
* The output of `pcdctl server list --status BUILD`.
* The output of `pcdctl compute service list`.
* The relevant lines from `/var/log/pf9/ostackhost.log` on an affected hypervisor.

### After Recovery: Clean Up Stuck VMs

After the messaging layer is healthy, VMs that were stuck in BUILD may not automatically recover. Check their status:

```bash
pcdctl server list --status BUILD
```

For each stuck VM, attempt to reset its state and delete it, then re-create:

```bash
pcdctl server delete <VM_UUID>
```

If the delete also hangs, reset the VM state first (Self-Hosted only, requires management-plane access — contact Platform9 Support for SaaS):

```bash
pcdctl server reset-state <VM_UUID>
pcdctl server delete <VM_UUID>
```

## Prevent Messaging Layer Failures at Scale

When creating a large number of VMs in a burst, the following practices reduce the risk of overwhelming the messaging layer:

* **Stage VM creation in batches.** Rather than creating hundreds of VMs simultaneously, create them in groups of 20–50 and wait for each batch to reach **ACTIVE** before proceeding.
* **Monitor VM creation rate against available messaging layer capacity.** In Self-Hosted deployments, review messaging layer resource allocation (CPU, memory) in the region namespace and increase limits if burst workloads regularly hit capacity.
* **Use tenant quotas to limit simultaneous creation.** See [Tenant Quotas, User Quotas and VM Leases](/private-cloud-director/identity-and-multi-tenancy/tenant-quotas-user-quotas-and-vm-leases) to set per-tenant instance and core limits.

## Related Pages

* [Failed to Deploy Virtual Machine](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/failed-to-deploy-virtual-machine)
* [Diagnose VM Scheduling Failures](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/diagnose-vm-scheduling-failures)
* [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state)
* [Tenant Quotas, User Quotas and VM Leases](/private-cloud-director/identity-and-multi-tenancy/tenant-quotas-user-quotas-and-vm-leases)


# Diagnose a Host Agent Stuck in Converging State

## Overview

When a host is authorized and roles are assigned, <code class="expression">space.vars.product\_acronym</code> downloads and applies service packages through the host agent. During this process, the host shows a `converging` status in the UI. Under normal conditions, convergence completes within a few minutes. If a host remains in `converging` state for more than 10–15 minutes, the host agent has likely encountered an error it cannot recover from on its own.

In this guide, you will locate the relevant log files, identify the most common causes of a stuck convergence, and restore the host to `applied` status.

## Check the Hostagent Log First

The hostagent log is the primary source of information when a host is stuck converging. SSH to the affected host and inspect the log:

```bash
tail -100 /var/log/pf9/hostagent.log
```

Or follow it live while monitoring:

```bash
tail -f /var/log/pf9/hostagent.log
```

Look for lines containing `ERROR`, `WARN`, `failed`, `timeout`, or `permission denied`. The section below maps the most common messages to their fixes.

## Common Causes and Fixes

### Service Start Failure

**Symptom in log:** Lines mentioning a specific `pf9-*` service (`pf9-ostackhost`, `pf9-neutron-ovn-controller`, etc.) failing to start or timing out.

The host agent applies roles by starting Platform9 services. If a service fails to start, the agent retries but eventually marks the convergence as failed.

**What to check:**

```bash
# Replace pf9-ostackhost with the service named in the log
sudo systemctl status pf9-ostackhost
sudo journalctl -u pf9-ostackhost --since "30 minutes ago" | tail -80
```

Common sub-causes:

* **Port conflict:** Another process is already bound to a port the service needs. Check with `sudo ss -tlnp`.
* **Missing dependency:** A required package was not installed. Check `/var/log/dpkg.log` and `/var/log/apt/history.log` for installation failures.
* **Configuration error written by a prior failed attempt:** Remove the stale role-state files and retry. See [Re-Onboarding Recovery](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-host-onboarding-issues#re-onboarding-recovery).

After addressing the root cause, restart the host agent to trigger re-convergence:

```bash
sudo systemctl restart pf9-hostagent
```

### Log-Rotation Permission Error Filling Up the Agent

**Symptom in log:** Messages about failing to write to `/var/log/pf9/hostagent.log`, or the log file stops updating while the host remains in `converging` state in the UI.

**Symptom on host:**

```bash
df -h /var/log
```

If `/var/log` is at or near 100%, the hostagent cannot write logs and may be stuck waiting for I/O.

**Fix:**

1. Free up log space:

```bash
sudo journalctl --vacuum-size=1G
sudo find /var/log/pf9 -name "*.log.*" -mtime +7 -delete
```

2. Check and fix log file permissions. The hostagent log must be writable by the `pf9` user or root:

```bash
ls -la /var/log/pf9/hostagent.log
sudo chown root:root /var/log/pf9/hostagent.log
sudo chmod 644 /var/log/pf9/hostagent.log
```

3. Restart the hostagent:

```bash
sudo systemctl restart pf9-hostagent
```

4. Monitor the log to confirm convergence resumes:

```bash
tail -f /var/log/pf9/hostagent.log
```

### Package Download or Installation Failure

**Symptom in log:** References to `pf9apps`, `apt-get`, `dpkg`, or HTTP errors when fetching packages from the management plane.

The hostagent downloads `.deb` packages into `/var/cache/pf9apps/` and installs them during convergence. Download failures cause the agent to retry until it gives up.

**What to check:**

```bash
# Confirm the management plane is reachable from the host
curl -sk https://<management-plane-fqdn>/resmgr/v1/hosts | head -c 200

# Check for stale or corrupted package files
ls -la /var/cache/pf9apps/
```

**Fix:**

1. Clear the package cache and let the agent re-download:

```bash
sudo rm -rf /var/cache/pf9apps/*
sudo systemctl restart pf9-hostagent
```

2. If the management plane is unreachable, resolve the network or proxy issue first. See [Troubleshooting Offline or Failed Hosts](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-offline-or-failed-hosts).

### Hostagent Crash Loop

**Symptom:** `pf9-hostagent` service shows `failed` or repeatedly restarts in `systemctl status pf9-hostagent`.

**What to check:**

```bash
sudo systemctl status pf9-hostagent
sudo journalctl -u pf9-hostagent --since "1 hour ago" | tail -100
```

Look for a panic or fatal error message. If the agent is crashing on startup, there may be a corrupted state file.

**Fix:**

```bash
# Stop the agent
sudo systemctl stop pf9-hostagent

# Remove state files that may be corrupted
sudo rm -f /opt/pf9/data/state/*

# Restart
sudo systemctl start pf9-hostagent
```

Monitor the log to verify the agent starts cleanly.

## Trigger a Re-Sync from the UI

If the hostagent log shows no ongoing activity and the host remains in `converging` state, trigger a re-sync from the UI:

1. Navigate to **Infrastructure > Cluster Hosts** in the <code class="expression">space.vars.product\_name</code> UI.
2. Select the affected host.
3. Click **Other > Re-sync Host** (if available).

This signals the management plane to re-send the desired role configuration to the host agent.

Alternatively, simply restarting the hostagent on the host achieves the same effect:

```bash
sudo systemctl restart pf9-hostagent
```

## When to Contact Support

If the host does not reach `applied` status within 15 minutes of the restart, collect the following and contact [Platform9 Support](https://support.platform9.com/):

* The full hostagent log: `/var/log/pf9/hostagent.log`
* The output of `sudo systemctl status pf9-hostagent`
* The output of `sudo journalctl -u pf9-hostagent --since "2 hours ago"`
* For Self-Hosted deployments only, the output of `airctl host-status`

{% hint style="info" %}
**Self-Hosted deployments only.** You can also check whether the management plane has a pending convergence task for the host:

```bash
airctl host-status --config /opt/pf9/airctl/conf/airctl-config.yaml
```

Look for hosts where `Status` is not `ok`. In SaaS deployments, contact Platform9 Support if host-level steps do not resolve the convergence failure.
{% endhint %}


# Diagnose Role Assignment Failures

## Overview

After a host is authorized in <code class="expression">space.vars.product\_name</code>, you assign roles (Hypervisor, Networking Service, Image Library Service, Persistent Storage Service) from the UI or API. Role assignment triggers the management plane to push the required configuration to the host agent, which then downloads packages and starts services. When this process fails, the host gets stuck in a `converging` or `error` state rather than reaching `applied`.

The two most common failure patterns are:

1. **HTTP 500 error when assigning or removing a role** — the management plane API returns an internal server error and the role change is not queued.
2. **"Interface None is missing an IP" error** — the Networking Service role fails to apply because the host's network interface configuration is incomplete.

In this guide, you will identify which failure pattern applies, verify the host's prerequisites, and resolve the error.

## Prerequisites Before Assigning Roles

Before assigning any role, confirm the following on the host:

* The host is in `online` connection status and `unauthorized` or `applied` role status in the UI (**Infrastructure > Cluster Hosts**).
* The host agent is running: `sudo systemctl status pf9-hostagent`
* The host can reach the management plane: `curl -sk https://<management-plane-fqdn>/resmgr/v1/hosts | head -c 200`
* NTP is synchronized: `sudo timedatectl status`
* Required network interfaces are present and have IP addresses (see below)

## HTTP 500 Error on Role Assignment or Removal

### Symptom

When you click **Edit Roles** in the UI and save, the UI displays an error or the role assignment never starts. In some cases, the API call to assign or remove a role returns an HTTP 500 response.

### What to Check

**1. Host state in the management plane**

A 500 error on role assignment often means the management plane has stale or conflicting state for the host — for example, a prior role removal that did not complete cleanly, leaving the host's record in an inconsistent state.

Check the hostagent log on the host for any errors that occurred just before the 500 response:

```bash
tail -100 /var/log/pf9/hostagent.log
```

Look for authentication failures, connection errors, or messages about the host not being found in the management plane.

**2. Host registration status**

Confirm the host's UUID in the local file matches what the management plane shows:

```bash
sudo cat /etc/pf9/host_id.conf
```

The UUID in this file must match the host UUID shown in the UI under **Infrastructure > Cluster Hosts**. A mismatch means the host was re-imaged or the identity file was changed without a clean decommission, and the management plane is seeing two conflicting records for the same host.

If there is a mismatch, decommission the stale host record from the management plane and re-onboard the host. See [Re-Onboarding Recovery](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-host-onboarding-issues#re-onboarding-recovery).

**3. Stale compute service entries**

If the host was previously decommissioned forcefully (using `pcdctl decommission-node -r`), stale compute service records may remain in the management plane database and block role re-assignment. Remove them:

```bash
pcdctl compute service list
pcdctl compute service delete <uuid>
```

After removing stale entries, restart the host agent:

```bash
sudo systemctl restart pf9-hostagent
```

Then retry the role assignment from the UI.

**4. Self-Hosted deployments: check the resmgr pod**

{% hint style="info" %}
**Self-Hosted deployments only.** If host-level steps do not resolve the 500 error, check the resource manager service logs on the management plane for the specific error:

```bash
kubectl logs -n <region-fqdn> -l app=resmgr --tail=200 | grep -i "500\|error\|<host-uuid>"
```

In SaaS deployments, contact Platform9 Support if the host-level steps do not resolve the error.
{% endhint %}

### Recovery Steps

1. Deauthorize any partially assigned roles from the UI (select the host, click **Edit Roles**, remove all roles, save).
2. Wait for the host to return to `unauthorized` status.
3. Verify the hostagent is running and the host is `online`.
4. Re-assign the roles from the UI.

If the 500 error persists after these steps, contact [Platform9 Support](https://support.platform9.com/) with the host UUID, the hostagent log, and the time the error occurred.

## "Interface None Is Missing an IP" Error

### Symptom

Role assignment starts but the Networking Service role fails to apply. The hostagent log or the UI role status shows an error message containing `Interface None is missing an IP` or `interface <name> has no IP address`.

### Cause

The Networking Service role requires at least one network interface on the host to have an IP address configured in the host's blueprint. This error means either:

* The interface selected in the host's network configuration (blueprint) has no IP address assigned to it on the host.
* The host blueprint references an interface name that does not exist on the host (for example, the interface was renamed after the blueprint was created).

### What to Check

**1. Verify the host's actual network interfaces and their IP addresses**

```bash
ip addr show
```

Confirm that the interface you intend to use for the management network (or the interface referenced in the cluster blueprint) has an IP address. If an interface shows `state DOWN` or has no `inet` entry, it has no IP.

**2. Verify the interface name matches what the blueprint expects**

Network interface names can change across OS upgrades or hardware changes (for example, `eth0` becoming `ens3`). Check the host configuration stored by the agent:

```bash
sudo cat /opt/pf9/data/state/network_config.json 2>/dev/null || echo "file not found"
```

Compare the interface names in this file against the output of `ip addr show`.

**3. Check the cluster blueprint**

In the UI, navigate to **Infrastructure > Cluster Blueprint** and review the network interface configuration for the cluster the host belongs to. Confirm the interface name and IP assignment match what is present on the host.

### Recovery Steps

1. If the interface has no IP, assign an IP address to it:

```bash
# Verify with
ip addr show <interface-name>
```

For a permanent IP assignment, update the Netplan configuration in `/etc/netplan/` and apply it:

```bash
sudo netplan apply
```

2. If the interface name in the blueprint does not match the actual interface on the host, update the host's network configuration to use the correct interface name. From the UI, navigate to **Infrastructure > Cluster Hosts**, select the host, and update the interface assignment.
3. After the interface has an IP and the blueprint references the correct interface name, retry the role assignment from the UI.
4. Monitor the hostagent log to confirm the Networking Service role applies successfully:

```bash
tail -f /var/log/pf9/hostagent.log
```

Look for `role pf9-neutron applied` or `converge successful`.

## General Role Assignment Checklist

Before retrying a role assignment after any failure, verify:

* [ ] Host shows `online` in **Infrastructure > Cluster Hosts**
* [ ] `sudo systemctl status pf9-hostagent` is `active (running)`
* [ ] `sudo cat /etc/pf9/host_id.conf` UUID matches the UI
* [ ] All network interfaces referenced by the cluster blueprint have IP addresses (`ip addr show`)
* [ ] No stale compute service entries for this host (`pcdctl compute service list`)
* [ ] Sufficient disk space on `/var`, `/opt`, and `/tmp` (`df -h`)

If all items pass and role assignment still fails, contact [Platform9 Support](https://support.platform9.com/) with the hostagent log and the exact error message.


# Troubleshoot Maintenance Mode Migration Failures

## Overview

VM migration failures during maintenance mode are most commonly caused by resource constraints on the destination host, affinity rule conflicts, or CPU model mismatches between the source and destination host. This guide explains how to identify the cause, safely abort or retry a failed migration, and recover any VMs left in an error state.

For how maintenance mode works and how to enable or disable it, see [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode).

In this guide, you will diagnose why migrations failed, resolve the blocking condition, and restore stranded VMs to a running state.

## Identify Which VMs Failed to Migrate

Open the migration progress panel for the host in maintenance mode:

1. Navigate to **Infrastructure > Cluster Hosts** in the <code class="expression">space.vars.product\_name</code> UI.
2. Select the host that is in maintenance mode.
3. Click **See Details** on the maintenance mode banner to open the **View Migration Progress** panel.

The panel lists each VM and its migration status. VMs with a `Failed` status are the ones to investigate.

For each failed VM, note the VM name and check the Compute Service log on the source host for the migration error:

```bash
grep -A 5 "<vm-name-or-uuid>" /var/log/pf9/ostackhost.log | tail -40
```

## Common Causes of Migration Failure

### Insufficient Resources on Destination Hosts

If no destination host has enough free CPU or memory to accept the VM, the migration fails with a "No valid host was found" error.

**What to check:**

```bash
pcdctl host list
```

Review the `vCPUs used` and `RAM used` columns for each host in the cluster. If every destination host is near capacity, you must either free up resources (shut down idle VMs) or add another host to the cluster before retrying maintenance mode.

### CPU Model Mismatch

Live migration requires that the source and destination host expose compatible CPU models to the VM. If a host in the cluster was recently upgraded and its effective CPU model differs from the source host, live migration will fail.

**What to check:**

```bash
# Run on both the source and destination host
virsh domcapabilities | grep "model usable='yes'" | sort
```

Compare the lists. The selected CPU model for the cluster (visible in `nova_override.conf` under `[libvirt] cpu_models`) must appear as usable on both hosts. See [Resolve CPU Baseline Mismatch After Host Upgrade](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/cpu-baseline-mismatch) for steps to correct a mismatch.

### Affinity or Anti-Affinity Rule Conflicts

Maintenance mode honors hard affinity and anti-affinity rules. If a VM has a hard affinity rule requiring it to be co-located with another VM that is also being migrated, and no destination host can satisfy both, the migration fails.

Review the VM's affinity group in the UI: navigate to **Compute > VM Affinity Anti-Affinity Rules** and confirm which group the VM belongs to. If the conflict cannot be resolved automatically, you may need to temporarily remove the hard rule, migrate the VM manually, and re-apply the rule.

### VMs in Error or Unmigratable States

As described in [VM States](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode#vm-states) in the Maintenance Mode guide, VMs in `Error`, `Suspended`, `Shutdown`, `Rescued`, or `Pending resize confirmation` states are skipped by maintenance mode. VMs in `Error` state must be recovered first.

To recover a VM in `Error` state, attempt a hard reboot:

```bash
pcdctl server reboot --hard <vm-id>
```

If the hard reboot does not resolve the error, see [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state).

## Abort Maintenance Mode and Retry

If maintenance mode is partially complete and you want to stop it, disable maintenance mode from the UI:

1. Navigate to **Infrastructure > Cluster Hosts**.
2. Select the host in maintenance mode.
3. Click **Other > Disable Maintenance Mode** or use the **Disable Maintenance Mode** button on the Host Details page.

Disabling maintenance mode marks the host as schedulable again. VMs that were already successfully migrated remain on their destination hosts; they are not migrated back automatically.

After resolving the blocking condition (freeing resources, fixing CPU model, recovering error-state VMs), re-enable maintenance mode to migrate the remaining VMs.

## Manually Migrate a Stranded VM

If a specific VM cannot be migrated by maintenance mode (for example, it has a Virtual TPM or an unresolvable affinity constraint), migrate it manually before enabling maintenance mode:

```bash
pcdctl server migrate <vm-id>
```

To target a specific destination host:

```bash
pcdctl server migrate --host <destination-host-name> <vm-id>
```

After the manual migration completes, the VM is no longer on the source host and maintenance mode can proceed without encountering it.

## Recover a VM Left in Error State After Migration

If a VM ended up in `Error` state during maintenance mode migration, recover it with a hard reboot:

```bash
pcdctl server reboot --hard <vm-id>
```

If the VM is reporting its hypervisor host as the host in maintenance mode (the source), the VM's record in the management plane may need to be reset. Contact [Platform9 Support](https://support.platform9.com/) with the VM UUID and the Compute Service log from the source host.

For a full procedure covering VMs that remain in `Error` after host maintenance events, see [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state).


# Resolve CPU Baseline Mismatch After Host Upgrade

## Overview

All hypervisor hosts in a virtualized cluster must expose the same CPU model to their VMs. The Compute Service selects the cluster's CPU baseline automatically when each host is first authorized, picking the newest approved model that every host in the cluster can report as `usable`. When a host is upgraded — especially when moving from one CPU generation to another, or when the libvirt version changes — the set of flags the host can advertise may change. If the upgraded host can no longer satisfy the flags required by the cluster's established CPU baseline, libvirt rejects the host and the Compute Service reports an error.

Common symptoms of a CPU baseline mismatch:

* A host that completed a <code class="expression">space.vars.product\_acronym</code> agent upgrade transitions to `error` or `offline` status even though the host itself is reachable.
* `/var/log/pf9/hostagent.log` or `/var/log/pf9/ostackhost.log` contains an error similar to `CPU model <model> is not supported on this host` or `unsupported configuration: CPU flag <flag> not found`.
* VM live migrations or DRR rebalancing fail with CPU compatibility errors after adding a new-generation host to a cluster that previously had uniform CPU hardware.

In this guide, you will identify the mismatch, choose a compatible baseline, and use `cpu_model_extra_flags` if needed to bridge the gap for a mixed-generation cluster.

## Step 1: Confirm the Cluster's Current CPU Baseline

Check the configured CPU model on one of the existing, healthy hosts:

```bash
grep -A 5 '\[libvirt\]' /opt/pf9/etc/nova/conf.d/nova_override.conf
```

Look for the `cpu_models` value. For example:

```
[libvirt]
cpu_models = Cascadelake-Server-noTSX
```

If `cpu_models` is not set in `nova_override.conf`, the Compute Service is using the auto-selected value. To see what model is actually in use, check the running configuration:

```bash
sudo grep -r "cpu_models" /opt/pf9/etc/nova/
```

## Step 2: Check What the Upgraded Host Reports as Usable

On the upgraded or newly added host, run:

```bash
virsh domcapabilities | grep "model usable='yes'" | sort
```

Compare the output against the cluster's current `cpu_models` value from Step 1. If the current baseline model does not appear in the output of this command, libvirt on this host cannot support that model and will reject the host.

## Step 3: Choose a New Common Baseline (If Required)

Run the usable-model command on every host in the cluster:

```bash
# Run on each host
virsh domcapabilities | grep "model usable='yes'" | sort > /tmp/usable-models-$(hostname).txt
```

Find the intersection — the most recent model that appears as `usable` across all hosts. Use the [Supported CPU Models](https://github.com/platform9/pcd-docs-gitbook/blob/main/private-cloud-director/2026.4/virtualized-clusters/troubleshooting-and-log-files/supported-cpu-models-list.md) list to confirm the model is approved.

If you need to lower the cluster baseline to accommodate a host with a less capable CPU (for example, an older CPU generation added to a cluster running on newer hardware), choose the newest model that all hosts can support.

## Step 4: Apply the New Baseline to All Hosts

Edit `nova_override.conf` on **each** hypervisor host in the cluster. The file must be updated on every host for the change to take effect consistently. Perform this change during a maintenance window, as each host's Compute Service must be restarted.

```bash
sudo vi /opt/pf9/etc/nova/conf.d/nova_override.conf
```

Set or update the `cpu_models` value under `[libvirt]`:

```ini
[libvirt]
cpu_models = <new-common-model>
```

After editing, restart the Compute Service on each host:

```bash
sudo systemctl restart pf9-ostackhost
```

See [Compute Service Advanced Configuration](/private-cloud-director/virtualized-clusters/nova-override#cpu-mode-and-model-configuration) for the full configuration reference.

## Use cpu\_model\_extra\_flags for Mixed-Generation Clusters

In some clusters, hosts from different CPU generations can support the same named model but differ in the specific flags they expose. libvirt may reject a host whose CPU is missing flags that the cluster baseline requires — even if the host reports the model itself as `usable`. This is common when mixing Intel CPU generations within the same cluster.

`cpu_model_extra_flags` is a Compute Service configuration option that explicitly adds or removes CPU feature flags from the model advertised to VMs. It lets you tailor the baseline to the lowest common denominator of flags present across all hosts, so that every host can pass libvirt's CPU compatibility check.

{% hint style="warning" %}
Modifying `cpu_model_extra_flags` affects the CPU feature set visible to all VMs on that host. Adding flags that the physical CPU does not support will cause VM launch failures. Removing flags may reduce VM performance or break workloads that require those flags. Make this change under guidance from [Platform9 Support](https://support.platform9.com/) or after thorough testing in a non-production environment.
{% endhint %}

### Identify Missing Flags

To see which flags the cluster baseline requires but the host is missing, run on the affected host:

```bash
# Get the flags the baseline model expects
cat /usr/share/libvirt/cpu_map/x86_<model-name>.xml | grep 'feature name'

# Get the flags the host actually reports as available
virsh domcapabilities | grep -A 200 '<model usable=.yes.><model-name>' | head -50
```

Alternatively, the Compute Service log on a failing host shows exactly which flag is missing:

```bash
grep -i "cpu flag\|unsupported.*cpu\|cpu.*not found" /var/log/pf9/ostackhost.log | tail -20
```

### Configure cpu\_model\_extra\_flags

Edit `nova_override.conf` on the hosts that are missing the flags. Use a minus prefix to remove a flag or a plus prefix (or no prefix) to add one:

```ini
[libvirt]
cpu_models = Cascadelake-Server-noTSX
cpu_model_extra_flags = -pcid,-ssbd
```

In this example, the flags `pcid` and `ssbd` are removed from the advertised CPU model on this host, bringing it in line with what the cluster's older-generation hosts can offer.

After editing, restart the Compute Service:

```bash
sudo systemctl restart pf9-ostackhost
```

Repeat on every host where the same flag adjustment is needed to ensure the cluster baseline is consistent.

### Verify the Adjusted Model

After restarting, confirm the Compute Service picks up the change:

```bash
sudo grep -A 10 '\[libvirt\]' /opt/pf9/etc/nova/conf.d/nova_override.conf
sudo systemctl status pf9-ostackhost
```

If the host returns to `applied` status in the UI, the CPU model is now accepted.

## Cross-Reference: Host Upgrades and CPU Baseline

When planning a host upgrade, check the CPU baseline compatibility before starting. A host OS upgrade (Ubuntu 22.04 to 24.04) can change the libvirt version and with it the set of flags reported as usable. Review the [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook) for the full pre-upgrade checklist, including the recommendation to verify CPU model consistency across the cluster before and after the upgrade.


# Troubleshoot VM HA

## Overview

This runbook helps you diagnose and recover from the most common VM High Availability (VM HA) failures in <code class="expression">space.vars.product\_name</code>. It covers five scenarios:

1. A host failed but VMs were not evacuated ("VM HA did not trigger").
2. Consul health prerequisites are not met, blocking failure detection (Self-Hosted only).
3. Shared or FC storage is not correctly reachable on all hosts, causing evacuation to fail.
4. Enabling or disabling VM HA returns an error.
5. VM HA needs to be re-validated after a host or management-plane upgrade.

For background on how VM HA works, the prerequisites it requires, and the services involved, see [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha). This runbook assumes you have already read that page.

In this guide, you will identify why VM HA did not trigger or is not functioning correctly, and restore it to a protected state.

## Prerequisites

* Access to the <code class="expression">space.vars.product\_name</code> UI or `pcdctl` CLI.
* SSH access to the hypervisor hosts in the cluster.
* For Self-Hosted deployments: access to the management cluster node and `kubectl` credentials for the region namespace.

***

## Diagnose "VM HA Did Not Trigger" <a href="#vm-ha-did-not-trigger" id="vm-ha-did-not-trigger"></a>

Use this section when a host failed but VMs were not evacuated to other hosts.

### Step 1: Confirm VM HA Is Enabled on the Cluster

In the <code class="expression">space.vars.product\_name</code> UI, navigate to **Infrastructure > Clusters** and open the affected cluster. Expand the **VM High Availability** card on the cluster details page. Confirm that:

* VM HA is toggled **on**.
* The cluster status shows **Protected**. If it shows **Degraded** or **Not Protected**, hover over the status indicator to see which prerequisite is not met.

If VM HA is off or the cluster is **Not Protected**, address the prerequisite failures shown in the UI before continuing.

### Step 2: Verify the Host Was Actually Detected as Down

VM HA will not evacuate VMs until the High Availability Manager has confirmed the host is down (after the 150-second cooldown). In the UI:

1. Navigate to **Infrastructure > Cluster Hosts** and find the affected host.
2. Confirm the host shows an **offline** or **failed** connection status. If the host shows as **online**, VM HA correctly did not evacuate — the host was never confirmed down from the management plane's perspective.
3. On the **Host Details** page, check the **VMHA Past Events** table for any evacuation events. If an event is present, the evacuation may have been attempted but failed; see [#diagnose-evacuation-failures](#diagnose-evacuation-failures "mention") below.

### Step 3: Check the pf9-ha-slave Role on All Hosts

The VM HA agent (installed via the `pf9-ha-slave` role) must be present on every hypervisor host in the cluster. A host missing this role does not participate in peer health probing, which can delay or prevent failure detection.

For each host in the cluster, SSH in and verify:

```bash
sudo apt list --installed 2>/dev/null | grep pf9-ha-slave
```

The package should be listed. If it is missing, the host has not received the role — re-authorize the host and ensure the `pf9-ha-slave` role is applied. From the UI, select the host under **Infrastructure > Cluster Hosts**, click **Edit Roles**, and confirm the HA slave role is assigned.

### Step 4: Verify the libvirt Exporter Is Running

The VM HA agent probes peer hosts using the libvirt exporter. If the libvirt exporter is not running on a host, that host cannot be probed and will appear healthy even if it is not.

On each host:

```bash
sudo systemctl status pf9-libvirt-exporter
```

The service should be `active (running)`. If it is stopped or failed, restart it:

```bash
sudo systemctl restart pf9-libvirt-exporter
```

Check for errors in the service journal:

```bash
sudo journalctl -u pf9-libvirt-exporter --since "1 hour ago" | tail -50
```

### Step 5: Inspect HA Agent Logs

The VM HA agent logs on each host are the primary source of failure information:

```bash
sudo tail -200 /var/log/pf9/ha-slave/ha-slave.log
```

Look for:

| Log pattern                           | Meaning                                                                       |
| ------------------------------------- | ----------------------------------------------------------------------------- |
| `peer list request failed`            | The agent could not reach the High Availability Manager to get its peer list. |
| `liveness check failed for host <IP>` | The agent detected a peer as down; this is normal if a host actually failed.  |
| `report submission failed`            | The agent could not post its health report to the High Availability Manager.  |
| `role not active`                     | The `pf9-ha-slave` role is not converged; re-sync the host.                   |

If the agent cannot reach the management plane, resolve connectivity first (see [Troubleshooting Offline Or Failed Hosts](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-offline-or-failed-hosts)).

### Step 6: Verify High Availability Manager Health

{% hint style="info" %}
**Self-Hosted deployments only**

The High Availability Manager runs as a pod in the region namespace of the management cluster. In SaaS deployments, Platform9 operates the High Availability Manager. If you suspect it is unhealthy, contact Platform9 Support.
{% endhint %}

From the management cluster node, list pods in the region namespace:

```bash
kubectl get pods -n <region-fqdn> | grep -i hamgr
```

The `hamgr` pod should be in `Running` state. If it is in `CrashLoopBackOff`, `Error`, or `Pending`, inspect its logs:

```bash
kubectl logs -n <region-fqdn> <hamgr-pod-name> --tail=100
```

Look for connection errors, database timeouts, or Consul-related failures. If the pod is crash-looping, escalate to Platform9 Support with the log output.

### Diagnose Evacuation Failures <a href="#diagnose-evacuation-failures" id="diagnose-evacuation-failures"></a>

If the **VMHA Past Events** table on the Host Details page shows an evacuation event with failed VMs:

1. Click **View Details** on the event to see per-VM evacuation status.
2. Note the fault message for each failed VM. Common fault messages:

   | Fault message                                        | Likely cause                                                                                                                                |
   | ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
   | `No valid host was found`                            | No destination host met placement constraints or had enough resources.                                                                      |
   | `Volume not found` or `connection to storage failed` | Shared storage was not reachable on the destination host. See [#validate-shared-and-fc-storage](#validate-shared-and-fc-storage "mention"). |
   | `KVM version mismatch`                               | Source and destination hosts run different OS versions. Mixed OS versions are not supported for evacuation.                                 |
   | `Flavor with host aggregate requirement not met`     | The VM's flavor pins it to a host aggregate with only one host.                                                                             |
3. For VMs that can be retried, click **Retry** on the evacuation detail page, or retry from the CLI:

   ```bash
   pcdctl server evacuate --host <DESTINATION_HOST_UUID> <VM_UUID>
   ```

For general VM ERROR-state recovery after evacuation failures, see [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state).

***

## Check Consul Health Prerequisites <a href="#consul-health" id="consul-health"></a>

Consul is used by the High Availability Manager for distributed coordination and service health detection. If Consul is unhealthy, the High Availability Manager cannot reliably detect host failures, and VM HA will not function correctly.

{% hint style="info" %}
**Self-Hosted deployments only**

Consul runs as pods inside the management cluster, so the checks and recovery steps in this entire section apply to Self-Hosted deployments only. In SaaS deployments, Platform9 operates the management plane (including Consul). If VM HA is not triggering and you suspect a management-plane problem, contact Platform9 Support.
{% endhint %}

### Identify Consul Symptoms

Symptoms of a Consul problem include:

* The `hamgr` pod logs show repeated `failed to connect to consul` or `consul agent not healthy` messages.
* VM HA events are not generated even though hosts are clearly down.
* The `hamgr` pod is restarting frequently.

### Verify Consul Health

From the management cluster node:

```bash
kubectl get pods -n <region-fqdn> | grep consul
```

All Consul pods should be in `Running` state. A quorum of Consul pods is required for the cluster to be healthy — if more than half the Consul pods are down, the Consul cluster loses quorum and all distributed locks and health checks fail.

Check Consul cluster health using the Consul CLI from inside one of the running Consul pods:

```bash
kubectl exec -n <region-fqdn> <consul-pod-name> -- consul members
```

All nodes should show `Status: alive`. Nodes in `Status: failed` or `Status: left` indicate that part of the Consul cluster is unavailable.

To check leader election status:

```bash
kubectl exec -n <region-fqdn> <consul-pod-name> -- consul operator raft list-peers
```

There must be exactly one leader. If no leader is elected, the cluster has lost quorum.

### Recover a Degraded Consul Cluster

If one Consul pod is down and the remaining pods still have quorum (a majority are running), the cluster is degraded but functional. Restart the failed pod:

```bash
kubectl delete pod -n <region-fqdn> <failed-consul-pod-name>
```

Kubernetes will automatically reschedule the pod. Wait for it to rejoin the Consul cluster (`consul members` shows it as `alive`).

If the Consul cluster has lost quorum (fewer than half the pods are running), recovery is more involved. Escalate to Platform9 Support immediately, as forcing a quorum recovery on a production Consul cluster carries risk.

After Consul is healthy, confirm the `hamgr` pod is also healthy and restart it if needed:

```bash
kubectl rollout restart deployment/hamgr -n <region-fqdn>
```

Wait for the pod to return to `Running` and confirm VM HA events resume.

***

## Validate Shared and FC Storage <a href="#validate-shared-and-fc-storage" id="validate-shared-and-fc-storage"></a>

VM HA requires that the storage backing each VM's root disk is reachable from the destination host, not just the source host. Evacuation will fail if a shared NFS mount is not mounted on the destination, or if an FC or iSCSI volume is not accessible on the destination's storage backend.

### Confirm Shared Storage Is Mounted on All Hosts

For NFS-backed ephemeral storage or image library storage, verify the mount is present on every host in the cluster:

```bash
mount | grep nfs
```

The expected NFS share should appear. If it is missing on any host, remount it and verify it persists across reboots by checking `/etc/fstab`.

For each host, also confirm the mount is writable:

```bash
touch /mnt/<shared-storage-path>/.ha_probe_$(hostname) && echo "writable" || echo "not writable"
```

Remove the probe file after confirming:

```bash
rm /mnt/<shared-storage-path>/.ha_probe_$(hostname)
```

### Confirm Block Storage Volumes Are Accessible on Destination Hosts

For VMs using block storage volumes (iSCSI, FC, or other backends), the volume must be attachable to a host other than the one it is currently attached to. Verify:

1. List the volume and confirm its type and backend:

   ```bash
   pcdctl volume show <VOLUME_UUID>
   ```

   Check the `volume_type` field. Confirm that at least one other host in the cluster has the same backend configured.
2. For FC (Fibre Channel) volumes specifically, ensure the destination host has:

   * A physical FC HBA (Host Bus Adapter) installed and enabled.
   * Zoning that allows the destination host to see the same storage array LUNs as the source host.

   On the destination host, scan for visible FC targets:

   ```bash
   sudo rescan-scsi-bus.sh
   lsscsi | grep -i fc
   ```

   If the LUN is not visible, contact your storage administrator to verify FC zoning includes the destination host.
3. Verify that the Persistent Storage Service role is applied to at least two hosts in the cluster:

   ```bash
   pcdctl volume service list
   ```

   All endpoints should show `enabled` and `up`. If only one endpoint is enabled, VM HA may have no valid storage backend to use on the destination host, causing evacuation to fail for block-volume VMs.

### Confirm the "This Is Shared Storage" Toggle for Ephemeral Storage

If your cluster uses ephemeral storage (VMs without block volume root disks), ephemeral shared storage must be enabled on the cluster blueprint. In the UI:

1. Navigate to **Infrastructure > Clusters** and open the cluster.
2. Open the **Cluster Blueprint** and expand **Customize Cluster Defaults**.
3. Confirm the **This is Shared Storage** toggle is enabled.

If the toggle is off, VM HA cannot evacuate VMs that use ephemeral root disks. See [Ephemeral Shared Storage](/private-cloud-director/storage/ephemeral-storage#ephemeral-shared-storage) for setup details.

***

## Diagnose Enable and Disable VM HA Failures <a href="#enable-disable-failures" id="enable-disable-failures"></a>

### 404 or 503 Error When Enabling VM HA

If the UI shows an error (for example, a 404 or 503) when you toggle VM HA on or off, the High Availability Manager is not reachable or is reporting an internal error.

**Verify prerequisites first:**

Before trying to enable VM HA, confirm the following are true. The toggle will silently fail or return an error if any prerequisite is not met:

* At least two hosts with the Hypervisor role are in the cluster and both are **online**.
* All hosts in the cluster belong to the same host aggregate, or none belong to any aggregate. Mixed-aggregate clusters cannot enable VM HA.
* The `pf9-ha-slave` role is applied and **converged** on all hosts (not just assigned — the role status must be `applied`, not `converging`).
* Consul is healthy (see [#consul-health](#consul-health "mention")).

**Check the High Availability Manager directly:**

{% hint style="info" %}
**Self-Hosted deployments only**

In SaaS deployments, contact Platform9 Support if you receive a persistent error when toggling VM HA.
{% endhint %}

From the management cluster node:

```bash
kubectl get pods -n <region-fqdn> | grep hamgr
```

If the `hamgr` pod is not `Running`, inspect its logs and restart it:

```bash
kubectl logs -n <region-fqdn> <hamgr-pod-name> --tail=100
kubectl rollout restart deployment/hamgr -n <region-fqdn>
```

After the pod is running, retry the toggle.

### VM HA Toggle Appears Enabled but Status Is "Not Protected"

After enabling VM HA, if the cluster immediately shows **Not Protected** or **Degraded**, hover over the VM HA status indicator in the UI to see which specific prerequisite is failing. Common causes:

| Reported prerequisite failure               | Action                                                                                                                                                                              |
| ------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Fewer than two hypervisor hosts             | Add a second host to the cluster.                                                                                                                                                   |
| Persistent Storage Service not redundant    | Add the Persistent Storage Service role to a second host.                                                                                                                           |
| Image Library Service not on shared storage | Configure the Image Library Service to use shared storage; see [Image Library High Availability](/private-cloud-director/images-and-image-library/image-library-high-availability). |
| `pf9-ha-slave` role not applied             | Re-sync hosts: navigate to the host, click **Other > Re-sync Host**.                                                                                                                |

### VM HA Remains Enabled but HA Manager Reports No Active Clusters

If VM HA is enabled but the High Availability Manager logs show no HA-enabled clusters being discovered, the management plane may have a stale cluster state. In the UI, disable VM HA on the cluster and re-enable it to trigger a fresh discovery cycle. If the problem persists in a Self-Hosted deployment, check the `hamgr` pod logs for database connectivity errors.

***

## Validate VM HA After an Upgrade <a href="#post-upgrade-validation" id="post-upgrade-validation"></a>

After upgrading host packages or the management plane, run the following checks before relying on VM HA for workload protection.

{% hint style="warning" %}
**Re-enable VM HA only after all hosts are on the same OS version.** If you disabled VM HA before a host OS upgrade, do not re-enable it while any hosts in the cluster are still running the old OS version. VM evacuation between hosts with different KVM versions will fail. See [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook) for the full upgrade sequence, and [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification#re-enable-vm-ha-and-drr) for the re-enablement steps.
{% endhint %}

### Check 1: Confirm pf9-ha-slave Is Re-Applied on All Hosts

After a host OS upgrade, the `pf9-ha-slave` role must be re-applied. Verify on each host:

```bash
sudo apt list --installed 2>/dev/null | grep pf9-ha-slave
sudo systemctl status pf9-ha-slave
```

The service should be `active (running)`. If it is not present or stopped, re-sync the host from the UI (**Other > Re-sync Host**) to re-apply the role, then confirm the service starts.

### Check 2: Confirm the libvirt Exporter Is Running on All Hosts

```bash
sudo systemctl status pf9-libvirt-exporter
```

If stopped, restart it:

```bash
sudo systemctl restart pf9-libvirt-exporter
```

### Check 3: Verify High Availability Manager Health

{% hint style="info" %}
**Self-Hosted deployments only**

After a management-plane upgrade, confirm the `hamgr` pod is running and healthy. In SaaS deployments, Platform9 performs management-plane upgrades; confirm VM HA status from the UI after Platform9 notifies you that the upgrade is complete.
{% endhint %}

```bash
kubectl get pods -n <region-fqdn> | grep hamgr
```

If the pod is not `Running`, inspect its logs and restart the deployment:

```bash
kubectl logs -n <region-fqdn> <hamgr-pod-name> --tail=100
kubectl rollout restart deployment/hamgr -n <region-fqdn>
```

### Check 4: Re-Enable VM HA and Confirm "Protected" Status

Once all hosts are on the same OS version and all services are healthy:

1. Navigate to **Infrastructure > Clusters**.
2. Select the cluster.
3. Toggle **VM High Availability** to on.
4. Confirm the cluster transitions to **Protected** status within one to two minutes.

If the cluster shows **Degraded** or **Not Protected** after re-enabling, hover over the status indicator and address each failing prerequisite. See [#enable-disable-failures](#enable-disable-failures "mention") for a reference table of common prerequisite failures.

### Check 5: Perform a Safe Validation Test

To confirm VM HA is functional without inducing a real failure, you can use maintenance mode as a proxy test:

1. Place a non-critical host into **Maintenance Mode** (see [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode)). This live-migrates its VMs to other hosts in the cluster.
2. Confirm the VMs migrate successfully and return to **ACTIVE** state.
3. Disable Maintenance Mode to bring the host back into service.

This validates that live migration works correctly on the current OS version and storage configuration. VM HA evacuation uses the same migration path, so a successful maintenance-mode drain confirms the underlying machinery is working.

{% hint style="info" %}
Maintenance mode uses **live migration**, whereas VM HA uses **evacuation** when a host is unresponsive. Live migration requires the source host to be online; evacuation does not. The maintenance-mode test validates the migration path but does not exercise the failure detection path. If you want to validate the full VM HA detection path, contact Platform9 Support for guidance on scheduling a controlled failover test.
{% endhint %}

***

## Related Pages

* [Virtual Machine High Availability (VM HA)](/private-cloud-director/virtualized-clusters/virtualized-cluster/virtual-machine-high-availability-vm-ha)
* [Recover VMs in ERROR State After Host Reboot or Patching](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/recover-vms-in-error-state)
* [Troubleshooting Offline or Failed Hosts](/private-cloud-director/virtualized-clusters/troubleshooting-and-log-files/troubleshooting-offline-or-failed-hosts)
* [Maintenance Mode](/private-cloud-director/virtualized-clusters/add-hosts-virtualized-cluster/maintenance-mode)
* [Block Storage High Availability](/private-cloud-director/storage/block-storage/block-storage-high-availability)
* [Host Upgrade Runbook](/private-cloud-director/upgrade/host-upgrade-runbook)
* [Post-Upgrade Verification](/private-cloud-director/upgrade/post-upgrade-verification)


# Resolve Live Migration Failures for Legacy Boot-from-Volume Hotplug VMs

## Overview

Live migration can fail for virtual machines that combine two specific characteristics:

* The VM was created using a hot-add-capable (zero-size) flavor. See [VM Hot Add CPU Or Memory](/private-cloud-director/virtualized-clusters/virtualmachine/vm-hot-add-cpu-or-memory) and [Flavorless VM Support](/private-cloud-director/virtualized-clusters/virtualmachine/flavorless-vms-with-hot-plug#zero-size-flavor).
* The VM boots from a zero-disk / boot-from-volume configuration, meaning its root disk is a persistent (block storage) volume rather than ephemeral storage. See [Live Migration of a VM using Volumes only](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration#live-migration).

This is a pre-existing issue and is not caused by upgrading. The <code class="expression">space.vars.product\_acronym</code> 2025.10 (October 2025) release corrected how the Compute Service generates the resource-request record for VMs of this type, but that correction only applies going forward — it is applied when the VM is created or when a hotplug operation is performed on it. VMs that were created with this combination of flavor and boot configuration before the 2025.10 release still hold the older, incorrect resource-request record, and live migration continues to fail for them until that record is corrected.

In this guide, you will identify whether a VM is affected and apply a one-time corrective step to unblock live migration for it.

## Symptoms

* A VM that uses a hot-add-capable flavor and boots from a persistent volume fails to live migrate, even though other VMs in the same virtualized cluster migrate successfully.
* The VM was created before the cluster was upgraded to the <code class="expression">space.vars.product\_acronym</code> 2025.10 release or later.
* The VM has not had a hotplug (hot-add CPU or memory) operation performed on it since being created.

## Identify Affected VMs

Check whether the VM uses a hot-add-capable flavor and a persistent-volume root disk:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl server show <VM_UUID>
```

{% endtab %}
{% endtabs %}

* Confirm the VM's flavor has zero vCPU and zero RAM configured at the flavor level (hot-add-capable), and that the VM currently reports non-zero vCPU/RAM values.
* Confirm the VM's root disk is a persistent volume rather than an ephemeral disk. See [Boot VM from New Volume / Boot VM from Existing Volume](/private-cloud-director/virtualized-clusters/virtualmachine) for the boot configurations that use a volume-backed root disk.

If the VM meets both criteria and was created before the cluster's upgrade to <code class="expression">space.vars.product\_acronym</code> 2025.10 or later, it is a candidate for the workaround below.

## Workaround: Trigger a No-Op Hotplug

Triggering a hotplug operation on the affected VM — without changing its current CPU or memory values — regenerates the VM's resource-request record using the corrected logic. This is a one-time step per affected VM; once applied, live migration works normally for that VM going forward.

{% hint style="info" %}
**NOTE**

Re-submitting the VM's current CPU and memory values causes no actual change to the VM's resources. The VM does not need to be power-cycled for this step.
{% endhint %}

1. Select the affected VM in the VM grid view.
2. Navigate to **▷ Other Actions ▷ Hotplug**.
3. Leave the vCPU and memory fields set to the VM's current values — do not change them.
4. Click **Hotplug VM**.

After this step completes, retry the live migration.

## Next Steps

Repeat this workaround for each affected VM in your virtualized cluster. VMs created after the cluster's upgrade to <code class="expression">space.vars.product\_acronym</code> 2025.10 or later, and any VM that already had a hotplug operation performed on it, are not affected and do not require this step. For general live migration behavior and prerequisites, see [Virtual Machine Migration](/private-cloud-director/virtualized-clusters/virtualmachine/vm-migration).


# Backup

TBD


# Performance Tuning

## Overview

This page explains how to tune VM performance in Private Cloud Director (PCD) by setting image properties. The recommendations listed below should yield positive results for the vast majority of use cases, but the actual outcome will depend on factors like NUMA node topology and the workloads running within the VM.

**Note**: Most image property changes apply only to new VMs created after making the update. For existing VMs, plan to rebuild the VM to see the effect.

## Recommended Image Properties

To update image properties, navigate to **Virtual Machines** in the left-hand navigation menu > **Images & VM Snapshots** > **Images** tab. Select an image and click on **Edit Properties**. You can enter one or more of the configurations from this section in a key-value format. Click on **Update Properties** after you complete your updates.

### OS Type

The `os_type` key indicates what OS is contained in the image.

**Values**

Set the value to one of the following, depending on the image OS.

* `linux`
* `windows`

Please note that the January 2026 release of PCD includes a required drop-down for setting this value in the new image upload wizard. You only need to set this value manually for images uploaded prior to the January 2026 release.

### Machine Type

The `hw_machine_type` key indicates the virtualized chipset used by VMs. Set the value to `q35` to use the newer Q35 chipset for modern x86 workloads.

**Values**

PCD supports two main variants of machine type for x86 hosts:

* `pc`, which corresponds to Intel's I440FX chipset (released in 1996)
* `q35`, which corresponds to Intel's 82Q35 chipset (released in 2007)

The `pc` machine type is considered legacy, and does not support many modern features. Some long-term stable Linux distributions (CentOS, RHEL, possibly others) are moving to support `q35` only.

### Paravirtualized I/O

#### Disk Controller

The `hw_disk_bus` key sets the **virtual disk controller type** that the VM will see when it boots from the image. Setting the value to `virtio` presents a **paravirtualized** disk device, which generally yields better performance than fully emulated legacy controllers.

**Value**

* `virtio` attaches disks using a paravirtualized controller/device model.

#### Virtual Network Interface Model

The `hw_vif_model` key sets the **virtual network adapter model** exposed to the guest operating system for VMs created from the image. Setting it to `virtio` presents a **paravirtualized** network device, which is commonly used for better network performance than fully emulated adapters.

**Value**

* `virtio` uses a paravirtualized network adapter model.

### QEMU Guest Agent

The `hw_qemu_guest_agent` controls whether QEMU Guest Agent support is enabled for VMs created from an image. When set to `yes`, the compute host can communicate with the guest operating system through a QEMU Machine Protocol (QMP) socket (assuming the QEMU guest agent is installed and running inside the VM). Set it to `no` to explicitly disable guest agent support.

**Values**

* `yes` to enable QEMU guest agent communication via QMP socket
* `no` to disable guest agent support


# Overview

You can use an image library to store and manage virtual machine images in your <code class="expression">space.vars.product\_name</code> environment. An image captures the state of a virtual machine, including the operating system, applications, data, and configurations.

Before you provision virtual machines, you must import images into your image library. You can then use these images to create new virtual machines.

## Prerequisites

Before configuring an image library, ensure that you meet the [Image Library Prerequisites](/private-cloud-director/getting-started/pre-requisites#image-library-prerequisites).

### Supported Backends

The image library service supports the following storage backends:

* **File** – Local file-based image storage.
* **Block storage** – <code class="expression">space.vars.product\_name</code> block storage service for image storage.

## Configure an Image Library

Each cluster requires at least one image library to host virtual machine images. Configure an image library by assigning the image library role to one or more hosts in your cluster.

Hosts with the image library role store virtual machine images on local, block, or file storage. These hosts serve as the image library endpoint when providing images to hosts during VM provisioning.

For high availability in production environments, assign the image library role to at least two hosts in your virtualized cluster.

To restrict a cluster's hosts to fetch images only from that cluster's own Image Library host instead of from any host in the region, see [Restrict the Image Library Service to a Specific Cluster](/private-cloud-director/images-and-image-library/restrict-image-library-to-cluster).

### Storage capacity considerations

Hosts with the Image Library role must have sufficient storage capacity for their VM image catalog. These hosts experience significant network I/O when serving an image for the first time during VM provisioning. After the first provision, hypervisors maintain a local cache of images to reduce I/O.

## Import Images

Before you create VMs, import images into your image library. You can import images using the <code class="expression">space.vars.product\_name</code> UI or the `pcdctl` CLI.

To upload an image through the UI without installing the CLI, see [Upload an Image Using the UI](/private-cloud-director/images-and-image-library/image-upload-via-ui).

The following steps show you how to import images using the CLI method.

### Step 1: Download and install the pcdctl CLI client

Download the `pcdctl` CLI client on a machine that has network access to your hypervisor hosts and your image files.

**To download and install pcdctl**

For detailed installation instructions, see [Installation](/private-cloud-director/reference/pcdctl-command-line#installation)

### Step 2: Configure authentication with environment variables

Configure the CLI to authenticate with your Private Cloud Director account by setting the required environment variables. The Private Cloud Director UI provides pre-populated environment variables based on your current domain, region, and tenant.

To set environment variables perform the following steps:

1. Sign in to your <code class="expression">space.vars.product\_name</code> account and select **Images**.
2. Choose **Import With CLI**.
3. Copy the environment variables displayed in the UI. The variables are pre-populated with values for your current domain, region, and tenant.
4. On your local machine, create a new `.sh` file (for example, `openstack-rc.sh`) and paste the copied variables.
5. In the file, replace the value of `OS_PASSWORD` with your <code class="expression">space.vars.product\_name</code> account password.
6. Save the file.
7. Run the following command to set the environment variables:

{% tabs %}
{% tab title="Bash" %}

```bash
source <NAME_OF_YOUR_RC_FILE>.sh
```

{% endtab %}
{% endtabs %}

### Step 3: Upload your VM image to the Image Library

After you configure authentication, use the `pcdctl` command to upload your VM image to the image library.

To upload an image, perform the following steps.

1. Run the following command:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl image create --insecure --container-format bare --disk-format qcow2 [--public | --private] [--protected | --unprotected] [--property <key=value>] --file <image-file-path> <image-name>
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**NOTE**

The `--insecure` flag is required because the image library service uses self-signed certificates.
{% endhint %}

2. (Optional) To make the image public and available to all tenants, include the `--public` flag:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl image create --insecure --container-format bare --disk-format qcow2 --public --file <image-file-path> <image-name>
```

{% endtab %}
{% endtabs %}

After you upload an image, you can use it to create virtual machines.

## Configure high availability

You can create a highly available image library by assigning the image library role to multiple hosts in your virtualized cluster. For production environments, we recommend configuring high availability for your image library.

For more information, see [Image Library High Availability](/private-cloud-director/images-and-image-library/image-library-high-availability).

## Manage the admin endpoint

The image library admin endpoint is the IP address of the image library service host used to upload images. <code class="expression">space.vars.product\_name</code> automatically configures this endpoint when you assign the image library role to a host.

When you configure high availability with multiple image library hosts, the last host assigned the image library role becomes the admin endpoint.

### Change the admin endpoint

In a highly available setup, if the host acting as the admin endpoint becomes unavailable, you must manually assign a different image library host as the admin endpoint. This change is required only for uploading images. VM provisioning continues to use any available image library host in a round-robin fashion.

For instructions to change the admin endpoint, see [Image Library High Availability](/private-cloud-director/images-and-image-library/image-library-high-availability).

## Control image visibility

You can control which tenants can access an image by setting its visibility. Images support the following visibility options:

* **Private** – Available only within the tenant where the image was created.
* **Public** – Available to all tenants in your domain (set using the `--public` option during CLI upload or in the UI)
* **Shared** – Available to specific tenants in your domain (default option for CLI uploads)

## Understand images and snapshots <a href="#images-vs-vm-snapshots" id="images-vs-vm-snapshots"></a>

You can use the image library to store both images and virtual machine snapshots.

For more information about VM snapshots, snapshot types, and behavior, see [Virtual Machine Snapshot](/private-cloud-director/virtualized-clusters/virtualmachine/virtual-machine-snapshot).

## Supported image formats

<code class="expression">space.vars.product\_name</code> supports the following image formats:

| **Format** | **Description**                                                                                                                                                                                                                                                                                                                                                                         |
| ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| raw        | An unstructured disk image format natively supported by the KVM hypervisor. A raw image is equivalent to a block device file created using the `dd` command (for example, copying `/dev/sda` to a file).                                                                                                                                                                                |
| qcow2      | QEMU copy-on-write version 2 format, commonly used with the KVM hypervisor. This format utilizes sparse representation, resulting in smaller image sizes compared to raw format files. The format supports dynamic expansion and copy-on-write. Files using this format typically have a `.img` extension. For more information, see [QCOW2](http://en.wikibooks.org/wiki/QEMU/Images). |
| iso        | A disk image formatted with the read-only ISO 9660 (ECMA-119) filesystem, commonly used for CDs and DVDs. For more information, see [ISO](http://www.ecma-international.org/publications/standards/Ecma-119.htm).                                                                                                                                                                       |


# Upload an Image Using the UI

## Overview

You can upload a virtual machine image to the Image Library Service directly from the <code class="expression">space.vars.product\_name</code> UI without installing the `pcdctl` CLI. The UI upload is suited for interactive, one-off imports; for scripted or bulk uploads use the `pcdctl` CLI method described in [Overview](/private-cloud-director/images-and-image-library/image-library---images).

In this guide, you will upload an image through the UI and verify that it reaches `active` status.

## Prerequisites

* You have `administrator` permissions in the target tenant (project).
* You have accepted the Image Library Service certificate. If the upload button is greyed out or a certificate warning appears, see [Image Library Service Certificate Configuration](/private-cloud-director/images-and-image-library/image-library-certificate-configuration) before continuing.
* The image file is in a [supported format](/private-cloud-director/images-and-image-library/image-library---images#supported-image-formats) (`raw`, `qcow2`, or `iso`).
* The Image Library Service host has sufficient free disk space to hold the image. Check available space on the Image Library host before uploading large images.

## Upload an Image

1. Sign in to the <code class="expression">space.vars.product\_name</code> UI and navigate to **Images**.
2. Select the **Import** button (or **Import Image**, depending on your UI version).
3. In the import dialog, complete the following fields:

   | Field                | Description                                                                                                                                 |
   | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------- |
   | **Image Name**       | A descriptive name for the image.                                                                                                           |
   | **Image File**       | Browse to and select the image file on your local machine.                                                                                  |
   | **Disk Format**      | Select the disk format that matches your file: `qcow2`, `raw`, or `iso`.                                                                    |
   | **Container Format** | Select `bare` for most images (the image file contains only the disk image, with no outer container).                                       |
   | **Visibility**       | Select **Public** to make the image available to all tenants (projects) in the domain, or **Private** to restrict it to the current tenant. |
   | **Protected**        | Toggle on to prevent accidental deletion of the image.                                                                                      |
4. Select **Import** to start the upload.
5. The image row appears in the **Images** table with a status of `queued` while the file transfers, then transitions to `saving`, and finally to `active` when the upload completes successfully.

{% hint style="info" %}
**Large images**

Large image files (several gigabytes) can take several minutes to upload depending on your network bandwidth to the Image Library host. The UI does not display a progress percentage for the file transfer. If the status stays at `queued` for more than a few minutes, see [Troubleshoot Image Upload Issues](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide#troubleshoot-image-upload-and-queued-status).
{% endhint %}

## Common Upload Errors

| Symptom                                                           | Likely Cause                                                                        | Resolution                                                                                                                                                                                              |
| ----------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Upload button is not active or the import dialog cannot be opened | Image Library Service certificate is not trusted in your browser                    | Follow [Image Library Service Certificate Configuration](/private-cloud-director/images-and-image-library/image-library-certificate-configuration).                                                     |
| Upload starts but the image stays in `queued` indefinitely        | Disk full on the Image Library host, service not running, or a network interruption | See [Troubleshoot Image Upload Issues](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide#troubleshoot-image-upload-and-queued-status). |
| Upload fails immediately with an HTTP error                       | Incorrect disk format or container format selected, or the image file is corrupted  | Verify the format with `qemu-img info <image-file>` on a local machine and retry with the correct format.                                                                                               |
| `403 Forbidden` error                                             | Your user account does not have `administrator` permissions in the current tenant   | Ask your administrator to grant the required role in the target tenant.                                                                                                                                 |

## Next Steps

After the image reaches `active` status, you can use it to create virtual machines. For advanced image configuration, such as setting image properties that control VM behavior, see [Image Properties](/private-cloud-director/images-and-image-library/image-properties).


# Image Library Service Certificate Configuration

## Overview

The Image Library Service uses a TLS certificate to secure image upload and retrieval traffic. If your browser or CLI client does not trust that certificate, image uploads from the UI will be blocked, and CLI uploads without the `--insecure` flag will fail.

<code class="expression">space.vars.product\_name</code> supports two deployment models with different certificate management paths:

* **SaaS** — Platform9 operates the management plane. The Image Library Service endpoint certificate is managed by Platform9. You accept the certificate in your browser; you do not modify management-plane certificates directly.
* **Self-Hosted** — You operate the management plane on-premise. You can supply your own custom certificate and apply it using `airctl`.

In this guide, you will trust or configure the Image Library Service certificate for your deployment model so that image uploads and VM provisioning succeed.

## SaaS Deployments

In SaaS deployments, the Image Library Service host uses a self-signed certificate for the image upload endpoint. You must accept this certificate in your browser before you can upload images through the UI.

### Accept the Certificate in Your Browser

1. Sign in to the <code class="expression">space.vars.product\_name</code> UI.
2. Navigate to **Images**. If the certificate has not been trusted, a banner or notification appears that reads **Action Required: Trust Certificate**.
3. Click the link in the notification. A new browser tab opens to the Image Library Service endpoint.
4. In the browser's security warning page, expand **Advanced** (or **Details**, depending on your browser) and click the link to proceed to the site.
5. After accepting the certificate, close the tab and return to the **Images** page. The upload controls should now be active.

For detailed steps for each browser, see [Accept Certificate Authority](/private-cloud-director/getting-started/accept-certificate-authority).

{% hint style="info" %}
**Certificate acceptance is per browser and per machine**

Each browser on each machine must accept the certificate separately. If a team member on a different workstation reports upload issues after you have already accepted the certificate, they must repeat these steps in their own browser.
{% endhint %}

### CLI Uploads (SaaS)

When using the `pcdctl` CLI to upload images in a SaaS deployment, include the `--insecure` flag to bypass certificate verification for the Image Library endpoint:

```bash
pcdctl image create --insecure --container-format bare --disk-format qcow2 \
  --file <image-file-path> <image-name>
```

This flag applies to the image endpoint only. It does not affect Identity Service authentication.

### Management-Plane Certificate Changes (SaaS)

{% hint style="warning" %}
**SaaS deployments only**

In SaaS deployments, Platform9 manages the management-plane certificates. You cannot modify these certificates directly. If you need a custom or CA-signed certificate for the Image Library Service endpoint in a SaaS deployment, contact [Platform9 Support](https://support.platform9.com/).
{% endhint %}

## Self-Hosted Deployments

In Self-Hosted deployments, you have full control over management-plane certificates and can supply a custom certificate.

### Accept the Default Self-Signed Certificate

If you are using the default self-signed certificate generated during installation, follow the same browser-acceptance steps as for SaaS deployments (see above). This is the quickest path for small or evaluation deployments.

For CLI uploads, use the `--insecure` flag with `pcdctl image create` as shown in the SaaS section above.

### Configure a Custom Certificate

{% hint style="info" %}
**Self-Hosted deployments only**

The steps in this section apply only to Self-Hosted <code class="expression">space.vars.self\_hosted\_product\_name</code>. In SaaS deployments, contact Platform9 Support for certificate changes.
{% endhint %}

To replace the default self-signed certificate with a CA-signed or custom certificate, use the `airctl renew-certs` command. This updates the management-plane certificate, which is then used by the Image Library Service and other services.

For full instructions, see [Using Custom Certificates](/private-cloud-director/getting-started/self-hosted/using-custom-certificates).

After applying a new certificate, verify that the Image Library Service endpoint is reachable:

```bash
curl -s https://<IMAGE_LIBRARY_HOST_FQDN>:9292/
```

If you configured a CA-signed certificate, this request should succeed without the `-k` flag. If it returns a certificate error, confirm that the certificate's Subject Alternative Names (SANs) include the Image Library host's FQDN or IP address.

### Verify Certificate Trust After Changes

After updating certificates in a Self-Hosted deployment, restart the Image Library Service on the affected host:

```bash
sudo systemctl restart pf9-glance-api
```

Then re-check the endpoint health and attempt a test upload to confirm the change took effect.

## Troubleshooting Certificate Issues

| Symptom                                                      | Likely Cause                                                           | Resolution                                                             |
| ------------------------------------------------------------ | ---------------------------------------------------------------------- | ---------------------------------------------------------------------- |
| UI upload button is not active; certificate warning shown    | Certificate not accepted in this browser                               | Follow the browser-acceptance steps above.                             |
| `SSL: CERTIFICATE_VERIFY_FAILED` in CLI output               | `--insecure` flag omitted, or a CA-signed cert whose CA is not trusted | Add `--insecure`, or add the CA certificate to the system trust store. |
| Certificate accepted, but upload still blocked               | Browser cached an older, untrusted state                               | Clear the browser cache or retry in a private window.                  |
| Custom certificate applied but endpoint still shows old cert | Service not restarted after cert change                                | Run `sudo systemctl restart pf9-glance-api` on the Image Library host. |

For general Image Library Service health checks, see [Image Library Service Endpoint Health](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide#image-library-service-endpoint-health).


# Image Properties

This page describes supported properties for Private Cloud Director images.

This page describes the various properties that are supported by <code class="expression">space.vars.product\_name</code> images.

### architecture

**Property**: `architecture`

**Type**: string

**Description**:

The CPU architecture that must be supported by the hypervisor. For example, `x86_64`, `arm`, or `ppc64`. Run **uname -m** to get the architecture of a machine. We strongly recommend using the architecture data vocabulary defined by the [libosinfo project](http://libosinfo.org/) for this purpose.

**Supported values:**

* `aarch64` - [ARM 64-bit](https://en.wikipedia.org/wiki/AArch64)
* `alpha` - [DEC 64-bit RISC](https://en.wikipedia.org/wiki/DEC_Alpha)
* `armv7l` - [ARM Cortex-A7 MPCore](https://en.wikipedia.org/wiki/ARM_architecture)
* `cris` - [Ethernet, Token Ring, AXis—Code Reduced Instruction Set](https://en.wikipedia.org/wiki/ETRAX_CRIS)
* `i686` - [Intel sixth-generation x86 (P6 micro architecture)](https://en.wikipedia.org/wiki/X86)
* `ia64` - [Itanium](https://en.wikipedia.org/wiki/Itanium)
* `lm32` - [Lattice Micro32](https://en.wikipedia.org/wiki/Milkymist)
* `m68k` - [Motorola 68000](https://en.wikipedia.org/wiki/Motorola_68000_family)
* `microblaze` - [Xilinx 32-bit FPGA (Big Endian)](https://en.wikipedia.org/wiki/MicroBlaze)
* `microblazeel` - [Xilinx 32-bit FPGA (Little Endian)](https://en.wikipedia.org/wiki/MicroBlaze)
* `mips` - [MIPS 32-bit RISC (Big Endian)](https://en.wikipedia.org/wiki/MIPS_architecture)
* `mipsel` - [MIPS 32-bit RISC (Little Endian)](https://en.wikipedia.org/wiki/MIPS_architecture)
* `mips64` - [MIPS 64-bit RISC (Big Endian)](https://en.wikipedia.org/wiki/MIPS_architecture)
* `mips64el` - [MIPS 64-bit RISC (Little Endian)](https://en.wikipedia.org/wiki/MIPS_architecture)
* `openrisc` - [OpenCores RISC](https://en.wikipedia.org/wiki/OpenRISC#QEMU_support)
* `parisc` - [HP Precision Architecture RISC](https://en.wikipedia.org/wiki/PA-RISC)
* `parisc64` - [HP Precision Architecture 64-bit RISC](https://en.wikipedia.org/wiki/PA-RISC)
* `ppc` - [PowerPC 32-bit](https://en.wikipedia.org/wiki/PowerPC)
* `ppc64` - [PowerPC 64-bit](https://en.wikipedia.org/wiki/PowerPC)
* `ppcemb` - [PowerPC (Embedded 32-bit)](https://en.wikipedia.org/wiki/PowerPC)
* `s390` - [IBM Enterprise Systems Architecture/390](https://en.wikipedia.org/wiki/S390)
* `s390x` - [S/390 64-bit](https://en.wikipedia.org/wiki/S390x)
* `sh4` - [SuperH SH-4 (Little Endian)](https://en.wikipedia.org/wiki/SuperH)
* `sh4eb` - [SuperH SH-4 (Big Endian)](https://en.wikipedia.org/wiki/SuperH)
* `sparc` - [Scalable Processor Architecture, 32-bit](https://en.wikipedia.org/wiki/Sparc)
* `sparc64` - [Scalable Processor Architecture, 64-bit](https://en.wikipedia.org/wiki/Sparc)
* `unicore32` - [Microprocessor Research and Development Center RISC Unicore32](https://en.wikipedia.org/wiki/Unicore)
* `x86_64` - [64-bit extension of IA-32](https://en.wikipedia.org/wiki/X86)
* `xtensa` - [Tensilica Xtensa configurable microprocessor core](https://en.wikipedia.org/wiki/Xtensa#Processor_Cores)
* `xtensaeb` - [Tensilica Xtensa configurable microprocessor core](https://en.wikipedia.org/wiki/Xtensa#Processor_Cores) (Big Endian)

***

### instance\_uuid

**Property**: `instance_uuid`

**Type**: string

**Description:**

For snapshot images, this is the UUID of the VM used to create this image. The value must be a valid VM UUID.

***

### img\_config\_drive

**Property:** `img_config_drive`

**Type:** string

**Description:**

Specifies whether the image needs a config drive.

**Supported values:**

* `mandatory`
* `optional` (default if property is not used)

***

### os\_admin\_user

**Property:** `os_admin_user`

**Type:** string

**Description:**

The name of the user with admin privileges. The value must be a valid username (defaults to `root` for Linux guests and `Administrator` for Windows guests).

***

### os\_distro

**Property:** `os_distro`

**Type:** string

**Description:**

The common name of the operating system distribution in lowercase (uses the same data vocabulary as the [libosinfo project](http://libosinfo.org/)). Specify only a recognized value for this field. Deprecated values are listed to assist you in searching for the recognized value.

**Supported values:**

* `arch` - Arch Linux. Do not use `archlinux` or `org.archlinux`.
* `centos` - Community Enterprise Operating System. Do not use `org.centos` or `CentOS`.
* `debian` - Debian. Do not use `Debian` or `org.debian`
* `fedora` - Fedora. Do not use `Fedora`, `org.fedora`, or `org.fedoraproject`.
* `freebsd` - FreeBSD. Do not use `org.freebsd`, `freeBSD`, or `FreeBSD`.
* `gentoo` - Gentoo Linux. Do not use `Gentoo` or `org.gentoo`.
* `mandrake` - Mandrakelinux (MandrakeSoft) distribution. Do not use `mandrakelinux` or `MandrakeLinux`.
* `mandriva` - Mandriva Linux. Do not use `mandrivalinux`.
* `mes` - Mandriva Enterprise Server. Do not use `mandrivaent` or `mandrivaES`.
* `msdos` - Microsoft Disc Operating System. Do not use `ms-dos`.
* `netbsd` - NetBSD. Do not use `NetBSD` or `org.netbsd`.
* `netware` - Novell NetWare. Do not use `novell` or `NetWare`.
* `openbsd` - OpenBSD. Do not use `OpenBSD` or `org.openbsd`.
* `opensolaris` - OpenSolaris. Do not use `OpenSolaris` or `org.opensolaris`.
* `opensuse` - openSUSE. Do not use `suse`, `SuSE`, or \`\` org.opensuse\`\`.
* `rhel` - Red Hat Enterprise Linux. Do not use `redhat`, `RedHat`, or `com.redhat`.
* `sled` - SUSE Linux Enterprise Desktop. Do not use `com.suse`.
* `ubuntu` - Ubuntu. Do not use `Ubuntu`, `com.ubuntu`, `org.ubuntu`, or `canonical`.
* `windows` - Microsoft Windows. Do not use `com.microsoft.server` or `windoze`.

***

### os\_version

**Property:** `os_version`

**Type:** string

**Description:**

The operating system version as specified by the distributor.

The value must be a valid version number (for example, `11.10`).

***

### os\_secure\_boot

**Property:** `os_secure_boot`

**Type:** string

**Description:**

Secure Boot is a security standard. When the VM starts, Secure Boot first examines software such as firmware and OS by their signature and only allows them to run if the signatures are valid.

Linux guests will require bootloader’s digital signature provided as `os_secure_boot_signature` and `hypervisor_version_requires>=10.0` image properties.

**Supported values:**

* `required` - Enable the Secure Boot feature.
* `disabled` or `optional` - (default if property not used) Disable the Secure Boot feature.

***

### os\_shutdown\_timeout

**Property:** `os_shutdown_timeout`

**Type:** integer

**Description:**

By default, guests will be given 60 seconds to perform a graceful shutdown. After that, the VM is powered off. This property allows overriding the amount of time (unit: seconds) to allow a guest OS to cleanly shut down before power off. A value of 0 (zero) means the guest will be powered off immediately with no opportunity for guest OS clean-up.

***

### trait:\<trait\_name>

**Property:** `trait:<trait_name>`

**Type:** string

**Description:**

Traits allow specifying a server to build on a compute node with the set of traits specified in the image. The traits are associated with the resource provider that represents the compute node in the Placement API.

The syntax of specifying traits is `trait:<trait_name>=value`, for example:

`trait:HW_CPU_X86_AVX2=required`

`trait:STORAGE_DISK_SSD=required`

The compute service scheduler will pass required traits specified on the image to the compute placement API to include only resource providers that can satisfy the required traits. Traits for the resource providers can be managed using the osc-placement plugin.

Image traits are used by the compute service scheduler even in cases of volume backed VMs, if the volume source is an image with traits.

**Supported values:**

* `required` - `<trait_name>` is required on the resource provider that represents the hypervisor node on which the image is launched. This is the only valid value; any other value is invalid.

***

### vm\_mode

**Property:** `vm_mode`

**Type:** string

**Description:**

The virtual machine mode. This represents the host/guest ABI (application binary interface) used for the virtual machine.

**Supported values:**

* `hvm` - Fully virtualized. This is the mode used by QEMU and KVM.
* `xen` - Xen 3.0 paravirtualized.
* `uml` - User Mode Linux paravirtualized.
* `exe` - Executables in containers. This is the mode used by LXC.

***

### hw\_cpu\_sockets

**Property:** `hw_cpu_sockets`

**Type:** integer

**Description:**

The preferred number of sockets to expose to the guest.

***

### hw\_cpu\_cores

**Property:** `hw_cpu_cores`

**Type:** integer

**Description:**

The preferred number of cores to expose to the guest.

***

### hw\_cpu\_threads

**Property:** `hw_cpu_threads`

**Type:** integer

**Description:**

The preferred number of threads to expose to the guest.

***

### hw\_cpu\_policy

**Property:** `hw_cpu_policy`

**Type:** string

**Description:**

Used to pin the virtual CPUs (vCPUs) of VMs to the host's physical CPU cores (pCPUs). Host aggregates should be used to separate these pinned VMs from unpinned VMs as the later will not respect the resourcing requirements of the former.

**Supported values:**

* `shared` - (default if property not specified) The guest vCPUs will be allowed to freely float across host pCPUs, albeit potentially constrained by NUMA policy.
* `dedicated` - The guest vCPUs will be strictly pinned to a set of host pCPUs. In the absence of an explicit vCPU topology request, the drivers typically expose all vCPUs as sockets with one core and one thread. When strict CPU pinning is in effect the guest CPU topology will be setup to match the topology of the CPUs to which it is pinned. This option implies an overcommit ratio of 1.0. For example, if a two vCPU guest is pinned to a single host core with two threads, then the guest will get a topology of one socket, one core, two threads.

***

### hw\_cpu\_thread\_policy

**Property:** `hw_cpu_thread_policy`

**Type:** string

**Description:**

Further refines `hw_cpu_policy=dedicated` by stating how hardware CPU threads in a simultaneous multithreading-based (SMT) architecture be used. SMT-based architectures include Intel processors with Hyper-Threading technology. In these architectures, processor cores share a number of components with one or more other cores. Cores in such architectures are commonly referred to as hardware threads, while the cores that a given core share components with are known as thread siblings.

**Supported values:**

* `prefer` - (default if property not specified) The host may or may not have an SMT architecture. Where an SMT architecture is present, thread siblings are preferred.
* `isolate` - The host must not have an SMT architecture or must emulate a non-SMT architecture. If the host does not have an SMT architecture, each vCPU is placed on a different core as expected. If the host does have an SMT architecture - that is, one or more cores have thread siblings - then each vCPU is placed on a different physical core. No vCPUs from other guests are placed on the same core. All but one thread sibling on each utilized core is therefore guaranteed to be unusable.
* `require` - The host must have an SMT architecture. Each vCPU is allocated on thread siblings. If the host does not have an SMT architecture, then it is not used. If the host has an SMT architecture, but not enough cores with free thread siblings are available, then scheduling fails.

***

### hw\_cdrom\_bus

**Property:** `hw_cdrom_bus`

**Type:** string

**Description:**

Specifies the type of disk controller to attach CD-ROM devices to. As for `hw_disk_bus`.

***

### hw\_disk\_bus

**Property:** `hw_disk_bus`

**Type:** string

**Description:**

Specifies the type of disk controller to attach disk devices to.

**Supported values:**

`scsi`, `virtio`, `uml`, `xen`, `ide`, `usb`, or `lxc`.

***

### hw\_firmware\_type

**Property:** `hw_firmware_type`

**Type:** string

**Description:**

Specifies the type of firmware with which to boot the guest.

**Supported values:**

* `bios`
* `uefi`

***

### hw\_mem\_encryption

**Property:** `hw_mem_encryption`

**Type:** boolean

**Description:**

Enables encryption of guest memory at the hardware level, if there are compute hosts available which support this. See compute service documentation on configuration of the KVM hypervisor for more details.

***

### hw\_pointer\_model

**Property:** `hw_pointer_model`

**Type:** string

**Description:**

Input devices that allow interaction with a graphical framebuffer, for example to provide a graphic tablet for absolute cursor movement. Currently only supported by the KVM/QEMU hypervisor configuration and VNC or SPICE consoles must be enabled.

**Supported values:**

* `usbtablet`

***

### hw\_rng\_model

**Property:** `hw_rng_model`

**Type:** string

**Description:**

Adds a random-number generator device to the image's VMs. This image property by itself does not guarantee that a hardware RNG will be used; it expresses a preference that may or may not be satisfied depending upon compute service configuration.

The cloud administrator can enable and control device behavior by configuring the VM's flavor. By default:

The generator device is disabled.

`/dev/urandom` is used as the default entropy source. To specify a physical hardware RNG device, use the following option in the `nova.conf` file:

`rng_dev_path=/dev/hwrng`

The use of a hardware random number generator must be configured in a flavor's extra\_specs by setting `hw_rng:allowed` to `True` in the flavor definition.

**Supported values:**

* `virtio`
* Other supported device.

***

### hw\_time\_hpet

**Property:** `hw_time_hpet`

**Type:** boolean

**Description:**

Adds support for the High Precision Event Timer (HPET) for x86 guests in the libvirt driver when `hypervisor_type=qemu` and `architecture=i686` or `architecture=x86_64`. The timer can be enabled by setting `hw_time_hpet=true`. By default HPET remains disabled.

***

### hw\_machine\_type

**Property:** `hw_machine_type`

**Type:** string

**Description:**

Enables booting an ARM system using the specified machine type. If an ARM image is used and its machine type is not explicitly specified, then Compute uses the `virt` machine type as the default for ARMv7 and AArch64. Valid types can be viewed by using the `virsh capabilities` command (machine types are displayed in the `machine` tag).

***

### os\_type

**Property:** `os_type`

**Type:** string

**Description:**

The operating system installed on the image. The libvirt API driver contains logic that takes different actions depending on the value of the `os_type` parameter of the image. For example, for `os_type=windows` images, it creates a FAT32-based swap partition instead of a Linux swap partition, and it limits the injected host name to less than 16 characters.

**Supported values:**

* `linux`
* `windows`

***

### hw\_scsi\_model

**Property:** `hw_scsi_model`

**Type:** string

**Description:**

Enables the use of VirtIO SCSI (virtio-scsi) to provide block device access for VMs; by default, VMs use VirtIO Block (virtio-blk). VirtIO SCSI is a para-virtualized SCSI controller device that provides improved scalability and performance, and supports advanced SCSI hardware.

**Supported values:**

* `virtio-scsi`

***

### hw\_serial\_port\_count

**Property:** `hw_serial_port_count`

**Type:** integer

**Description:**

Specifies the count of serial ports that should be provided. If `hw:serial_port_count` is not set in the flavor's extra\_specs, then any count is permitted. If `hw:serial_port_count` is set, then this provides the default serial port count. It is permitted to override the default serial port count, but only with a lower value.

***

### hw\_video\_model

**Property:** `hw_video_model`

**Type:** string

**Description:**

The graphic device model presented to the guest. `none` disables the graphics device in the guest and should generally be used when using GPU passthrough.

**Supported values:**

* `vga`
* `cirrus`
* `vmvga`
* `xen`
* `qxl`
* `virtio`
* `gop`
* `none`
* `bochs`

***

### hw\_video\_ram

**Property:** `hw_video_ram`

**Type:** integer

**Description:**

Maximum RAM in MB for the video image. Used only if a `hw_video:ram_max_mb` value has been set in the flavor's extra\_specs and that value is higher than the value set in `hw_video_ram`

***

### hw\_watchdog\_action

**Property:** `hw_watchdog_action`

**Type:** string

**Description:**

Enables a virtual hardware watchdog device that carries out the specified action if the server hangs. The watchdog uses the i6300esb device (emulating a PCI Intel 6300ESB). If `hw_watchdog_action` is not specified, the watchdog is disabled.

**Supported values:**

* `disabled` - (default) The device is not attached. Allows the user to disable the watchdog for the image, even if it has been enabled using the image's flavor.
* `reset` - Forcefully reset the guest.
* `poweroff` - Forcefully power off the guest.
* `pause` - Pause the guest.
* `none` - Only enable the watchdog; do nothing if the server hangs.

***

### hw\_vif\_model

**Property:** `hw_vif_model`

**Type:** string

**Description:**

Specifies the model of virtual network interface device to use.

Options:

* `e1000`, `e1000e`, `ne2k_pci`, `pcnet`, `rtl8139`, `virtio`, and `vmxnet3`.

***

### hw\_vif\_multiqueue\_enabled

**Property:** `hw_vif_multiqueue_enabled`

**Type:** boolean

**Description:**

If true, this enables the virtio-net multiqueue feature. In this case, the driver sets the number of queues equal to the number of guest vCPUs. This makes the network performance scale across a number of vCPUs.

***

### hw\_boot\_menu

**Property:** `hw_boot_menu`

**Type:** boolean

**Description:**

If true, enables the BIOS bootmenu. In cases where both the image metadata and Extra Spec are set, the Extra Spec setting is used. This allows for flexibility in setting/overriding the default behavior as needed.

***

### hw\_pmu

**Property:** `hw_pmu`

**Type:** boolean

**Description:**

Controls emulation of a virtual performance monitoring unit (vPMU) in the guest. To reduce latency in realtime workloads disable the vPMU by setting `hw_pmu=false`.


# Get Images

This document offers links to download virtual machine image files for popular operating systems.

## CirrOS

CirrOS is a minimal Linux distribution that was designed for use as a test image.

* Here's a link to [CirrOS download page](https://download.cirros-cloud.net/) for all available CirrOS images.
* Download the 64-bit qcow2 image for CirrOS verion 0.6.2 here - [cirros-0.6.2-x86\_64-disk.img](https://download.cirros-cloud.net/0.6.2/cirros-0.6.2-x86_64-disk.img)

## Red Hat Enterprise Linux

Red Hat maintains official Red Hat Enterprise Linux cloud images. A valid Red Hat Enterprise Linux subscription is required to download these images.

* [Red Hat Enterprise Linux 7 KVM Guest Image](https://access.redhat.com/downloads/content/69/ver=/rhel---7/x86_64/product-downloads)
* [Red Hat Enterprise Linux 8 KVM Guest Image](https://access.redhat.com/downloads/content/479/ver=/rhel---8/x86_64/product-downloads)
* [Red Hat Enterprise Linux 9 KVM Guest Image](https://access.redhat.com/downloads/content/479/ver=/rhel---9/x86_64/product-downloads)

Note - In a RHEL cloud image, the login account is `cloud-user`.

## Ubuntu

* Official site for Ubuntu based images maintained by Canonical - [Ubuntu-based images](https://cloud-images.ubuntu.com/).
* Download qcow2 image for Ubuntu 24.04 from here - [noble-server-cloudimg-amd64.img](https://cloud-images.ubuntu.com/noble/current/noble-server-cloudimg-amd64.img).


# Image Library High Availability

You can create a highly available deployment of <code class="expression">space.vars.product\_name</code> image library service in a region by enabling image library role on multiple hosts. This feature requires the use of shared storage backend for image library storage. Once configured, this feature enables seamless access to images across host failures or maintenance events for hosts assigned with image library role.

## Prerequisites

Following are the pre-requisites for enabling high availability for image library service:

Networking and firewall rules allow image-related traffic between image library hosts and all other hosts configured with rest of the

services.

* **A shared storage backend is required.** A shared storage backend (e.g., NFS or another supported block storage volume backend) is a mandatory prerequisite for enabling high availability (HA) of the Image Library Service. Assigning the Image Library role to multiple hosts without using a shared storage backend is not supported may lead to issues such as inconsistent image discovery, image deletion failures, or orphaned image data.
* **Shared storage must be accessible by all Image Library hosts.** The shared backend must be mounted and available to all hosts that will be assigned the Image Library role with the same path.
* **All Image Library hosts must be in the same region.** Cross-region image library deployments are not supported. All hosts assigned the Image Library role must reside within the same region.
* **Network connectivity must allow image-related traffic.** Firewall and network policies must permit image-related traffic between all Image Library hosts and all other hosts configured with rest of the <code class="expression">space.vars.product\_name</code> services.

{% hint style="info" %}
**Info**

A shared storage backend is a mandatory requirement for enabling high availability for Image Library Service. Assigning Image Library role to multiple hosts without using shared storage backend may result in issues with image deletion or discovery.
{% endhint %}

{% hint style="info" %}
**Info**

The ability to assign the Image Library role to multiple hosts without using a shared storage backend is deprecated and will be removed in a future release. All new and existing deployments should migrate to a shared storage backend to ensure compatibility and continued support.
{% endhint %}

## Supported Storage Backends

Following table describes the supported and unsupported backends for configuring high availability for image library service.

| **Backend Type**       | **HA Support**                                                                                                                                                                                                      |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| File-based (e.g., NFS) | Supported, but it must be mounted on **every image library host at the same mount point/path**. This gives consistency between image library hosts to prevent corruption, deletion mismatches, or discovery issues. |
| Block Storage Volume   | Recommended for production because it offers **better scalability, reliability and image transfer performance** compared to file-based storage. Still needs to be accessible by all image library hosts.            |

## How High Availability Works in Image Library

* When images are stored on shared storage, any image library host can serve them for VM or volume creation.
* During image creation, <code class="expression">space.vars.product\_name</code> dynamically selects an available and healthy image library host. If one of the image library hosts is offline, the system transparently retries with another active image library host to create the image, ensuring uninterrupted service.

## Deployment Steps

### Step 1: Configure Shared Storage Backend

The image library can be configured to use block storage as it's backend. You can do this by providing the name of the volume type for the block storage backend you'd like to use while specifying image library location as part of the cluster blueprint configuration. This informs the image library service to use block storage as the persistent backend to save and retrieve images.

Alternatively, if using file-based shared storage (e.g., NFS), it must be mounted or attached on each host where the image library role will be enabled with the exact same path\*\*.\*\*

{% tabs %}
{% tab title="Bash" %}

```bash
mount -t nfs <NFS_SERVER>:_<EXPORTED_PATH> _var_opt_imagelibrary_data
ls -l _var_opt_imagelibrary_data
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
**Info**

When using file-based shared storage, each image library host must mount the storage at the exact same path location.
{% endhint %}

### Step 2: Enable Image Library Role on Multiple Hosts

1. Navigate to Infrastructure > Hosts in the UI.
2. Select the target host.
3. Click Edit Roles and assign the Image Library role.
4. Repeat this for all hosts that should be part of your highly available image library service setup.

### Step 3: Validate that image library service is running

Run the following command on each host enabled with the image library role:

{% tabs %}
{% tab title="Bash" %}

```yaml
systemctl status pf9-glance-api
```

{% endtab %}
{% endtabs %}

Check that:

* The pf9-glance-api service is active.
* No errors are reported in `/var/log/pf9/glance-api.log`.

You can also validate from the UI by checking the Settings > API Access > API Endpoints and check that `image-cluster` service is available with multiple image library endpoints.

## Image Library Admin Endpoint

Read more about [Image Library Admin Endpoint here.](/private-cloud-director/images-and-image-library/image-library---images#image-library-admin-endpoint) In case of highly available image library setup with multiple hosts having image library role assigned, the last host to get the image library service role assigned is selected to be the admin endpoint.

The admin endpoint is primarily used to upload images to the image library. When creating a new virtual machine, the compute and the block storage service are configured to round robin across all available image library hosts to fetch the required image. To restrict image fetches to a single cluster, see [Restrict the Image Library Service to a Specific Cluster](/private-cloud-director/images-and-image-library/restrict-image-library-to-cluster).

If an image library host that is also acting as an admin endpoint goes down, the admin endpoint is not automatically assigned to one of the other surviving image library hosts today. You will need to manually change the admin endpoint to a different image library host (by following the steps below).

Note that this limitation only impacts your ability to upload new images to the image library. It does not impact new virtual machine provisioning. The admin endpoint is only used to upload new images to the image library. When creating a new VM, the compute and block storage service are designed to use any of the available image library hosts in a round robin fashion the fetch the virtual machine image.

{% hint style="danger" %}
**Important**

If an image library host that is also acting as an admin endpoint goes down, the admin endpoint is not automatically assigned to one of the other surviving image library hosts today. You will need to manually change the admin endpoint to a different image library host (by following the steps below). This is required so you can continue to upload images to the image library service.
{% endhint %}

You can manually configure or change the image library admin endpoint assignment by running the following `pcdctl` command:

### Step 1 - Get the admin endpoint UUID

Run the following command to get UUID of the admin endpoint. This command will list the ID of the current admin endpoint. Copy it.

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl endpoint list --service glance --interface admin
```

{% endtab %}
{% endtabs %}

Example output:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl endpoint list --service glance --interface admin
+----------------------------------+--------+--------------+--------------+---------+-----------+----------------------------+
| ID                               | Region | Service Name | Service Type | Enabled | Interface | URL                        |
+----------------------------------+--------+--------------+--------------+---------+-----------+----------------------------+
| 4bf27ff9f8a146d59dcce04bcedb7mz0 | SJC    | glance       | image        | True    | admin     | https:__111.11.33.138:9494 |
+----------------------------------+--------+--------------+--------------+---------+-----------+----------------------------+
```

{% endtab %}
{% endtabs %}

### Step 2 - Set the Admin Endpoint

Run the following command to set the new admin endpoint. Replace with the ip address of your alternate image library host. Replace with the ID that you copied from the command above.

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl endpoint set --url https:__<IP>:9494 <UUID>
```

{% endtab %}
{% endtabs %}

Using our previous example, and say the IP address of your second image service host is 111.11.33.139, then the following command will set this host to be the image library service admin endpoint:

{% tabs %}
{% tab title="Bash" %}

```bash
pcdctl endpoint set --url https:__111.11.33.139:9494 4bf27ff9f8a146d59dcce04bcedb7mz0
```

{% endtab %}
{% endtabs %}


# Restrict the Image Library Service to a Specific Cluster

## Overview

By default, the Image Library Service operates at the region level: the Compute Service and the Persistent Storage Service round robin across all Image Library hosts in the region to fetch an image. In a multi-cluster deployment, you may want the hosts in a cluster to use only that cluster's Image Library host, for isolation and better performance. <code class="expression">space.vars.product\_name</code> does not expose a blueprint-level setting for this, so you restrict it with a per-host configuration override.

In this guide, you will apply a configuration override on the compute and persistent storage hosts in a cluster so that they fetch images only from that cluster's Image Library host.

{% hint style="danger" %}
**Important**

This is an advanced operation that should only be performed by Administrators, ideally under guidance from the Platform9 support or solution architect teams.
{% endhint %}

## Before You Begin

* Apply the override on every compute host and every persistent storage host in the target cluster. If you miss a host, it continues to fetch images from any Image Library host in the region.
* Changes persist across upgrades and require a service restart to take effect.
* Identify the cluster-local Image Library endpoint first, in the form `https://<image-library-host>:9494`. Use `https://localhost:9494` only when the host you are editing also has the Image Library role assigned.

## Apply the Override

On each compute host and each persistent storage host in the cluster, add the following block to the service's override file:

{% tabs %}
{% tab title="Bash" %}

```bash
[glance]
endpoint_override = https://localhost:9494

[pf9_glance]
glance_cluster = false
```

{% endtab %}
{% endtabs %}

* `endpoint_override` sets the single Image Library endpoint this host uses. Replace `https://localhost:9494` with the cluster-local endpoint (`https://<image-library-host>:9494`) on any host that does not also hold the Image Library role.
* `glance_cluster = false` disables the default region-wide behavior, so the service does not round robin to other Image Library hosts.

The Compute Service reads this override when booting a VM from an image, and the Persistent Storage Service reads it when creating a volume from an image. Add the block to both override files and restart the corresponding service:

| Service            | Override file                                                    | Restart command                              |
| ------------------ | ---------------------------------------------------------------- | -------------------------------------------- |
| Compute            | `/opt/pf9/etc/nova/conf.d/nova_override.conf`                    | `systemctl restart pf9-ostackhost`           |
| Persistent Storage | `/opt/pf9/etc/pf9-cindervolume-base/conf.d/cinder_override.conf` | `sudo service pf9-cindervolume-base restart` |

## Verify the Configuration

Create a test VM, and a test volume from an image, in the cluster. Confirm that both are served by the cluster-local Image Library host, for example by checking `/var/log/pf9/glance-api.log` on that host.


# Troubleshooting And Log Files

## Troubleshoot Image Library Issues

If the Service Health dashboard shows that the Image Library Service is unhealthy, the issue may be caused by one of the following:

* You do not have certificate authorization to access the Image Library Service — see [Image Library Service Certificate Configuration](/private-cloud-director/images-and-image-library/image-library-certificate-configuration).
* The Image Library Service is not responding on one or more hosts configured with the image library role — see [Image Library Service Endpoint Health](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide#image-library-service-endpoint-health).
* An image upload is stuck in `queued` status — see [Troubleshoot Image Upload and Queued Status](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide#troubleshoot-image-upload-and-queued-status).

To debug the issue, check the log files on the affected hosts.

### Log Files

The Image Library Service logs are located at:

`/var/log/pf9/glance-api.log`

Audit logs (upload outcomes, user identity, image UUID) are at:

`/var/log/pf9/glance-audit.log`

## Troubleshooting Guides

* [Image Service Troubleshooting Guide](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-service-troubleshooting-guide) — general upload flow, queued-status diagnosis, endpoint health, and bulk-deploy caching guidance.
* [Image Creation Failed Using CLI](/private-cloud-director/images-and-image-library/troubleshooting-and-log-files/image-creation-failed-using-cli) — step-by-step resolution for CLI upload failures.


# Image Creation Failed Using CLI

## Problem

> When image creation failures occur during VM image uploads through the OpenStack or [pcdctl](https://platform9.com/docs/private-cloud-director/private-cloud-director/pcdctl-command-line) command-line tool, this troubleshooting guide provides resolution steps.

## Environment

* Private Cloud Director - v2025.4 and Higher.
* Self-Hosted Private Cloud Director Virtualisation – v2025.4 and Higher.

## Procedure

{% stepper %}
{% step %}

#### Prerequisites

Ensure the the [prerequisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/pre-requisites#image-library-prerequisites) are met. Also refer to the documentation about [importing image via CLI](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#importing-images-via-the-cli).
{% endstep %}

{% step %}

#### Source OpenStack credentials

Source the OpenStack `admin.rc` file. Ensure that `OS_USERNAME` and `OS_PASSWORD` are correct as per your <code class="expression">space.vars.product\_name</code> account.
{% endstep %}

{% step %}

#### Verify permissions

Ensure you have `administrator` permission to upload the image.
{% endstep %}

{% step %}

#### Set OS\_INTERFACE

Make sure that the `OS_INTERFACE` variable is set to the `admin` value as specified in the UI as mentioned in the [documentation](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#step-3---upload-the-image).
{% endstep %}

{% step %}

#### Run OpenStack CLI in debug/verbose mode

Use command debug mode during image creation to get detailed output.

{% tabs %}
{% tab title="Command" %}
{% code title="openstack debug" %}

```bash
$ openstack --debug image create 
or 
$ openstack --verbose image create
```

{% endcode %}
{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Review glance logs (Self-Hosted PCD)

Review glance server logs in the `glance-api` pod logs for Self-Hosted PCD. On the host, check `/var/log/pf9/glance-api.log` to track relevant events against a specific image ID.
{% endstep %}

{% step %}

#### Get glance endpoints

Get the glance API public endpoint:

{% tabs %}
{% tab title="Command" %}
{% code title="list glance endpoints" %}

```bash
$ openstack endpoint list --service glance
$ openstack endpoint list --service glance-cluster
```

{% endcode %}
{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Verify glance API

Verify the glance API service by connecting to the glance endpoint. It will return the version, status and other details.

{% tabs %}
{% tab title="Command" %}
{% code title="curl glance" %}

```bash
$ curl -s https://<PCD_FQDN>/glance/
```

{% endcode %}
{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Contact Support

If these steps do not resolve the issue, reach out to the [Platform9 Support Team](https://support.platform9.com/) for additional assistance.
{% endstep %}
{% endstepper %}

## Most common causes

* Glance service on the underlying host is down.
* Glance host is unreachable from the VM from where the upload is being performed.
* OpenStack `admin.rc` file does not have the `OS_INTERFACE` variable set to the `admin`.
* The `--insecure` flag was not used as the Glance node uses self-signed certificates.
* The user does not have sufficient permissions to perform the image upload.
* Ensure that the glance [Pre-Requisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#prerequisites) are met.


# Image Service Troubleshooting Guide

## Problem

A troubleshooting guide for image services is needed to address frequent issues with image management in cloud environments, such as image upload failures, slow performance, and incorrect metadata. The guide must provide clear, actionable steps for diagnosing and resolving common errors to ensure the reliability and availability of the image service.

## Environment

* Private Cloud Director Virtualization - v2025.4 and Higher
* Self-Hosted Private Cloud Director Virtualization - v2025.4 and Higher
* Component - PCD Image service

## Deep Dive

### Image Creation Flow

The image creation process in `space.vars.product_name` is managed by the **Glance** service. The flow begins when a user uploads a new image file, which is then processed and stored.

{% stepper %}
{% step %}

#### User Request

A user initiates an image upload via the OpenStack CLI, `space.vars.product_name` dashboard, or direct API call. The request includes the image file and metadata (e.g., name, format, disk format).
{% endstep %}

{% step %}

#### API Service

The **Glance API** service receives the request, validates the user's authentication token with **Keystone**, and checks for permissions and quotas. Below Glance API logs show the token is being used.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO glance.api.v2.image_data [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] Unable to create trust: no such option collect_timing in group [keystone_authtoken] Use the existing user token.
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Image Validation and Upload Request

The glance service validates the [image format](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#image-formats) and confirms that its virtual size (can be fetched using the command `qemu-img info <image_name>.qcow2` on glance host) meets the requirements. Then the Image `PUT /v2/images/<IMAGE_UUID>/file` request is placed to upload image data to a temporary staging area, ideally at the default staging location (If default directory changed, then check the custom location) `/var/lib/glance/os_glance_staging_store/`.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO glance.location [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] Image format matched and virtual size computed: 41126400
INFO eventlet.wsgi.server [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] 127.0.0.1 - - [..] "PUT /v2/images/[IMAGE_UUID]/file HTTP/1.0" 204 468 2.400140
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Glance API to Registry

The API service then communicates with the **Glance Registry**, which creates a new entry for the image in the Glance database. The status is set to `queued`.
{% endstep %}

{% step %}

#### Glance Service

The Glance API hands off the request to the **pf9-glance-api** service, which moves the image data from the staging area to the backend storage (e.g., Swift, Ceph, or a local file system) default image file storage location `/var/opt/imagelibrary/data/glance/`.
{% endstep %}

{% step %}

#### Status Update

Once the image is successfully stored, the Glance Store service updates the image's status in the database from `queued` to `active`. The image is now ready for use. The host glance audit logs (`/var/log/pf9/glance-audit.log`) show information about the request Username, Image UUID, outcome, etc.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO oslo.messaging.notification.audit.http.response [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] {"message_id": "[Audit_Message_ID]", "publisher_id": "glance-api", "event_type": "audit.http.response", "priority": "INFO", "payload": {"typeURI": "http://schemas.dmtf.org/cloud/audit/1.0/event", "eventType": "activity", "id": "[Activity_ID]", "eventTime": "[..]", "action": "update", "outcome": "success", "observer": {"id": "target"}, "initiator": {"id": "[..]", "typeURI": "service/security/account/user", "name": "[USER_NAME]", "credential": {"token": "***", "identity_status": "Confirmed"}, "host": {"address": "127.0.0.1", "agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/139.0.0.0 Safari/537.36"}, "project_id": "[PROJECT_UUID]"}, "target": {"id": "unknown", "typeURI": "unknown", "name": "unknown"}, "requestPath": "/v2/images/[IMAGE_UUID]/file", "tags": ["correlation_id?value=[..]"], "reason": {"reasonType": "HTTP", "reasonCode": "204"}, "reporterchain": [{"role": "modifier", "reporterTime": "[..]", "reporter": {"id": "target"}}]}, "timestamp": "[..]"}
```

{% endtab %}
{% endtabs %}
{% endstep %}
{% endstepper %}

***

### Image Deletion Flow

The deletion process also uses the Glance services to remove the image's data and its database entry.

{% stepper %}
{% step %}

#### User Request

A user sends a deletion request via the OpenStack CLI, `space.vars.product_name` dashboard, or direct API call. The request includes the image information (e.g Image ID or Name).
{% endstep %}

{% step %}

#### API Service

The **Glance API** service receives the request, validates the user's authentication token with **Keystone**, and performs a permission check, and changes the image's status in the database to `deleting`. Below Glance API logs show the token is being used.

{% tabs %}
{% tab title="Sample Logs" %}

```dart
INFO glance.api.v2.image_data [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] Unable to create trust: no such option collect_timing in group [keystone_authtoken] Use the existing user token.
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Glance Service

The **pf9-glance-api** service receives a message to delete the image from the backend storage. The Glance API hands off the request to the **pf9-glance-api** service, which moves the image data from backend storage (e.g., Swift, Ceph, or a local file system) default image file storage location `/var/opt/imagelibrary/data/glance/`. This is a crucial step that frees up disk space.

{% tabs %}
{% tab title="Sample logs" %}

```dart
INFO eventlet.wsgi.server [None [REQ-ID] [USER_ID] [TENANT_ID] - - default default] 127.0.0.1 - - [..] "DELETE /v2/images/[IMAGE_UUID] HTTP/1.0" 204 468 1.924124
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Final Status Update

Once the data is confirmed to be deleted from the backend store, the Glance API removes the image's database entry, completing the deletion process.
{% endstep %}
{% endstepper %}

## Procedure

The following steps outline how to troubleshoot the image issue.

{% stepper %}
{% step %}

#### Review image details

Review image details like status and any errors.

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack image show <IMAGE_UUID>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Validate Glance endpoints

Validate if the glance image endpoints are available and the public endpoint is responding using a curl request. This curl request should return the glance information.

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack endpoint list --service glance
$ openstack endpoint list --service glance-cluster
$ curl -s https://<FQDN>/glance/
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Check if the image service is enabled

{% tabs %}
{% tab title="Command" %}

```bash
$ openstack service list | grep -i image
$ openstack service show <GLANCE/GLANCE_CLUSTER_UUID>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Check glance-api pod (Self-Hosted only)

The management plane has a **glance-api** pod to provide the image service. Check if the glance-api pod is running in the workload region namespace. Review this pod:

{% hint style="info" %}
**Info**

Step 4 is applicable only for Self-Hosted Private Cloud Director
{% endhint %}

* Check if they are in "CrashLoopBackOff/OOMkilled/Pending/Error/Init" state.
* Also, verify if all containers in the pods are Running.
* See the events section in pod describe output.
* Review pods logs using `REQ_ID` or `VM_UUID` for relevant details.

{% tabs %}
{% tab title="Command" %}

```bash
$ kubectl get pods -o wide -n <WORKLOAD_REGION> | grep -i "glance"

$ kubectl describe -n <WORKLOAD_REGION> <GLANCE_API_POD>

$ kubectl logs -n <WORKLOAD_REGION> <GLANCE_API_POD>
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Validate pf9-glance-api service on host

Validate if the **pf9-glance-api** service is running on the host where glance role is applied.

{% tabs %}
{% tab title="Command" %}

```bash
$ sudo systemctl status pf9-glance-api
```

{% endtab %}
{% endtabs %}
{% endstep %}

{% step %}

#### Review host logs

On the host, review the `/var/log/pf9/glance-api.log` to track the relevant events against a specific image ID.
{% endstep %}

{% step %}

#### Escalation

If these steps prove insufficient to resolve the issue, kindly reach out to the [Platform9 Support Team](https://support.platform9.com/hc/en-us) for additional assistance.
{% endstep %}
{% endstepper %}

## Most Common Causes

* Ensure that the Image Library Service [Pre-Requisites](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#prerequisites) are met.
* While uploading an image, the `admin.rc` file does not have the `OS_INTERFACE` variable set to `admin`.
* Incorrect image format. See [Supported Image Formats](https://platform9.com/docs/private-cloud-director/private-cloud-director/image-library---images#image-formats).
* The `pf9-glance-api` service is down on the Image Library host.
* In Self-Hosted deployments, the `--insecure` flag was not used with the `pcdctl` command and the Image Library host uses a self-signed certificate.
* Uploading an image to a volume whose type is both encrypted and backed by an NFS storage backend — see [Diagnose Image Upload Failures on Encrypted NFS-Backed Volumes](/private-cloud-director/storage/troubleshooting-and-log-files/troubleshooting-cinder-issues#diagnose-image-upload-failures-on-encrypted-nfs-backed-volumes).

***

## Troubleshoot Image Upload and Queued Status <a href="#troubleshoot-image-upload-and-queued-status" id="troubleshoot-image-upload-and-queued-status"></a>

### Overview

When you upload an image, it passes through the following status sequence: `queued` → `saving` → `active`. An image that stays in `queued` indefinitely means the Image Library Service accepted the metadata but was unable to receive or store the file data. The sections below walk through the most common causes and how to resolve them.

### Check Disk Space on the Image Library Host

The Image Library Service writes image data to `/var/opt/imagelibrary/data/glance/` by default. If the filesystem holding that path is full, the upload stalls in `queued` and the service logs an I/O or disk-full error.

1. On the Image Library host, check available disk space:

```bash
df -h /var/opt/imagelibrary/data/glance/
```

2. If the filesystem is at or near 100%, free space by removing unused images or expanding the filesystem before retrying the upload.
3. Check the Image Library Service log for disk-related errors:

```bash
sudo grep -i "no space\|disk full\|errno\|IOError" /var/log/pf9/glance-api.log | tail -50
```

### Verify the Image Library Service Is Running

A stopped or crashed `pf9-glance-api` service will cause all uploads to stall in `queued`.

```bash
sudo systemctl status pf9-glance-api
```

If the service is not active, start it and then check for errors that might have caused it to stop:

```bash
sudo systemctl start pf9-glance-api
sudo journalctl -u pf9-glance-api -n 100
```

### Verify Image Format

An unsupported or mismatched image format can prevent the service from processing the upload. Confirm the actual format of the image file before uploading:

```bash
qemu-img info <image-file>
```

Use the format reported by `qemu-img info` when specifying `--disk-format` in the `pcdctl` command or selecting the format in the UI. Supported formats are `raw`, `qcow2`, and `iso`.

### Check for Network or Timeout Issues

Large image uploads over slow or intermittent network connections can time out before the data fully transfers.

* Upload from a machine that has a direct, low-latency network path to the Image Library host.
* Check for network errors in the Image Library Service log:

```bash
sudo grep -i "timeout\|connection reset\|broken pipe" /var/log/pf9/glance-api.log | tail -50
```

### Recover a Stuck Image

If an image is stuck in `queued` and you have verified that none of the above conditions apply, delete the stuck image record and re-upload:

```bash
pcdctl image delete <IMAGE_UUID>
```

Then retry the upload. If the upload consistently stalls on re-upload, contact [Platform9 Support](https://support.platform9.com/) with the image UUID and the relevant entries from `/var/log/pf9/glance-api.log`.

### Differences Between UI and CLI Upload Failures

| Upload Path                                                                 | What to Check                                                                                                                                                                                                       |
| --------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| UI upload button is inactive                                                | Image Library Service certificate is not trusted in your browser — see [Image Library Service Certificate Configuration](/private-cloud-director/images-and-image-library/image-library-certificate-configuration). |
| UI upload starts but stalls                                                 | Check disk space and service health as described above.                                                                                                                                                             |
| CLI upload returns `SSL: CERTIFICATE_VERIFY_FAILED`                         | Add `--insecure` to the `pcdctl image create` command.                                                                                                                                                              |
| CLI upload returns `Connection refused` or `Unable to establish connection` | The Image Library endpoint is not reachable — see [Image Library Service Endpoint Health](#image-library-service-endpoint-health).                                                                                  |

***

## Image Library Service Endpoint Health <a href="#image-library-service-endpoint-health" id="image-library-service-endpoint-health"></a>

### Overview

When the Image Library Service endpoint is unreachable, image uploads fail and VM provisioning that requires fetching an image from the Image Library host will also fail. The most common symptoms are:

* `pcdctl image create` returns a connection error or HTTP 502 / 503.
* A `curl` request to the Image Library endpoint returns an NGINX error page (502 Bad Gateway, 503 Service Unavailable).
* The **Service Health** dashboard in the UI shows the Image Library Service as unhealthy.

### Check the Image Library Service Endpoints

List the registered Image Library Service endpoints and verify that they are enabled:

```bash
pcdctl endpoint list --service glance
pcdctl endpoint list --service glance-cluster
```

An endpoint with `Enabled: False` will not serve requests. If any endpoint shows as disabled, re-enable it or reassign the Image Library role to the host.

### Test the Endpoint Directly

Use `curl` to check whether the Image Library endpoint responds:

```bash
curl -sk https://<IMAGE_LIBRARY_HOST_IP_OR_FQDN>:9292/
```

* A JSON response that includes version information indicates the endpoint is healthy.
* An NGINX error page (HTML with "502 Bad Gateway" or "503 Service Unavailable") means the NGINX reverse proxy on the Image Library host is running but the upstream `pf9-glance-api` process is not accepting connections.

### Check the Service on the Image Library Host

When the endpoint returns an NGINX error, the `pf9-glance-api` service on the Image Library host has likely stopped or is not listening:

```bash
sudo systemctl status pf9-glance-api
sudo systemctl status pf9-imagelibrary
```

Restart both services if either is not active:

```bash
sudo systemctl restart pf9-glance-api
sudo systemctl restart pf9-imagelibrary
```

After restarting, check the log for startup errors:

```bash
sudo tail -100 /var/log/pf9/glance-api.log
```

### Review Logs

The primary log for Image Library Service issues is on the Image Library host:

```
/var/log/pf9/glance-api.log
```

For audit events (upload outcomes, user identity, image UUID):

```
/var/log/pf9/glance-audit.log
```

{% hint style="info" %}
**Self-Hosted deployments only**

In Self-Hosted deployments, the management-plane `glance-api` pod in the region namespace provides the endpoint that the Image Library host proxies through. If the host-level service appears healthy but the endpoint still returns errors, check the management-plane pod:

```bash
kubectl get pods -o wide -n <REGION_NAMESPACE> | grep glance
kubectl logs -n <REGION_NAMESPACE> <GLANCE_API_POD>
```

In SaaS deployments, the management-plane components are operated by Platform9. If endpoint errors persist after verifying host-level service health, contact [Platform9 Support](https://support.platform9.com/).
{% endhint %}

### Escalation

If service restarts do not resolve the issue and the endpoint continues to return errors, contact [Platform9 Support](https://support.platform9.com/) with:

* Output of `pcdctl endpoint list --service glance`
* Output of `sudo systemctl status pf9-glance-api`
* The last 200 lines of `/var/log/pf9/glance-api.log`

***

## Persistent Storage Backend and Image Caching at Scale <a href="#persistent-storage-backend-and-image-caching-at-scale" id="persistent-storage-backend-and-image-caching-at-scale"></a>

### Overview

The Image Library Service supports two storage backends: a local file store and a Persistent Storage Service (block storage) volume. When the block storage backend is configured, image data is stored on a block storage volume rather than on the Image Library host's local filesystem. This is the recommended configuration for production environments and for Image Library High Availability.

For background on configuring the block storage backend and enabling Image Library High Availability, see [Image Library High Availability](/private-cloud-director/images-and-image-library/image-library-high-availability) and [Block Storage High Availability](/private-cloud-director/storage/block-storage/block-storage-high-availability).

### How Hypervisor Image Caching Works

When a virtual machine is provisioned for the first time using a given image, the hypervisor (compute host) must fetch the image data from the Image Library host over the network. After the first successful boot from that image on a given hypervisor, the hypervisor retains a local copy of the image in its image cache. Subsequent VM deployments using the same image on the same hypervisor use the cached copy and avoid the full image transfer.

This caching behavior has important implications for bulk and parallel VM deployments:

* **First deployment of an image on a hypervisor is the most I/O-intensive.** If many VMs are deployed simultaneously using an image that has never been cached on those hypervisors, all of them will attempt to fetch the full image from the Image Library host at the same time.
* **Image Library host bandwidth is the bottleneck.** The Image Library host (or hosts, in an HA setup) must serve the full image to every hypervisor that does not yet have a cached copy. With many hypervisors requesting simultaneously, this can saturate the Image Library host's network interface and cause VM provisioning to time out or fail.
* **Block storage backend improves throughput.** When the Image Library Service uses a block storage volume as its backend, the image data is fetched from the storage array rather than from the Image Library host's local disk, which can improve parallelism for large deployments.

### Recommendations for Bulk Deployments

When you need to deploy many VMs simultaneously — for example, during a large-scale workload rollout — consider the following:

* **Pre-warm the cache.** Before the bulk deployment, deploy a single VM on each hypervisor using the target image. This populates the cache on all hypervisors. Subsequent bulk deployments will use the cached copy and avoid the simultaneous image-fetch bottleneck.
* **Use Image Library HA with shared storage.** An Image Library deployment with multiple hosts sharing a block storage backend distributes read requests across hosts, reducing the load on any single host during the initial image fetch. See [Image Library High Availability](/private-cloud-director/images-and-image-library/image-library-high-availability).
* **Stage bulk deploys in waves.** If pre-warming is not practical, deploy VMs in smaller batches rather than all at once, allowing each batch to complete its image fetch before the next batch starts.

### Troubleshoot Slow or Failing Bulk Deploys

If VM provisioning fails or times out during a large bulk deployment and image fetching is suspected:

1. Check the Image Library host's network utilization during the deployment window. Sustained high network I/O concurrent with provisioning failures points to an image-fetch bottleneck.
2. Check `/var/log/pf9/glance-api.log` on the Image Library host for errors or timeouts corresponding to the failed provisioning requests.
3. Verify that the Image Library Service endpoint is healthy before the deployment: see [Image Library Service Endpoint Health](#image-library-service-endpoint-health).
4. If the Image Library uses a block storage backend, verify that the storage volume is healthy and that the Persistent Storage Service is operating normally.

If issues persist, contact [Platform9 Support](https://support.platform9.com/) with the Image Library host logs and a description of the deployment scale (number of VMs, number of hypervisors, image size).




---

[Next Page](/llms-full.txt/1)

