Troubleshooting And Log Files
Troubleshooting Compute Service Problems
If your Private Cloud Director Service Health dashboard indicates that the Compute Service is unhealthy, it may be because a large percentage of your hosts with the Hypervisor role assigned are either offline or the Compute service is unresponsive on those hosts. Refer to the Log Filesto debug the issue further.
Important Directories
/var – All logs go under /var/log/pf9. The only exception is the pcdctl log files which go under /var/log
/opt – All the packages for services installed by Private Cloud Director go under /opt/pf9
/var/opt/pf9 - Subdirectories for the Platform9 host agent and networking service go under here with temp files or state files.
Log Files
Essential log files for debugging:
Log files for all services - Each host stores all its log files for the various components running on it at
/var/log/pf9. Here you will find logs for compute, image library, storage, networking, and other services, depending on the roles assigned to that host. See the documentation for each service for more information about its log files.Compute service log - The log file for the compute service is located
/var/log/pf9/ostackhost.logon all hosts with the hypervisor role assigned. Useful for debugging issues with virtual machine creation or updates.Host agent log - The log file for the Platform9 host agent that is installed on each host is located at
/var/log/pf9/hostagent.log. This is helpful for debugging issues with host agent install or connectivity with the management plane.Communication agent log - Located at
/var/log/pf9/comms/comms.log. Log file for the Platform9 communications agent, which is responsible for ensuring the health and uptime of the host agent. Helpful for debugging issues regarding connectivity.Libvirt logs - Located at
/var/log/libvirt/qemu/<vm-id>where<vm-id>should be the UUID of the VM, and at/var/log/libvirt/libvirtd.log. Libvirt logs help with debugging any resource allocation or other issues with virtual machine instances.
Troubleshooting Steps
Verify that the host has outbound network connectivity to the internet and the management plane controller:
$ curl -s https://<FQDN>$ ping www.google.com$ telnetwww.google.com443
If your environment has a proxy server, make sure that you have configured the Platform9 host agent installed on the host to route traffic via the proxy:
$ sudo bash <path to host agent installer> --proxy=<proxy server>:<proxy port>
Verify that the hostname on the host matches the hostname shown in the GUI.
The UUID given in the
host_idfile should match the host UUID shown on the GUI.$ sudo cat /etc/pf9/host_id.confIf NTP is enabled, make sure the NTP servers are configured correctly across all the nodes within the configuration file
/etc/systemd/timesyncd.conf.d/conf.d. You can verify whether the host's time is in sync using$ sudo timedatectl status.Check that the Platform9 packages are installed on the host and the version which is shown on the GUI.
$ sudo apt list | grep -i pf9*Confirm that the host agent is running on your host:
$ sudo systemctl status pf9-hostagent$ sudo systemctl status pf9-ostackhost$ sudo systemctl status pf9-comms$ sudo systemctl status pf9-sidekick
Check the below service logs for any “
errors/timeouts/warnings”, during the time of issue. verify if any recent network changes might have impacted the connectivity.$ sudo cat /var/log/pf9/hostagent.log$ sudo cat /var/log/pf9/ostackhost.log$ sudo cat /var/log/pf9/comms/comms.log$ sudo cat /var/log/pf9/sidekick/sidekick.log
If these steps prove insufficient to resolve the issue, kindly reach out to the Platform9 Support team for additional assistance.
Most Common causes
Connectivity to the management plane controller is broken or unreachable.
Dependent services are down.
Packages are corrupted or not installed.
NTP is not synchronized.
Insufficient disk space available.
Ensure that the host that you are trying to add meets the pre-requisites.
Operational Recovery Runbooks
The following runbooks cover the most common operational recovery scenarios for the Compute Service:
Recover VMs in ERROR State After Host Reboot or Patching — diagnose and restore VMs that land in ERROR after a host reboot or maintenance event, including hard reboot, evacuation, and rebuild options.
Recover libvirt and Compute Service Failures — diagnose and restart
libvirtdand the Platform9 Compute Service when a hypervisor host becomes unresponsive.Diagnose VM Scheduling Failures ("No Valid Host Was Found") — step-by-step checklist for resolving VM placement failures caused by resource exhaustion, disabled hosts, host aggregate mismatches, or image property constraints.
Recover from Messaging Layer Failures Affecting VM Creation — identify and recover from messaging layer problems that cause VM creates to hang or fail silently at scale.
Diagnose a Host Agent Stuck in Converging State — locate the cause when a host remains in
convergingstatus after role assignment, including service start failures and log-rotation permission errors.Diagnose Role Assignment Failures — resolve HTTP 500 errors on role assignment and "Interface None is missing an IP" errors caused by incomplete host network configuration.
Troubleshoot Maintenance Mode Migration Failures — diagnose and recover when VMs fail to migrate or are left stranded or in an error state during maintenance mode.
Resolve CPU Baseline Mismatch After Host Upgrade — correct cluster CPU baseline issues after a host upgrade, including how to use
cpu_model_extra_flagsfor mixed-generation clusters.Troubleshoot VM HA — step-by-step diagnostics for VM HA not triggering, Consul health prerequisites, shared and FC storage validation, enable/disable failures, and post-upgrade re-validation.
Resolve Live Migration Failures for Legacy Boot-from-Volume Hotplug VMs — one-time corrective step for boot-from-volume, hot-add-capable VMs created before the 2025.10 release that still fail live migration.
Last updated
Was this helpful?
