Upgrade All: how one button left my VCF 9.1.1 lab deadlocked for two weeks

On September 3rd I upgraded my lab’s management components from VCF 9.1 to 9.1.1. In the Fleet Lifecycle UI there is a button labelled Upgrade All, and I used it, because I assumed it would run the components in the order the release notes prescribe.

It does not. And assuming it did was, in hindsight, fairly careless of me — the release notes spell out the dependencies in plain prose, I had them open at some point, and I still let a button decide the sequence for me. That assumption cost me two weeks and two separate deadlocks.

Two components failed that day: VCF Automation and the Migration Service Engine. Both ended up on “Ready for upgrade, prechecks failed, upgrade failed”, and there they sat.

TL;DR

PointWhat I learned
Upgrade All Button does not respect dependenciesIt starts components in parallel. Where one depends on another, that breaks things
The order that matters hereVCF Automation before Migration Service Engine — stated verbatim in the release notes
A half-finished upgrade is not “mostly working”I assumed I could keep using VCFA. I could not. It was blocked, just not visibly
Status fields lieServices reported Healthy while holding empty value sets and blocked deployments
The symptom sat four layers above the causeThree patches at three layers, all silently reverted
Don’t leave backups on the applianceThe upgrade replaces the VM. /root is empty afterwards

The causal chain

Sep 3    "Upgrade All" in Fleet Lifecycle
           |
           +-- Migration Service Engine starts BEFORE VCF Automation
           |     +-- Helm upgrade vcd-migrator 9.1.0.0200 -> 9.1.1
           |           +-- PostgresInstance to move from v14 to v17
           |                 +-- Webhook vpostgresinstance.kb.io denies it
           |                       (old platform only permits v13-v15)
           |                       +-- Flux: rollback -> retry -> ...
           |                             ENDLESS LOOP (13 days)
           |                             +-- PackageDeployment stuck "Progressing"
           |                                   +-- Backup precheck times out
           |                                         +-- VCFA upgrade impossible
           |                                               +-- Platform stays old
           |                                                     +-- Webhook stays old
           |
           +-- VCF Automation upgrade fails as well

Sep 4    DSM shows up as unhealthy -> I start an uninstall
           +-- 3 ConfigurationHandlers block one another
                 +-- Service Manager reconcile blocked

Sep 16   I try to install ArgoCD and nothing works
           +-- and only now do I start looking properly

Why I only noticed two weeks later

Here is the part I find most instructive in hindsight.

After the failed upgrade I carried on using the lab. VCF Automation was up, the UI responded, services were listed as Healthy. My working assumption was that an unfinished upgrade leaves you on the old version, which is inconvenient but harmless. You finish it later.

That assumption was wrong. The environment was not on the old version and fine. It was blocked — with a reconcile loop running in the background that made every subsequent upgrade attempt impossible. Nothing in the UI told me that.

What eventually made me look properly was ArgoCD. I wanted to install the ArgoCD Supervisor Service, and it threw a signature verification warning:

Connectivity to image depot.kube-system.svc/vcf/vcf-service-argocd/ga/1.2.0/
argocd-service:v1.2.0_vmware.1 cannot be verified: Get "https://depot.kube-system.svc/v2/":
dial tcp: lookup depot.kube-system.svc on 127.0.0.53:53: no such host

depot.kube-system.svc is the Supervisor’s internal software depot, and it only exists once a Regional Harbor registry has been deployed through VCF Automation. So: deploy Harbor. Except Harbor wouldn’t deploy either, and that is when I finally stopped assuming and started digging.

Harbor says Healthy, Harbor does nothing

Harbor was listed as installed and Healthy, but the Regions field was empty, the wizard showed no Service Properties, and Submit produced no task at all. I ran it several times convinced I was misclicking something.

The YAML view gave the whole story in one line:

regions: []

My region was selected. The checkbox was ticked. The resulting spec was empty anyway.

VCF Automation’s UI is an Angular SPA, and its API responses carry far more than the interface shows. This snippet in the browser console captures the relevant traffic:

window.__cap = [];
const OO = XMLHttpRequest.prototype.open, OS = XMLHttpRequest.prototype.send;
XMLHttpRequest.prototype.open = function(m,u){ this.__m=m; this.__u=u; return OO.apply(this,arguments); };
XMLHttpRequest.prototype.send = function(b){
  if (this.__u && /service-manager/.test(this.__u)) {
    this.__b = b;
    this.addEventListener('load', () => {
      window.__cap.push({ m:this.__m, u:this.__u, st:this.status,
                          req:String(this.__b||''), res:String(this.responseText||'') });
    });
  }
  return OS.apply(this, arguments);
};

Step through the wizard, then read window.__cap. I can recommend this for any VCFA problem where the UI stays quiet.

The interesting call is POST .../render-effective-values. The request carried my selection correctly:

#@ allowed_supervisors = { "k5re": "*" }

The response did not:

{"data":"regions: []\n"}

The request payload also contains the package’s ytt overlay, base64 encoded. Decoding it explained everything:

#@ def get_inventory_regions():
#@   filtered_supervisors = filter_supervisors_by_version(region["supervisors"], 9, 1)
#@   #@ Only include regions that have at least one valid supervisor
#@   if len(filtered_supervisors) > 0:

The Harbor package requires Supervisor 9.1 or higher. If no supervisor in a region qualifies, the region is dropped silently. No message, no warning, nothing in the logs I had checked.

Mine was on 9.0.2.0:

v1.32.9+vmware.2-fips-vsc9.0.2.0-25129014
                        ^^^^^^^

Takeaway: if a REGIONAL service shows Regions: –, the region failed a filter. It does not mean your selection wasn’t submitted.

Deadlock one: three handlers fighting over one slot

While looking at the service list I noticed VMware Data Services Manager still sitting in Uninstalling — since September 4th. I remembered that one. It had shown up as unhealthy, I had started an uninstall, it never finished, and I had filed it under “look at it later”.

Later had arrived.

data-services.broadcom.com | 9.1.1.0.25651575 | inactive | Busy | Uninstalling
"errorDetails": [{
  "type": "InstallTask",
  "severity": "ERROR",
  "message": "1 Installation error(s); Details: [3 CR reported error(s)]"
}]

Three CRs reporting errors, and no running task anywhere in VCF Operations. The uninstall task was long gone, only the service entry was still on Busy.

The three turned out to be ConfigurationHandlers in the prelude namespace:

CRHandler
21b92e70-4288-51b8-88cc-86f2588216a0self-managed-base-handler
65c5aa19-a3c7-5050-a08c-911a3326792fvcfa-info-handler
6e4f52c9-5de5-525e-9989-54217db25cdevcf-managed-base-handler

All Unhealthy, each complaining about one of the others:

ConfigurationHandler '65c5aa19-...' already exists for service ID '73e5e9e3-...'.
Only one ConfigurationHandler per service is allowed

All three shared an identical creationTimestamp, so the package created three handlers at once into a platform that permits exactly one. That is a packaging bug, not an operator error.

Back up first:

mkdir -p /root/ch-backup && cd /root/ch-backup
kubectl -n prelude get configurationhandlers.services.vcfa.broadcom.com -o yaml \
  > all-configurationhandlers-$(date +%F).yaml

Then delete two of the three:

kubectl -n prelude delete configurationhandlers.services.vcfa.broadcom.com \
  21b92e70-4288-51b8-88cc-86f2588216a0 \
  65c5aa19-a3c7-5050-a08c-911a3326792f \
  --timeout=90s

All three carry a finalizer, so I expected this to hang. It didn’t, which told me the controller was alive. If yours does hang:

kubectl -n prelude patch configurationhandlers.services.vcfa.broadcom.com <name> \
  --type=merge -p '{"metadata":{"finalizers":null}}'

The remaining handler stayed Unhealthy afterwards, still referencing an object I had just deleted. Stale status, fixed with a reconcile trigger:

kubectl -n prelude annotate configurationhandlers.services.vcfa.broadcom.com \
  6e4f52c9-5de5-525e-9989-54217db25cde reconcile.trigger="$(date +%s)" --overwrite

All three went Healthy, the DSM uninstall completed within minutes, and Harbor promptly dropped from Healthy to Not Installed — considerably more honest.

One deadlock down. I thought I was nearly there.

Deadlock two: the carousel

To get Harbor I needed Supervisor 9.1. For that I needed VCF Automation on 9.1.1. And VCF Automation refused:

Backup precheck failed. Timed out waiting for in-progress package deployments
to complete: vcd-migrator.

[com.vmware.vcfms.system.precheck.backup.PackageDeploymentsCompleteWaitTimeout]

That precheck is not a preliminary step you can skip. It is a subtask inside the upgrade workflow, and the component’s “…” menu offers nothing but Run Prechecks.

kubectl get packagedeployments.releases.vmsp.vmware.com -A
prelude         vcfa-bundle     Successful
vcd-migrator    vcd-migrator    Progressing   package deployment is in progress
vmsp-platform   vmsp-platform   Successful

My first guess was a stale status, same as before. Turns out to be wrong:

kubectl get packagedeployments.releases.vmsp.vmware.com vcd-migrator -n vcd-migrator \
  -o jsonpath='gen={.metadata.generation} observed={.status.observedGeneration}{"\n"}'

gen=4 observed=4. The controller was current and working. The pods were healthy and had been running for days. Something was genuinely in progress and never finishing.

Then the Helm release made me look twice:

vcd-migrator  REVISION 19905  deployed  vcd-migrator-9.1.0-0200-25556825

Nineteen thousand nine hundred and five, on a namespace 22 days old. Roughly 1,200 deployments a day, one every seventy seconds, for two weeks. Not a stuck deployment — a carousel.

The helm-controller log shows one full lap:

HelmChart ... is in-sync
release out-of-sync with desired state: release chart changed
running 'upgrade' action with timeout of 30m0s
release is in a failed state
running 'rollback' action with timeout of 30m0s
error: failed to wait for object to sync in-cache after patching: context deadline exceeded
-> repeat

And the reason it fails:

Helm upgrade failed for release vcd-migrator/vcd-migrator with chart
vcd-migrator@9.1.1-0-25714559: cannot patch "vcd-migrator-postgres" with kind
PostgresInstance: admission webhook "vpostgresinstance.kb.io" denied the request:
postgres versions below 13 and above 15 are not supported

The 9.1.1 chart wants PostgreSQL 17. I was on 14. The admission webhook from the old platform allows 13 through 15, because that is what this platform generation supports.

So the migrator needs the new platform. The platform arrives with the VCFA upgrade. The VCFA upgrade is blocked by the migrator. I stared at that for a while.

What did not work

Three attempts to stop the loop, all reverted within seconds by the vmsp-operator. I’m listing them because they are exactly what anyone would try first:

# Suspend the HelmRelease. Standard Flux move.
kubectl patch helmrelease vcd-migrator -n vcd-migrator \
  --type=merge -p '{"spec":{"suspend":true}}'

# Suspend the PackageDeployment. There is a purpose-built field for it.
kubectl patch packagedeployments.releases.vmsp.vmware.com vcd-migrator -n vcd-migrator \
  --type=merge -p '{"spec":{"suspend":{"enable":true,"cascade":true}}}'

# Reset the target version so there is nothing to upgrade to.
kubectl patch packagedeployments.releases.vmsp.vmware.com vcd-migrator -n vcd-migrator \
  --type=merge -p '{"spec":{"packages":[{"name":"vcd-migrator","version":"9.1.0.0200.25556825"}]}}'

The PackageDeployment controller also kept resetting the Helm retry counter:

"message":"overriding HelmRelease upgrade retries for all namespaces","retries":10

Flux was never going to give up, because something kept telling it not to. I had spent an evening treating the symptom.

Finding the source

The next day I went looking for whoever was doing the overwriting:

kubectl logs -n vmsp-platform deploy/vmsp-operator --tail=300 | grep -i migrator

Every line carried the same object:

"bundle": {"name":"vcd-migrator-9.1.1.0.25714559","namespace":"vcd-migrator"}

A Bundle. That was the layer I hadn’t looked at. The control hierarchy in VCFA runs like this, and everything I had patched sat in the bottom half:

Bundle  ->  PackageDeployment  ->  HelmRelease (Flux)  ->  Helm release
kubectl get bundles.bundle.vmsp.vmware.com -A
prelude         vcfa-bundle-9.1.0.0200.25556825    Successful
prelude         vcfa-bundle-9.1.1.0.25714559       Pushed
vcd-migrator    vcd-migrator-9.1.0.0200.25556825   Pushed
vcd-migrator    vcd-migrator-9.1.1.0.25714559      Pushed     <- this one
vmsp-platform   vmsp-platform-9.1.1.0.25714471     Pushed

Before deleting anything, check whether something owns it and would recreate it:

kubectl get bundles.bundle.vmsp.vmware.com vcd-migrator-9.1.1.0.25714559 \
  -n vcd-migrator -o yaml | sed -n '/^metadata:/,/^spec:/p'
metadata:
  creationTimestamp: "2026-09-04T15:20:45Z"
  finalizers:
  - bundle.vmsp.vmware.com/finalizer
  # no ownerReferences

No ownerReferences. Nothing brings it back on its own.

The fix

# 1. back up, and scp it off the appliance (see Learnings)
mkdir -p /root/bundle-backup && cd /root/bundle-backup
kubectl get bundles.bundle.vmsp.vmware.com vcd-migrator-9.1.1.0.25714559 \
  -n vcd-migrator -o yaml > bundle-vcd-migrator-9.1.1.yaml

# 2. delete the 9.1.1 bundle only. The 9.1.0.0200 one has to stay.
kubectl delete bundles.bundle.vmsp.vmware.com vcd-migrator-9.1.1.0.25714559 \
  -n vcd-migrator --timeout=120s

# 3. now reset the target version. Same patch as the night before.
kubectl patch packagedeployments.releases.vmsp.vmware.com vcd-migrator -n vcd-migrator \
  --type=merge -p '{"spec":{"packages":[{"name":"vcd-migrator","version":"9.1.0.0200.25556825"}]}}'

Two minutes later:

9.1.0.0200.25556825
vcd-migrator   Successful   successful package deployment

The version stuck. The revision counter stopped. After thirteen days and something north of 19,000 revisions, the carousel stood still.

Order matters here. Deleting the bundle alone is not enough, the PackageDeployment keeps its 9.1.1 target. Patching alone is not enough either, the operator rewrites it from the bundle. You need both, bundle first.

The upgrade itself

The precheck passed on the first attempt. Then I clicked Upgrade on ‘one’ component – VCF Automation.

It ran for about six and a half hours across thirteen subtasks. A few things worth knowing:

The appliance VM gets replaced. My hostname went from ...-r8hpx to ...-l5z2s, and SSH greeted me with a host key warning:

ssh-keygen -R vcf-automation-vip.mb.lab

One subtask logged Deploying VCF Component. Status: Failed mid-run — as Info, not Error, and the workflow carried on. Helm retries are set to 10 here. Do not intervene, however tempting.

For watching progress from the appliance:

kubectl get machines -A
kubectl get nodes
kubectl get packagedeployments.releases.vmsp.vmware.com -A

helm list -n prelude -a -o json | python3 -c "
import json,sys
from collections import Counter
d=json.load(sys.stdin)
print(Counter(r['status'] for r in d))
print('updated today:', sum(1 for r in d if r['updated'].startswith('$(date +%F)')), 'of', len(d))
"

Use -a. Without it, helm list only shows deployed releases, and during the upgrade they pass through pending-upgrade and vanish. My counter climbed to 47, dropped to 28, and I briefly assumed something was rolling back. Nothing was.

End result:

vcd-migrator   9.1.0.0200.25556825   Running     <- deliberately held back
vcfa           9.1.1.0.25714559      Running
vsp            9.1.1.0.25714471      Running

Learnings

Don’t use Upgrade All. The whole article in one line. The button starts components in parallel, and where one depends on another, that breaks. There is also a known issue where concurrent starts cause upgrade plans to interfere, leaving a component on “Ready for upgrade” with matching builds on both sides and no customer-side way to correct it.

Read the dependencies in the release notes, not in blog posts. Mine included. The 9.1.1 release notes don’t give a numbered sequence, they give rules:

#Rule
1Fleet Lifecycle to 9.1.1 before any other VCF component
2VCF Operations must not be patched in parallel with others
3VCF Operations and license servers before ESX patches
4Management Services Runtime before Identity Broker
5Management Services Runtime before Salt RaaS
6VCF Automation before Migration Service Engine
7Software Depot blocks other patches while it runs

Rule 6 is the one I broke. And you’ll find a simplified chain circulating on various blogs that includes “SDDC Lifecycle” as a step — it isn’t in the release notes. I had it in my own notes until I went back and checked.

A failed upgrade is not a harmless pause. This was my most expensive assumption. I kept using the environment for two weeks, and it looked fine from the outside while a reconcile loop made every further upgrade impossible. Clean up failed upgrades promptly, even if nothing appears broken.

Status fields lie. Harbor said Healthy with an empty value set. DSM said Healthy at service level with three deadlocked CRs. The migrator’s Helm release said deployed while being redeployed every seventy seconds. Look at status.conditions, observedGeneration and actual timestamps instead.

Watch revision numbers. A four-digit Helm revision after three weeks is a red flag you can spot at a glance.

Find the source, don’t sedate the symptom. Three patches at three layers, all reverted, because the controlling layer was above all of them. I should have gone to the operator logs an hour earlier.

Webhooks exist for a reason. I was briefly tempted to exclude the namespace from the Postgres validating webhook to force the upgrade through. Good thing I checked what the chart actually wanted first: Postgres 17, on a platform whose operator tops out at 15. That would have moved the failure one layer down and possibly taken the database with it.

Don’t leave backups on the appliance. I dutifully backed up the bundle YAML to /root/bundle-backup before deleting it. The upgrade then replaced the entire VM. You can guess the rest.


But finally


Next steps

Not done yet. The Supervisor is still on 9.0.2.0 and won’t upgrade, because VKS 3.4.1-embedded blocks the precheck and the newer VKS is embedded in the Supervisor OVA I don’t have yet. A proper chicken and egg. The way out appears to be registering a VKS version manually from the support portal rather than using the embedded one, and that is next on my list.

After that, Harbor, and then finally ArgoCD — which is where this whole thing started.

If you’ve hit the same problem, or found a better way around the VKS situation, I’d be glad to hear about it in the comments. And if you’re about to patch your own management components: one at a time, wait for Completed, verify the version. It costs you an afternoon and saves you a fortnight.

No responses yet

Leave a Reply

Your email address will not be published. Required fields are marked *