Skip to content

Troubleshooting

A short FAQ. For the full VM-over-SSH host-side troubleshooting table (permission-denied, host-key verification, missing template variables and more), see the VM backup over SSH guide on GitHub.

Something is not wired up correctly

Open /spike in the web UI. The host integration check probes every mount and CLI (Docker socket, libvirt, restic, qemu-img, rclone) and reports any missing pieces. Start here before assuming a bug: a missing mount or an unreachable host shows up immediately.

I cannot reach the web UI

BombVault serves HTTPS out of the box on port 3443 (self-signed certificate), so open https://<your-unraid-ip>:3443. Accept the self-signed certificate warning, or put BombVault behind a reverse proxy with your own certificate. If you run with HTTP_ONLY=true, it serves plain HTTP on port 3000 instead (intended for use behind a TLS-terminating proxy).

I lost my APP_KEY

APP_KEY derives the restic repository password. Without it (and without the encryption-key recovery kit), encrypted backups cannot be recovered. This is why the Dashboard nags you to download the recovery kit. See Off-site & recovery. Generate a key with openssl rand -hex 32 and store it off the server before you rely on any backup.

VM backup will not connect

VM backup talks to libvirt over SSH, never a mount.

  • Confirm SSH is enabled on the host and BombVault's public key is authorized in /root/.ssh/authorized_keys (Settings, Integrations, Host SSH shows the key and a Test connection button).
  • On a custom br0.x network, set LIBVIRT_HOST to your Unraid LAN IP (the container cannot reach the host via host.docker.internal there). Enable Settings, Docker, Host access to custom networks.
  • If you changed Unraid's SSH port, set LIBVIRT_SSH_PORT to match.
  • Full step-by-step diagnosis (reachability test, VLAN routing, Permission denied (publickey), Host key verification failed) is in the VM backup over SSH guide.

A live VM snapshot did not run

Live snapshots need the qemu guest agent installed in the VM and the disk on /mnt/cache (or /mnt/diskX), not /mnt/user. On a shut-off VM, live automatically falls back to graceful. A graceful backup shuts the VM down, backs up the disks, then restarts it, so it is always consistent.

A backup failed with "repository is already locked"

This is usually an orphaned restic lock left behind when the container was updated or restarted mid-operation. BombVault detects a provably orphaned lock, force-clears it and retries once, automatically. If it persists, use Settings, Integrity, Unlock for the affected domain to clear a stale lock by hand. A genuine problem still surfaces rather than being hidden. After a restart BombVault waits until such a lock has gone ten minutes without a refresh. A restic that is still running, for example in a second BombVault on the same repository, refreshes its lock every five minutes.

My off-site copy did not happen after a backup

Off-site replication is best-effort by design, so an off-site hiccup never fails the local backup. Check the off-site schedule for that domain (Settings, Schedules): a blank schedule replicates after every local backup, while a cadence ships less often. Use Replicate now on the Off-site page for an on-demand run, and watch the replication indicator on the Dashboard.

A restore aborted before it started

Before anything is stopped or removed, restore runs a pre-flight conflict check: it verifies the container's static IP and published host ports are free. If another container already holds one, it aborts with a clear, actionable message instead of leaving a half-finished restore. Free the conflicting port or IP, then retry.

A plain export failed instead of writing a file

If age encryption is on (Settings) but no valid recipient is set, an export fails with a clear error instead of writing plaintext. Add a valid recipient (an age public key or an SSH public key), or turn encryption off if you intend the export to be plaintext. See Features.

A database dump failed

A failed dump never fails the backup around it; it is recorded as its own failed run, and the reason names what to fix.

  • Login refused. The dump signs in with the container's own password variables (POSTGRES_PASSWORD, MARIADB_ROOT_PASSWORD, MYSQL_ROOT_PASSWORD or their _FILE versions). Check them on the database container. A _FILE variable pointing at a secret the container's own user cannot read fails the same way.
  • Missing privileges. With a random root password the dump can only sign in as the app user, so it holds that one database, and MySQL 8.4 and newer may refuse it outright. Give the container a real root password, or switch the dump off for it.
  • The system tables need an upgrade. MariaDB refuses to dump when its system tables come from an older version (error 1558). Add the variable MARIADB_AUTO_UPGRADE=1 and restart the container, or run mariadb-upgrade inside it once.
  • No dump tool. A slim or self-built image without pg_dump, mysqldump or mariadb-dump cannot be dumped. Use the official image, or switch the dump off.
  • A time limit. A dump gets DB_DUMP_MAX_HOURS (6 by default), the backup around it gets BACKUP_MAX_HOURS, and a dump that stops making progress is cut after BACKUP_STALL_HOURS. The usual cause of the last one is a lock the application holds. Raise the limit that fired, or dump while the application is quiet.
  • The container is paused or restarting. The dump talks to the running server. If the container keeps restarting, its own log says why.
  • A damaged dump could not be removed. A dump BombVault could not finish is deleted again. When that delete fails, the dump stays in the list marked as damaged, and you can delete it there.

An import failed

An import stops the container, moves its data folder aside and lets the image create an empty one in its place. If a step before the import itself fails, the old folder is put back automatically. If the import fails, the container keeps the fresh folder and the old one stays beside it as <data folder>.bombvault-before-import-<timestamp>; the run's error message names the exact path.

To put it back by hand: stop the container, rename the current data folder out of the way, rename the kept folder back to the original name, and start the container. On Unraid, the file manager does this from the Shares tab.

A ZFS dataset backup failed or skipped a dataset

Every problem carries a reason code in brackets, and the ZFS datasets page lists all of them with the fix. The three most common:

  • snapshot-loop: the snapshot did not reach BombVault because Host Data does not pass new mounts through. Edit the container, set the Access Mode of Host Data to Read/Write - Slave and restart BombVault.
  • key-not-loaded: an encrypted dataset whose key is not loaded is skipped. Load the key with zfs load-key and mount the dataset; the next backup includes it.
  • ssh-auth: the server refused BombVault's key. The connection card on the ZFS page shows the command that authorizes it; run it once on the server.

An item stays at "Learning N/10"

Most anomaly checks start after 10 successful backups of an item, and the count starts again after Mark as expected and after the item's selection changed. An item that is not scheduled does not learn, and a container without appdata has nothing to learn from, which its badge says.

Retention stopped deleting old backups of one item

An open critical anomaly is holding them: the item's source is almost empty, shrank sharply, or one backup stored most of its data again. Open the anomaly from the badge on the item. If data is missing or was encrypted, restore from the linked last good backup first. Then acknowledge the anomaly, or mark it as expected if the change was yours, and the next run prunes as usual. The retention preview marks such an item as kept. For a ZFS item only the dataset named in the anomaly keeps its old backups, and the other datasets of the tree are pruned as usual.

Manual prune says some items were kept

The same cause: prune leaves the old backups of an item with such an anomaly alone and names the item in its message. Everything else is pruned as usual.

History import says a repository could not be read

After the upgrade BombVault reads the sizes of earlier backups from each repository once. A repository it could not reach at that time, such as an off-site target that was down or a share that was not mounted, is counted on the Anomalies card under Settings, Integrity and tried again once a day. Its items learn from new backups in the meantime.

The disk-space warning does not match the Unraid dashboard

On the Unraid user share (/mnt/user) the free space is that of the whole array, not of one disk. Remote repositories are measured only through rclone remotes that report their free space; S3, B2, REST and SFTP repositories have no figure and are listed as not measured on the Anomalies card.

An AI assistant cannot connect

The MCP server page lists what each status code and each refusal of the MCP endpoint means and what to do about it.

The container keeps restarting or looks unhealthy

BombVault reports healthy/unhealthy from its own /api/health. An auto-heal tool (such as Autoheal) can restart it automatically if the engine ever wedges. Check the container log and the /spike report for the underlying cause.

Still stuck?