Connecting¶
Four things can talk to a KnightLoader: the web UI it serves itself, a desktop build, the Android app, and the browser extension. Any of them can also talk to a second KnightLoader. This is how each of those connections is made, and what happens when one of them cannot be.
The short version: on one network, nothing needs configuring. Across networks, twelve words connect every instance you run. The relay section after that is the same machinery reached the manual way, for anyone who wants to name the relay and the key themselves.
On one network¶
An instance announces itself every five seconds over UDP multicast, to an administratively-scoped group that routers do not forward off the local network. Every other instance listening picks it up.
The Instances page shows what it heard under Found on your network, with an Add button on each row.
Being on the same network is not consent, so nothing is connected until that button is pressed. A guest laptop and an IoT device sit on that network too.
Adding stores an address. It exchanges nothing: no credential travels in either direction. A credential exchange triggered by an announce that anything on the LAN can send would let any device there help itself to a token.
So if the instance you add has a password set, it will refuse the calls that follow, and the page says so. An address is not a credential; the connection phrase is, and it covers every instance in the group at once instead of one peer at a time. An instance with no password works immediately.
Nothing in an announce is trusted. Anything on the network can send one, so the fields are length-capped on arrival and the list is bounded, so a device announcing endless invented instances cannot grow it without limit. What a row shows is a claim. That is why the address is displayed next to the name, and why nothing is added until you decide to add it.
Multicast blocked, or no network at all, means an empty list. Everything below still works.
Across networks: the connection phrase¶
Settings → Access → Connect instances and remote access → Create a phrase. Twelve words come back. Type them into every other KnightLoader you run, and they find each other across networks and behind NAT, with no port forward, no domain, no account and nothing to log into.
The words are the credential. Whoever has them reaches every instance in the group, so keep that in mind before you read them out over the phone.
The words carry a 128-bit secret and nothing else: no address, no name. The relay's address is compiled into the binary, which is what keeps this to twelve words instead of a URL plus a key.
The list they are drawn from is BIP39's, the one hardware wallets use. That has nothing to do with cryptocurrency, and the UI never calls this a wallet seed: those 2048 words were chosen so no two share their first four letters, none are near-homophones and none carry accents, which is what makes a phrase safe to read down a phone line and to type on a mobile keyboard. Four checksum bits ride along, so a mistyped or swapped word is refused on the spot, naming the word and its position, rather than becoming a connection that silently never finds its sibling.
The relay never learns the phrase. What an instance sends is
SHA-256("knightloader/relay/group-key/v1" || secret). The relay matches
connections that present the same derived key and forwards frames between
them; it has no account list, no registration step and no database. So whoever
runs it (us at relay.halleluja.design, or you) cannot reconstruct anybody's
words.
And it cannot read what it forwards. A second key comes out of the same
secret under its own domain, SHA-256("knightloader/relay/frame-key/v1" ||
secret), and every proxy frame is sealed with AES-256-GCM under it. The relay
sees which instance a frame is addressed to and which request it answers,
because it routes on those, and nothing else: not the path, not the body, not
the API token a phone attaches. The two domains are what makes this work: the
relay is handed the group key in every hello frame, so a frame key derived
from that would be one it already holds.
The routing fields are bound into the seal as additional data, so a relay cannot take a frame addressed to one instance and deliver it as another's. The one field it does author is the error it returns when nobody is connected under a target id, which it has to, holding no key. That lets a hostile relay claim an instance is absent (a denial of service it could equally perform by dropping the frame), but not fabricate an answer: a response with no sealed payload is never mistaken for one.
If you configure a relay by hand-entered key instead of by phrase, there is no secret to derive from and the frame key comes from the relay key itself. That still seals the traffic against anything sitting between you and the relay (a reverse proxy, a TLS terminator, a captured log), but not against the relay operator, who is handed that key. That trade fits the case it exists for: somebody hosting the relay themselves.
To run your own, put its address in relayUrl under Settings → Advanced
on every instance in the group; the same phrase then works against it, because
the phrase carries the secret and not the address. That page lists every
setting this instance has, so the self-hosting knobs live there and not on the
card. The card is for the twelve words, which is what almost everybody needs.
relayServe is on the same page, for the case where one instance is the
relay.
Showing the phrase again needs the instance password re-entered, when one is set. A live session is not enough: it may have been opened hours ago on a screen nobody is sitting at, and what is behind that button is not this instance's password but the key to every instance in the group.
An instance with no password says so, loudly, before it mints anything, and then mints it anyway if you tell it to. The phrase reaches every instance you connect with it, so an unprotected one is a door into all of them. That is your call to make about your own network, so it is not refused on your behalf.
Leaving forgets the secret and stops dialling. The other instances keep going without it; a phrase is a group, not a pairing.
When neither can reach the other: the relay¶
Both ends dial out to a relay and meet there, so neither needs a public address, a port forward or a domain. Traffic is JSON frames over one WebSocket each.
This is the same machinery the connection phrase uses, reached the manual way: a URL and a key you choose and type into both ends, rather than twelve words that carry the key and already know the address. The phrase is the shorter road to the same place; this one exists for anyone who wants to name the relay and the key themselves.
There is an official relay, wss://relay.halleluja.design/relay/connect,
which is what a phrase points at unless you override it, and running your own
is a first-class option, not a fallback. docker compose it anywhere both ends
can reach, put the same key in both, done. Set KL_RELAY_DOMAIN and it
terminates TLS itself, getting and renewing its own certificate over
TLS-ALPN-01: no reverse proxy, no certbot, no renewal cron, and no port 80. The
challenge completes inside a handshake on 443, so the firewall in front of it
opens one port.
KL_RELAY_DOMAIN takes a comma-separated list, and the certificate covers
every name in it. That is for one situation and it is worth knowing before you
need it: moving a relay to a new address. Old clients keep dialling the old
name, and a whitelist of one would stop issuing a certificate for it the moment
you switch, so those clients fail in the handshake instead of moving across.
Run both names for as long as anything still dials the old one, then drop it.
Or let one instance be the relay¶
Settings → Advanced → relayServe. The relay then answers under
/relay/connect on the address that instance already uses, behind the same
reverse proxy and the same certificate, and the other instances put that
address in their own relayUrl. It needs no second container, no second port
and no second certificate.
For an instance a proxy serves under a path, enter the address with the path,
such as https://example.com/kl, and the others dial /relay/connect below
it. A relayUrl whose path already ends in /connect is dialled as it stands,
for a relay a proxy has mounted somewhere else.
It admits only the relay key that instance stores, so switching it on does not
turn a published address into a meeting place for whoever finds it. With the
switch off, /relay/connect answers 404, the same as any build that never had
the feature.
What this does not change is the one requirement a relay has: it is the third point both sides dial out to, so it has to be reachable by both. Turning it on inside a desktop install that nothing can reach from outside gives the other instances nothing to dial. The instance that hosts it is the one with the address (a server, a NAS, anything already behind a domain), and the ones behind NAT are what it exists to connect.
The relay operator carries your frames, so they see who is talking and when, and a relay you do not run is a relay you are trusting with that. Run your own if it matters. They cannot read your phrase: they only ever receive a hash of it.
Being on the relay is what authenticates a sibling. A request arriving this way came off a socket the relay only joins to other connections presenting the same group key, so the sender has already proved it holds the phrase before any handler runs. There is no second credential to exchange, which is what retired the pairing code: whoever can present the group key could have joined the group and been handed one anyway.
That makes the reachable surface the thing to bound, and it is an allowlist rather than a property each route happens to have. A sibling may read and drive tasks, links and the queue, answer or skip the captchas holding them up, load a widget captcha's page for the phone and say when it will not load, and may read whether a password is set, who else is in the group, and the instance's own accent and corner shape. It cannot read the settings or the accounts, change the password, mint an API token, or ask for the phrase back. A route added later is outside the list until somebody puts it in.
A relay peer is never written to disk. It exists for as long as the relay sees it, and one remembered across a restart would be a peer that cannot be reached and cannot be explained.
The Android app¶
- The phrase is the one way in, as it is in the browser extension: twelve words, typed or scanned from the QR the web UI shows beside them, and every instance in the group appears at once, with no address, no token and nothing to look up. The phone derives the same key its siblings do and dials the same relay, which is what authenticates it, so a password on an instance costs nothing extra here.
- A connection saved by address in an earlier build of the app keeps working, but the app no longer makes one. Those builds took the address typed in, scanned from the Access tab's QR or found with Find on this network, plus a token where the instance had a password.
Find on this network, in those builds, asks every address on the phone's own
/24 for /api/health over HTTP and fills in the address of anything that
answers as a KnightLoader. React Native has no UDP socket, so the app cannot
join the multicast group the servers use.
That route is frozen because of this. It answers {"status":"ok","version":…}
on a 200 for as long as the process is up, whatever is actually wrong with the
instance, and those builds compare that string literally, so an instance that
answered degraded would stop being findable by every phone still running
one, and phones update on their own schedule rather than with the
container. The container's own HEALTHCHECK and the Click'n'Load bridge read
it the same way. New fields may be added to it; the two that are there may not
move, and the status may not stop being ok.
The real state is a second, session-guarded readout: GET /api/health/detail
answers every part of the instance with a state of its own, the queue by why
it is waiting and why it failed, room on the target folders and how long the
process has been up. GET /api/metrics is the same reading as Prometheus
exposition text and answers 404 until the switch on the Health settings page
is turned on. Neither is open, and neither is forwarded to a peer over the
federation or the relay: both describe this machine's disks and sidecar, and
a row of them drawn under a peer's name would name the wrong box.
- Captchas are answered on the phone. A card on the instance's downloads,
a count on its overview card and a banner over the open screen lead to a list
of what is waiting, and picture and click captchas are answered right there.
reCAPTCHA, hCaptcha and Cloudflare Turnstile open in a window of their own,
on either kind of connection. The app fetches the instance's widget page,
over the relay as well, and shows it under the address of the hoster's page
the captcha came from. The captcha service then sees the hoster's website,
as it would in a browser on that page. Many hosters tie their captcha to
their own domains, and every Turnstile key is tied that way, so anywhere
else the service refuses to run.
The banner also says when a captcha timed out or was answered somewhere else,
as the web UI's messages do.
The app watches the instance it has open, and only while it is in front; the overview counts what waits on the others. A captcha that arrives while the app is in the background is announced when you come back to it, as long as Android has kept the app in memory. After Android has closed it, the card on the downloads still shows what is waiting, but no banner comes up. The app sends no notification while it is closed. While it watches, the instance counts you as watching for the captchas the app can answer, which is every kind except a captcha service KnightLoader does not know. With Only when nobody is watching switched on on the Captcha settings page, the paid solvers wait for your answer on those first, and a captcha the app cannot answer does not wait for it. Nor does a widget captcha that will not load in the app, until Refresh loads it after all. The card says what the solvers are doing, as the web UI's captcha window does.
The browser extension¶
The extension takes the same twelve words and nothing else. There is no address
field, no name field, no token field and no sync button on the options page any
more. The extension carries its own relay client (extension/src/relay.js) and derives
the group key from the phrase with a WebCrypto port of internal/seedphrase
(extension/src/phrase.js), so it is a group member in its own right rather
than a guest of one configured instance.
That replaces the whole previous shape and everything that hung off it:
- The roster is read live, when a window opens, instead of being stored and synced. An instance that is switched off is not offered; one that came online a minute ago is. Nothing tells this browser anything. It asks.
- A peer with no address of its own (a desktop build, or one reachable only through a relay) is now reached directly through the relay, not forwarded on its behalf by a sibling that happens to have an address.
- Sends are
POST /api/linksover the relay, admitted because membership is the credential. No window opens, no session cookie is involved, and thesameOriginguard is not worked around, because it is not on that path.
The site access the extension asks for at install time is for Click'n'Load and
for nothing else; see docs/browser-tools.md.
API tokens and their rights¶
A script, Sonarr, a dashboard, or the phone app connected by address rather than by phrase signs in with an API token: Settings → Remote access → API tokens → New token. Every token has a name, can be revoked on its own, and carries some of four rights:
| Right | What it covers |
|---|---|
| Read | the download list, the queue, the history, the statistics, the health readout and the live stream |
| Add | links, torrents, NZB files handed over by Sonarr or Radarr, and containers |
| Control | pausing, resuming, removing and reordering downloads, the queue's own switches, captchas, unpacking, and the accent, corners and rainbow the apps take over |
| Admin | settings, accounts, tokens, instances, scripts, the logs, backups and restarts, and picking the folder a download goes to |
The window that creates a token offers four presets: Full access (all four), Add and read, Read only, and Custom, which shows one switch per right.
Add and read is enough for Sonarr and Radarr, through either of their
doors. Through SABnzbd's they send their key in the address (?apikey=), where
it ends up in the log of any proxy in between, and a key that can only add and
read is a smaller loss than one that can change the password. When Sonarr clears a finished download after importing it, the
bridge stops reporting it. The download itself stays in your list unless the
token also has Control; then it is removed, with its files if Sonarr asks for
that. Removing a download that is still running needs Control: without it
Sonarr shows the refusal, and the download carries on and stays in Sonarr's
queue.
Through qBittorrent's door Sonarr also resumes every torrent right after adding it, which Add covers as long as nothing in the torrent was paused. Resuming a paused torrent needs Control, and so do the qBittorrent settings in Sonarr that act on a torrent after the add: Priority "First", Initial State "Force Started" and a Post-Import Category. With Add and read the torrent stays as it was.
Such a key cannot pick where files land either. A dir in POST /api/links
or POST /api/tasks/options, or a save path in qBittorrent's torrents/add
or torrents/createCategory, needs Admin, whatever the route needs otherwise,
because a folder of the caller's choosing could be any folder the instance can
write to. Without one, links go where the download folder, a category or a
Packagizer rule puts them.
Read only suits a dashboard, or a monitoring system reading /api/metrics.
The phone app uses Read, Add and Control and never needs Admin. The Modules
page warns when no token has the rights the Sonarr bridge or /api/metrics
needs.
A call the token has no right to is answered with a 403 that names the missing right:
{"error": "this API token does not have the \"control\" right", "code": "tokenScope", "params": {"scope": "control"}}
The SABnzbd door puts the same sentence in SABnzbd's own error document,
because that is what Sonarr shows, and the qBittorrent door sends it as the
text of its 403. GET /api/help lists the right each route
needs as its scope. The table behind it is internal/api/scopes.go, and a
test fails for any route missing from it, so a new route cannot be reached with
a narrowed token until somebody has decided which right it needs.
Only tokens are narrowed. A browser signed in with the password, a sibling that came in over the relay with the phrase, and the tokens instances hand each other when they pair all have every right. A token made before tokens had rights keeps all four, so scripts and apps set up earlier keep working. A token that may manage tokens can only issue ones with rights it holds itself. A call a token forwards to another instance is checked here, against the right it would need on this instance, before it leaves; forwarding needs no right of its own, so a token that may only add can add links on a peer too.
The rights are kept in token-scopes.json beside tokens.json in the data
directory. A token stays as narrow as it was even after going back to a
release without rights for a while and upgrading again.
The secret is shown once, when the token is made. The instance keeps a hash of it and never sends it back.