Distills patterns from the Teleport access topology work: verified-data diagram building, as-displayed HTML-to-PDF export via headless Chromium, and container-scoped/host/database onboarding into an existing Teleport cluster. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R3ZTfgQrkR3q8DSEoZAmvs
8.9 KiB
name, description
| name | description |
|---|---|
| teleport-onboard | Onboard a new resource into a Teleport cluster — a VPS/droplet, an EC2 (or any fresh VM) host, a container-scoped least-privilege app login on a multi-app host, or a database (db_service) — following an existing team's established patterns rather than inventing a new access model. Use when asked to register, join, or onboard something into Teleport, add scoped/restricted access to a specific app or container, or when Teleport resources (nodes/databases/roles) don't match what's actually running. |
Teleport resource onboarding
First: discover the existing pattern, don't assume one
Before creating anything, SSH into a host that's already onboarded the way you intend to onboard the new one, and read what's actually there:
cat /etc/passwd # look for non-standard login shells — a sign of scoped access
sudo cat /etc/sudoers.d/<candidate-user> # the exact commands that login is allowed to run
sudo cat /etc/teleport.yaml # or find it: /opt/teleport/config/, /etc/teleport/ — it varies
ps aux | grep teleport # more than one teleport process on a host is a real bug (see gotcha below)
Do not invent a new access model if one already exists — mirror it exactly, including its exact sudoers syntax, its wrapper-script conventions, and its Teleport role shape.
Where tctl actually runs: the cluster's auth server is often just one specific host's Docker container (look for an image like teleport-distroless in docker ps), not a separate admin machine. sudo docker exec teleport tctl ... gives full cluster admin without a separate tctl binary or identity file — check for this before assuming you need new credentials.
Pattern A — container-scoped forced-login-shell (multi-app host, least privilege per app)
Use when one host runs several unrelated apps as containers and different people need access to only their own app's container, never the host shell or a sibling's container.
-
Wrapper script
/usr/local/bin/enter-<app>.sh(root:root,0755):#!/bin/bash set -euo pipefail # Forced login shell — must never fall back to a real host shell. if [ ! -t 0 ]; then echo "Interactive TTY required." >&2; exit 1; fi echo "=== <app> - container access ===" echo "1) <container-a>" echo "2) <container-b>" read -rp "Pilih [1-2]: " choice case "$choice" in 1) exec sudo /usr/bin/docker exec -it <container-a> sh ;; 2) exec sudo /usr/bin/docker exec -it <container-b> sh ;; *) echo "Pilihan tidak valid."; exit 1 ;; esacOne numbered menu entry per container this login may reach. Never a generic
docker exec -it "$1"— the container name must be hard-coded per case, orsudoerscommand-matching below becomes meaningless. -
OS user:
useradd -m -s /usr/local/bin/enter-<app>.sh <login>— mind the 32-character Linux username limit (useradd: invalid user nameis exactly that error); shorten rather than truncate blindly so the name stays meaningful. -
Sudoers, one line per allowed container, no wildcards:
<login> ALL=(root) NOPASSWD: /usr/bin/docker exec -it <container-a> sh <login> ALL=(root) NOPASSWD: /usr/bin/docker exec -it <container-b> shWrite to
/etc/sudoers.d/<login>,chmod 440, and alwaysvisudo -c -f <file>before trusting it — a syntax error here can be silent until someone tries to log in, or can break sudo more broadly if written wrong. -
Teleport role (via
tctl create -f, from wherever the auth server actually runs):kind: role version: v7 metadata: name: access-<app> spec: allow: logins: ["<login>"] node_labels: {"*": "*"} # or {name: ["<Exact Droplet Display Name>"]} to scope to one specific hostCheck how existing analogous roles are assigned (
tctl get users --format=json) — new roles usually get created unassigned, with assignment to a real person happening later as a separate, explicit step. Don't assume you should assign it to anyone. -
Verify — don't just trust that the files look right:
sudo -l -U <login> # must show exactly the intended docker exec commands, nothing elseAlso spot-check that
shactually exists in each target container image (docker exec <container> sh -c 'echo ok') — a container built on a shell-less base image will make the wrapper script fail at the worst time.
Pattern B — single-app host (simpler, no wrapper needed)
When a host only runs one app's containers, there's nothing to scope within the host — a plain restricted login is enough:
useradd -m -s /bin/bash -G docker,adm <login> # docker group for the app's containers, adm for logs
No sudoers file, no wrapper script. Confirm this is really the pattern in use (Pattern A hosts and Pattern B hosts can coexist across a fleet) before picking one over the other.
Pattern C — onboarding a brand-new host (EC2 or any fresh VM)
- Bootstrap plain SSH access first — the host isn't in Teleport yet, so Teleport can't get you in. Try available keys, confirm passwordless sudo:
ssh -i ~/.ssh/<key> <user>@<host-ip> "sudo -n true && echo ok" - Install the Teleport agent at the same major.minor version as the cluster (check with
tctl statusfirst):curl https://apt.releases.teleport.dev/gpg -o /tmp/teleport-pubkey.asc sudo tee /etc/apt/keyrings/teleport-archive-keyring.asc < /tmp/teleport-pubkey.asc echo "deb [signed-by=/etc/apt/keyrings/teleport-archive-keyring.asc] https://apt.releases.teleport.dev/ubuntu $(. /etc/os-release; echo $VERSION_CODENAME) stable/v<MAJOR>" | sudo tee /etc/apt/sources.list.d/teleport.list sudo apt-get update -qq && sudo apt-get install -y teleport - Generate a short-lived join token from the auth server:
tctl tokens add --type=node,db --ttl=15m --format=json - Write
/etc/teleport.yamlon the new host:Add aversion: v3 teleport: nodename: <display-name> data_dir: /var/lib/teleport proxy_server: <proxy>:443 join_params: { token_name: "<token>", method: token } auth_service: { enabled: "no" } proxy_service: { enabled: "no" } ssh_service: enabled: "yes" labels: { name: "<Display Name>", tier: infra, provider: <aws|...> }db_serviceblock too if this host also fronts a database (see Pattern D). sudo systemctl enable teleport && sudo systemctl start teleport, then confirm from your own client —tsh lsshould show the new node within seconds.- Also create the broad host-level login (
devops/whatever the fleet convention is, with passwordless sudo) on the new host so it inherits the same team-wide access as every other host — a new scoped role is additive, never a replacement for existing broad access. Double-check the login actually exists on the OS (a cloud image may only shipubuntu, notdevops).
Pattern D — registering a database (db_service)
db_service:
enabled: "yes"
databases:
- name: <name>
protocol: postgres # or mysql
uri: <host>:<port>
tls:
mode: verify-full # has a real, verifiable TLS cert (managed/cloud DB)
ca_cert_file: /etc/teleport/db-ca/<name>.crt
# OR, for a self-hosted DB with no real cert (a plain dev postgres container, etc):
# mode: insecure
static_labels: { env: prod|dev, tier: infra }
For AWS RDS, download the provider's CA bundle rather than guessing at insecure mode:
curl -s https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem -o /etc/teleport/db-ca/aws-rds-global-bundle.pem
Restart the agent (systemctl restart teleport) after any teleport.yaml edit. If you're SSHed in over Teleport itself, this drops your own session — that's expected, not a failure; reconnect and check systemctl is-active teleport + journalctl -u teleport -n 40 --no-pager for real errors.
Known gotcha: orphaned teleport processes shadow new resources
A host can end up running two teleport processes — the current systemd-managed one, and an older one left over from a migration (started manually or via an old init method, its config file possibly already deleted, invisible to systemctl). The old process keeps heartbeating a stale copy of a resource under the same name, and the auth server can end up not surfacing your freshly-registered same-named resource in tsh ls / tsh db ls even though your new agent's own log says "started successfully."
Symptom: a database or node you just configured doesn't show up, no errors anywhere obvious. Before chasing TLS or RBAC theories, check for a second process:
ps aux | grep '[t]eleport start'
If found and clearly orphaned (no matching systemd unit, config path that no longer exists), confirm with the user, then stop it (kill -TERM <pid>, escalate to -KILL only if it doesn't exit) and restart the real service — the resource should appear cleanly.