Files
Adnan Zahir 98dca237f4 Add infra-ops-toolkit plugin: topology diagram, html-to-pdf, teleport onboarding skills
Distills patterns from the Teleport access topology work: verified-data diagram
building, as-displayed HTML-to-PDF export via headless Chromium, and
container-scoped/host/database onboarding into an existing Teleport cluster.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R3ZTfgQrkR3q8DSEoZAmvs
2026-09-08 16:11:36 +07:00

8.9 KiB

name, description
name description
teleport-onboard Onboard a new resource into a Teleport cluster — a VPS/droplet, an EC2 (or any fresh VM) host, a container-scoped least-privilege app login on a multi-app host, or a database (db_service) — following an existing team's established patterns rather than inventing a new access model. Use when asked to register, join, or onboard something into Teleport, add scoped/restricted access to a specific app or container, or when Teleport resources (nodes/databases/roles) don't match what's actually running.

Teleport resource onboarding

First: discover the existing pattern, don't assume one

Before creating anything, SSH into a host that's already onboarded the way you intend to onboard the new one, and read what's actually there:

cat /etc/passwd                          # look for non-standard login shells — a sign of scoped access
sudo cat /etc/sudoers.d/<candidate-user>  # the exact commands that login is allowed to run
sudo cat /etc/teleport.yaml               # or find it: /opt/teleport/config/, /etc/teleport/ — it varies
ps aux | grep teleport                    # more than one teleport process on a host is a real bug (see gotcha below)

Do not invent a new access model if one already exists — mirror it exactly, including its exact sudoers syntax, its wrapper-script conventions, and its Teleport role shape.

Where tctl actually runs: the cluster's auth server is often just one specific host's Docker container (look for an image like teleport-distroless in docker ps), not a separate admin machine. sudo docker exec teleport tctl ... gives full cluster admin without a separate tctl binary or identity file — check for this before assuming you need new credentials.

Pattern A — container-scoped forced-login-shell (multi-app host, least privilege per app)

Use when one host runs several unrelated apps as containers and different people need access to only their own app's container, never the host shell or a sibling's container.

  1. Wrapper script /usr/local/bin/enter-<app>.sh (root:root, 0755):

    #!/bin/bash
    set -euo pipefail
    # Forced login shell — must never fall back to a real host shell.
    if [ ! -t 0 ]; then echo "Interactive TTY required." >&2; exit 1; fi
    echo "=== <app> - container access ==="
    echo "1) <container-a>"
    echo "2) <container-b>"
    read -rp "Pilih [1-2]: " choice
    case "$choice" in
      1) exec sudo /usr/bin/docker exec -it <container-a> sh ;;
      2) exec sudo /usr/bin/docker exec -it <container-b> sh ;;
      *) echo "Pilihan tidak valid."; exit 1 ;;
    esac
    

    One numbered menu entry per container this login may reach. Never a generic docker exec -it "$1" — the container name must be hard-coded per case, or sudoers command-matching below becomes meaningless.

  2. OS user: useradd -m -s /usr/local/bin/enter-<app>.sh <login> — mind the 32-character Linux username limit (useradd: invalid user name is exactly that error); shorten rather than truncate blindly so the name stays meaningful.

  3. Sudoers, one line per allowed container, no wildcards:

    <login> ALL=(root) NOPASSWD: /usr/bin/docker exec -it <container-a> sh
    <login> ALL=(root) NOPASSWD: /usr/bin/docker exec -it <container-b> sh
    

    Write to /etc/sudoers.d/<login>, chmod 440, and always visudo -c -f <file> before trusting it — a syntax error here can be silent until someone tries to log in, or can break sudo more broadly if written wrong.

  4. Teleport role (via tctl create -f, from wherever the auth server actually runs):

    kind: role
    version: v7
    metadata:
      name: access-<app>
    spec:
      allow:
        logins: ["<login>"]
        node_labels: {"*": "*"}   # or {name: ["<Exact Droplet Display Name>"]} to scope to one specific host
    

    Check how existing analogous roles are assigned (tctl get users --format=json) — new roles usually get created unassigned, with assignment to a real person happening later as a separate, explicit step. Don't assume you should assign it to anyone.

  5. Verify — don't just trust that the files look right:

    sudo -l -U <login>     # must show exactly the intended docker exec commands, nothing else
    

    Also spot-check that sh actually exists in each target container image (docker exec <container> sh -c 'echo ok') — a container built on a shell-less base image will make the wrapper script fail at the worst time.

Pattern B — single-app host (simpler, no wrapper needed)

When a host only runs one app's containers, there's nothing to scope within the host — a plain restricted login is enough:

useradd -m -s /bin/bash -G docker,adm <login>   # docker group for the app's containers, adm for logs

No sudoers file, no wrapper script. Confirm this is really the pattern in use (Pattern A hosts and Pattern B hosts can coexist across a fleet) before picking one over the other.

Pattern C — onboarding a brand-new host (EC2 or any fresh VM)

  1. Bootstrap plain SSH access first — the host isn't in Teleport yet, so Teleport can't get you in. Try available keys, confirm passwordless sudo:
    ssh -i ~/.ssh/<key> <user>@<host-ip> "sudo -n true && echo ok"
    
  2. Install the Teleport agent at the same major.minor version as the cluster (check with tctl status first):
    curl https://apt.releases.teleport.dev/gpg -o /tmp/teleport-pubkey.asc
    sudo tee /etc/apt/keyrings/teleport-archive-keyring.asc < /tmp/teleport-pubkey.asc
    echo "deb [signed-by=/etc/apt/keyrings/teleport-archive-keyring.asc] https://apt.releases.teleport.dev/ubuntu $(. /etc/os-release; echo $VERSION_CODENAME) stable/v<MAJOR>" | sudo tee /etc/apt/sources.list.d/teleport.list
    sudo apt-get update -qq && sudo apt-get install -y teleport
    
  3. Generate a short-lived join token from the auth server:
    tctl tokens add --type=node,db --ttl=15m --format=json
    
  4. Write /etc/teleport.yaml on the new host:
    version: v3
    teleport:
      nodename: <display-name>
      data_dir: /var/lib/teleport
      proxy_server: <proxy>:443
      join_params: { token_name: "<token>", method: token }
    auth_service: { enabled: "no" }
    proxy_service: { enabled: "no" }
    ssh_service:
      enabled: "yes"
      labels: { name: "<Display Name>", tier: infra, provider: <aws|...> }
    
    Add a db_service block too if this host also fronts a database (see Pattern D).
  5. sudo systemctl enable teleport && sudo systemctl start teleport, then confirm from your own client — tsh ls should show the new node within seconds.
  6. Also create the broad host-level login (devops/whatever the fleet convention is, with passwordless sudo) on the new host so it inherits the same team-wide access as every other host — a new scoped role is additive, never a replacement for existing broad access. Double-check the login actually exists on the OS (a cloud image may only ship ubuntu, not devops).

Pattern D — registering a database (db_service)

db_service:
  enabled: "yes"
  databases:
  - name: <name>
    protocol: postgres   # or mysql
    uri: <host>:<port>
    tls:
      mode: verify-full          # has a real, verifiable TLS cert (managed/cloud DB)
      ca_cert_file: /etc/teleport/db-ca/<name>.crt
    # OR, for a self-hosted DB with no real cert (a plain dev postgres container, etc):
    #   mode: insecure
    static_labels: { env: prod|dev, tier: infra }

For AWS RDS, download the provider's CA bundle rather than guessing at insecure mode:

curl -s https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem -o /etc/teleport/db-ca/aws-rds-global-bundle.pem

Restart the agent (systemctl restart teleport) after any teleport.yaml edit. If you're SSHed in over Teleport itself, this drops your own session — that's expected, not a failure; reconnect and check systemctl is-active teleport + journalctl -u teleport -n 40 --no-pager for real errors.

Known gotcha: orphaned teleport processes shadow new resources

A host can end up running two teleport processes — the current systemd-managed one, and an older one left over from a migration (started manually or via an old init method, its config file possibly already deleted, invisible to systemctl). The old process keeps heartbeating a stale copy of a resource under the same name, and the auth server can end up not surfacing your freshly-registered same-named resource in tsh ls / tsh db ls even though your new agent's own log says "started successfully."

Symptom: a database or node you just configured doesn't show up, no errors anywhere obvious. Before chasing TLS or RBAC theories, check for a second process:

ps aux | grep '[t]eleport start'

If found and clearly orphaned (no matching systemd unit, config path that no longer exists), confirm with the user, then stop it (kill -TERM <pid>, escalate to -KILL only if it doesn't exit) and restart the real service — the resource should appear cleanly.