--- name: teleport-onboard description: Onboard a new resource into a Teleport cluster — a VPS/droplet, an EC2 (or any fresh VM) host, a container-scoped least-privilege app login on a multi-app host, or a database (db_service) — following an existing team's established patterns rather than inventing a new access model. Use when asked to register, join, or onboard something into Teleport, add scoped/restricted access to a specific app or container, or when Teleport resources (nodes/databases/roles) don't match what's actually running. --- # Teleport resource onboarding ## First: discover the existing pattern, don't assume one Before creating anything, SSH into a host that's already onboarded the way you intend to onboard the new one, and read what's actually there: ```bash cat /etc/passwd # look for non-standard login shells — a sign of scoped access sudo cat /etc/sudoers.d/ # the exact commands that login is allowed to run sudo cat /etc/teleport.yaml # or find it: /opt/teleport/config/, /etc/teleport/ — it varies ps aux | grep teleport # more than one teleport process on a host is a real bug (see gotcha below) ``` Do not invent a new access model if one already exists — mirror it exactly, including its exact `sudoers` syntax, its wrapper-script conventions, and its Teleport role shape. **Where `tctl` actually runs**: the cluster's auth server is often just one specific host's Docker container (look for an image like `teleport-distroless` in `docker ps`), not a separate admin machine. `sudo docker exec teleport tctl ...` gives full cluster admin without a separate tctl binary or identity file — check for this before assuming you need new credentials. ## Pattern A — container-scoped forced-login-shell (multi-app host, least privilege per app) Use when one host runs several unrelated apps as containers and different people need access to *only their own* app's container, never the host shell or a sibling's container. 1. **Wrapper script** `/usr/local/bin/enter-.sh` (root:root, `0755`): ```bash #!/bin/bash set -euo pipefail # Forced login shell — must never fall back to a real host shell. if [ ! -t 0 ]; then echo "Interactive TTY required." >&2; exit 1; fi echo "=== - container access ===" echo "1) " echo "2) " read -rp "Pilih [1-2]: " choice case "$choice" in 1) exec sudo /usr/bin/docker exec -it sh ;; 2) exec sudo /usr/bin/docker exec -it sh ;; *) echo "Pilihan tidak valid."; exit 1 ;; esac ``` One numbered menu entry per container this login may reach. Never a generic `docker exec -it "$1"` — the container name must be hard-coded per case, or `sudoers` command-matching below becomes meaningless. 2. **OS user**: `useradd -m -s /usr/local/bin/enter-.sh ` — mind the **32-character Linux username limit** (`useradd: invalid user name` is exactly that error); shorten rather than truncate blindly so the name stays meaningful. 3. **Sudoers**, one line per allowed container, no wildcards: ``` ALL=(root) NOPASSWD: /usr/bin/docker exec -it sh ALL=(root) NOPASSWD: /usr/bin/docker exec -it sh ``` Write to `/etc/sudoers.d/`, `chmod 440`, and **always** `visudo -c -f ` before trusting it — a syntax error here can be silent until someone tries to log in, or can break sudo more broadly if written wrong. 4. **Teleport role** (via `tctl create -f`, from wherever the auth server actually runs): ```yaml kind: role version: v7 metadata: name: access- spec: allow: logins: [""] node_labels: {"*": "*"} # or {name: [""]} to scope to one specific host ``` Check how existing analogous roles are assigned (`tctl get users --format=json`) — new roles usually get created *unassigned*, with assignment to a real person happening later as a separate, explicit step. Don't assume you should assign it to anyone. 5. **Verify** — don't just trust that the files look right: ```bash sudo -l -U # must show exactly the intended docker exec commands, nothing else ``` Also spot-check that `sh` actually exists in each target container image (`docker exec sh -c 'echo ok'`) — a container built on a shell-less base image will make the wrapper script fail at the worst time. ## Pattern B — single-app host (simpler, no wrapper needed) When a host only runs one app's containers, there's nothing to scope *within* the host — a plain restricted login is enough: ```bash useradd -m -s /bin/bash -G docker,adm # docker group for the app's containers, adm for logs ``` No sudoers file, no wrapper script. Confirm this is really the pattern in use (Pattern A hosts and Pattern B hosts can coexist across a fleet) before picking one over the other. ## Pattern C — onboarding a brand-new host (EC2 or any fresh VM) 1. Bootstrap plain SSH access first — the host isn't in Teleport yet, so Teleport can't get you in. Try available keys, confirm passwordless sudo: ```bash ssh -i ~/.ssh/ @ "sudo -n true && echo ok" ``` 2. Install the Teleport agent at the **same major.minor version as the cluster** (check with `tctl status` first): ```bash curl https://apt.releases.teleport.dev/gpg -o /tmp/teleport-pubkey.asc sudo tee /etc/apt/keyrings/teleport-archive-keyring.asc < /tmp/teleport-pubkey.asc echo "deb [signed-by=/etc/apt/keyrings/teleport-archive-keyring.asc] https://apt.releases.teleport.dev/ubuntu $(. /etc/os-release; echo $VERSION_CODENAME) stable/v" | sudo tee /etc/apt/sources.list.d/teleport.list sudo apt-get update -qq && sudo apt-get install -y teleport ``` 3. Generate a short-lived join token from the auth server: ```bash tctl tokens add --type=node,db --ttl=15m --format=json ``` 4. Write `/etc/teleport.yaml` on the new host: ```yaml version: v3 teleport: nodename: data_dir: /var/lib/teleport proxy_server: :443 join_params: { token_name: "", method: token } auth_service: { enabled: "no" } proxy_service: { enabled: "no" } ssh_service: enabled: "yes" labels: { name: "", tier: infra, provider: } ``` Add a `db_service` block too if this host also fronts a database (see Pattern D). 5. `sudo systemctl enable teleport && sudo systemctl start teleport`, then confirm from your own client — `tsh ls` should show the new node within seconds. 6. Also create the broad host-level login (`devops`/whatever the fleet convention is, with passwordless sudo) on the new host so it inherits the same team-wide access as every other host — a new scoped role is **additive**, never a replacement for existing broad access. Double-check the login actually exists on the OS (a cloud image may only ship `ubuntu`, not `devops`). ## Pattern D — registering a database (`db_service`) ```yaml db_service: enabled: "yes" databases: - name: protocol: postgres # or mysql uri: : tls: mode: verify-full # has a real, verifiable TLS cert (managed/cloud DB) ca_cert_file: /etc/teleport/db-ca/.crt # OR, for a self-hosted DB with no real cert (a plain dev postgres container, etc): # mode: insecure static_labels: { env: prod|dev, tier: infra } ``` For AWS RDS, download the provider's CA bundle rather than guessing at `insecure` mode: ```bash curl -s https://truststore.pki.rds.amazonaws.com/global/global-bundle.pem -o /etc/teleport/db-ca/aws-rds-global-bundle.pem ``` Restart the agent (`systemctl restart teleport`) after any `teleport.yaml` edit. **If you're SSHed in over Teleport itself, this drops your own session** — that's expected, not a failure; reconnect and check `systemctl is-active teleport` + `journalctl -u teleport -n 40 --no-pager` for real errors. ## Known gotcha: orphaned teleport processes shadow new resources A host can end up running **two teleport processes** — the current systemd-managed one, and an older one left over from a migration (started manually or via an old init method, its config file possibly already deleted, invisible to `systemctl`). The old process keeps heartbeating a stale copy of a resource under the same name, and the auth server can end up not surfacing your freshly-registered same-named resource in `tsh ls` / `tsh db ls` even though your new agent's own log says "started successfully." Symptom: a database or node you just configured doesn't show up, no errors anywhere obvious. Before chasing TLS or RBAC theories, check for a second process: ```bash ps aux | grep '[t]eleport start' ``` If found and clearly orphaned (no matching systemd unit, config path that no longer exists), confirm with the user, then stop it (`kill -TERM `, escalate to `-KILL` only if it doesn't exit) and restart the real service — the resource should appear cleanly.