Commit Graph

937 Commits

Author SHA1 Message Date
Kristoffer Dalby eeaac680be mapper: drop DNSConfig from policy responses
It forced every client into a full netmap rebuild; the resolver race it
guarded against was a client bug (tailscale/tailscale#19749).
2026-09-30 19:03:13 +02:00
Kristoffer Dalby 357f34a778 state: send a node's DNS config with its own refresh
Drained at dispatch for nodeAttrs changes from any policy, user or node
update, and for hostname/OS changes feeding NextDNS device metadata.
2026-09-30 19:03:13 +02:00
Kristoffer Dalby 59030f380d servertest: cover a node's DNS config following its NextDNS inputs
Profile via policy reload or tag change, device metadata via hostname;
a hostname change today waits for the next policy response to reach DNS.
2026-09-30 19:03:13 +02:00
Kristoffer Dalby 60d42808b8 state: keep node removal apart from policy refresh in DeleteNode
Peers get the deletion as an explicit PeersRemoved change, independent of
their sent-peers tracking.

Updates tailscale/tailscale#15660
2026-09-30 19:03:13 +02:00
Kristoffer Dalby 1bd62737b9 mapper: send removed peers as their own delta
Removals derived from a policy change, full update or reconnect rode a
response whose DNSConfig, SSHPolicy, Node or Peers force a full rebuild.

Updates tailscale/tailscale#15660
2026-09-30 19:03:13 +02:00
Kristoffer Dalby baca8309f1 servertest: reproduce removed peers missing from IPN bus deltas
Delta-only IPN bus watchers (the Android app since 1.100) miss removals
riding a full-rebuild response: deleted, policy-hidden, or while offline.

Updates tailscale/tailscale#15660
2026-09-30 19:03:13 +02:00
Kristoffer Dalby 90d3e0dd73 policy/v2: check cached per-node results against a fresh compile
Random node writes; every FilterForNode, MatchersForNode and SSHPolicy
read must match a PolicyManager built from the same nodes.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 4fd766a431 policy/v2: leave the policy manager unchanged when a recompile fails
updateLocked runs every fallible step before writing pm, so a failed
SetUsers or SetNodes no longer leaves half a new filter live.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby edf5cc994e policy/v2: drop per-node filter caches when users change
autogroup:self sources resolve users by name outside the filter hash,
so SetUsers left stale self rules cached and reported no change.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 9526d74d61 state: send PolicyChange from SetApprovedRoutes only when visibility moved
Otherwise NodeAdded. Any SubnetRoutes/ExitRoutes change still bumps
NodesGeneration, so in practice only unannounced approvals narrow.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby c26f6e5255 state: approve routes inside the map request write
One NodeStore write, one peer build, one row update per
auto-approved map request, instead of a second SetApprovedRoutes write.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 21f6e46fb8 state: refresh policy nodes inside the peer map build
One peer build per tag/user/IP/route write; callers detect policy moves
via NodesGeneration. Per-node caches only store results for the node
pm holds, so a mapper reading mid-build cannot pin a stale filter.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 311d9323e0 policy/v2: count node-driven recompiles
Lets a caller detect that its node write moved the policy when the
SetNodes ran on another goroutine.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 05fc17bc85 state: decide whole-peer updates at each caller, not in persist
persistNodeAndRefreshPolicy no longer fabricates NodeAdded; RenameNode
and SetNodeTags add it since both are peer visible.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 86b4f7c430 state,policy: add peer map and NodeStore write benchmarks
Cover BuildPeerMap, SetNodes, NodeStore writes and
UpdateNodeFromMapRequest across node counts and policy shapes.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby d813144f3b state: fail when a Node field is not classified for peer map reuse
Every exported types.Node field is pinned to relation/election/payload
in nodeFieldImpact; a new field fails until classified, and each
relation/election field gets a mutation check against updateChanges.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 1613023be9 state: check NodeStore adjacency against a full peer map build
Rapid test drives random writes through a real NodeStore and policy
manager, checking adjacency before and after syncPolicy against a
BuildPeerMap from a fresh policy manager.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 9ce160ed5b policy/v2: keep the live policy when SetPolicy fails to compile
SetPolicy assigned the new policy before compiling it, so a compile that
failed partway left the rejected filter live while the stored policy stayed
old.
2026-09-30 17:21:02 +02:00
Kristoffer Dalby 3417e4cb77 policy/v2: add via exit node captures with unrelated rules
SaaS keeps via exit steering unless a rule reaches the internet, and
sends exit nodes every rule, with or without via.

Updates #3493
2026-09-30 17:20:46 +02:00
Kristoffer Dalby adbf4ba0f0 servertest: compare via capture filters as rule sets
Headscale merges rules sharing sources; SaaS does not. Compare
(src, dst, ports) triples, keyed by every captured node's addresses.
2026-09-30 17:20:46 +02:00
Kristoffer Dalby 7f5fdbe5d2 policy/v2: send exit nodes every user's autogroup:self rules
Same as other rules: exit routes contain every self destination.

Updates #3493
2026-09-30 17:20:46 +02:00
Kristoffer Dalby 61d1599393 policyutil: send approved exit nodes every rule
Tailscale SaaS treats exit routes like subnet routes when reducing
filters; 0.0.0.0/0 contains every dst, so exit nodes get all rules.

Updates #3493
2026-09-30 17:20:46 +02:00
Kristoffer Dalby 7efd22d0bb policy/v2: let autogroup:internet rules lift via exit steering
A regular rule to the internet allows every exit node, but
autogroup:internet resolves to no prefix, so it never matched.

Updates #3493
2026-09-30 17:20:46 +02:00
Kristoffer Dalby c700b32eec policy/v2: stop narrow rules from undoing via exit steering
Every address overlaps 0.0.0.0/0, so any regular rule matching the
viewer dropped the exclusion; only a wildcard dst now does.

Fixes #3493
2026-09-30 17:20:46 +02:00
Kristoffer Dalby 35a90e0018 derp: serve a self-signed TLS listener for tests
HEADSCALE_DEBUG_INSECURE_TLS_LISTEN_ADDR serves the router over throwaway TLS
and marks the embedded region InsecureForTests: TLS DERP for CA-less clients
(tailscale-rs, Android), plus the :443 noise fallback Go redials use.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby fbb04c5611 testcapture: read captures through util.ReadFileByExt
Drops its private HuJSON decoder.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby 906016beb4 dns: pick the extra_records_path format by file extension
HuJSON and YAML records work too; an unknown extension fails instead of being
parsed as JSON.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby dd8ce695cf derp: pick the derp.paths format by file extension
Adds JSON and HuJSON maps. JSON read as YAML decoded to an empty map; unknown
extensions and empty maps now fail at startup.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby d20a412079 util: decode files by their extension
UnmarshalByExt picks JSON, HuJSON or YAML from the file name; content can't
tell them apart, since YAML parses JSON syntax.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby 90732bdaaf derp: drop regions set to null in derp.paths
19d5d9de cloned regions while merging and skipped nil ones, so the documented
null-removal recipe silently kept the region.
2026-09-28 21:42:19 +02:00
Kristoffer Dalby 393dd3e2d9 db: store all credentials in one SHA-256-hashed table
API keys, pre-auth keys and OAuth clients/tokens share one table and verify
path. Secrets carry 256 bits of crypto/rand entropy, so a SHA-256 digest
needs no stretching; bcrypt/argon2id rows rehash on use until 0.32.
2026-09-26 00:33:12 +02:00
Kristoffer Dalby e90500e3a9 db: drop migrations predating 0.29
Only 0.29.x -> 0.30 upgrades are supported; checkMinimumMigration refuses
older databases instead of silently skipping the removed steps.
2026-09-26 00:33:12 +02:00
Michael Lopez 04d1e3c83f db: name the nodes that block a user deletion
The error only said the user still has nodes, which the CLI prompt did
not mention at all. Wrap ErrUserStillHasNodes with the ID and hostname
of every blocking node so the operator knows what to remove. Run the
DestroyUser test table on Postgres as well as SQLite, since the two
schemas define different foreign-key actions; the Postgres variant
skips without a local server.
2026-09-25 23:56:02 +02:00
MsfPablo d60bac5c79 db: reject unknown OAuth client scopes at creation (#3424)
Unknown scopes grant nothing, so clients were silently under-privileged.
Report every invalid scope at once.

Fixes #3406
2026-09-25 21:24:37 +02:00
Kristoffer Dalby eebeab3c1e docs: document joining nodes with an OAuth client
Prefix swap, baseURL, every attribute; examples for tailscale up,
container, tsnet, GitHub Action.
2026-09-25 17:09:19 +02:00
Kristoffer Dalby bb3c0b4cd6 servertest: drive tailscale client oauthkey flow against v2 API
Real feature/oauthkey resolver via tsnet: tskey-client- secret,
?baseURL attributes, CreateKey, register re-advertising key tags.
2026-09-25 17:09:19 +02:00
Kristoffer Dalby 61a626eab7 state: accept advertise-tags subset of a pre-auth key's tags
tailscale client's OAuth authkey flow re-advertises the key's tags.
Reject only tags the key lacks, for new nodes and re-registrations.
2026-09-25 17:09:19 +02:00
Kristoffer Dalby 69e84c356d db: accept tskey-client- prefix for OAuth client auth
The upstream tailscale client only runs its OAuth client-credentials exchange
for secrets prefixed tskey-client-, so accept it as an alias for hskey-client-.
The prefix is only a label sliced off before lookup, so the same stored client
authenticates under either; lets the official client and GitHub Action mint auth
keys against headscale.
2026-09-25 17:09:19 +02:00
Igor Serganov 0661c6f540 db: return non-NotFound errors from user lookups (#3480) 2026-09-23 15:24:57 +02:00
Michael Lopez fdcdebb392 db: delete pre-auth key on its own savepoint 2026-09-23 15:24:22 +02:00
Kristoffer Dalby 4cad0b564e types: lift configValidator across LoadServerConfig sub-builders
Sub-builders called after validateServerConfig (prefixV4, prefixV6,
allocation strategy, dns, oidc client secret/path, isSafeServerURL)
each returned the first error. So an operator that fixed one
issue saw the next one only on the next startup.

Lift one *configValidator across the entire LoadServerConfig flow.
validateServerConfigInto(v) populates it; each sub-builder's error
is wrapped in a structured *ConfigError (with the original sentinel
kept on the Cause field, so errors.Is against errOidcMutuallyExclusive,
errServerURLSame/Suffix, ErrNoPrefixConfigured, and
ErrInvalidAllocationStrategy keeps working). v.Err() is checked once,
right before the &Config{} construction.

TestReadConfig/base-domain-in-server-url-err matched the old sentinel
wording; flipped to the new structured Reason. Added
TestLoadServerConfig_CollectsAcrossSubBuilders to lock the wiring:
four sub-builder failures from a single config produce four
*ConfigErrors in the report.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 8a12c572ef types: route deprecation fatals through configValidator
deprecator.Log() called log.Fatal directly, so a single deprecated
key killed the process before any other validation rule could
report. Operators fixing one issue at a time.

Replace Log() with Apply(*configValidator). Warns still go to
log.Warn; fatals are pushed onto the validator as *ConfigError so
they merge with the rest of the config-load report. The free-form
strings collected via set.Set are gone; deprecator now stores
typed deprecation{OldKey, NewKey} records and Apply renders them
into the structured ConfigError shape.

Move the early-return PKCE check into a new validatePKCEConfig
modular validator for the same reason.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 2114ced0d5 cli: classify Serve errors with operator hints
A *ListenerBindError that wraps syscall.EADDRINUSE now ends with a
"sudo ss -tlnp 'sport = :PORT'" pointer, and one wrapping
syscall.EACCES with a CAP_NET_BIND_SERVICE / setcap pointer. The
underlying chain is preserved via fmt.Errorf("%w"), so errors.Is /
errors.As continue to walk to the typed bind error and the syscall
errno.

Drop the "headscale ran into an error and had to shut down" wrap,
which only restated the symptom.

Export types.PortFromAddr so the classifier can render the port
number in the hint.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 61021c739a app: route ACME HTTP-01 listener through errgroup
The HTTP-01 challenge listener was launched in an orphan goroutine that
called log.Fatal on bind failure, which bypassed the signal-handler
shutdown path and made bind errors look identical to the main HTTP
listener.

Bind eagerly via net.ListenConfig in getTLSSettings, return the listener
+ http.Server in a tlsBundle, register the Serve loop with the existing
errgroup, and call Shutdown alongside the other servers. Bind failures
now surface as ListenerBindError(Listener:"ACME HTTP-01 challenge")
through the normal error chain.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 37e1904924 hscontrol: name listener bind failures via ListenerBindError
Replace the three "binding to TCP address" wraps with typed
ListenerBindError values that carry the listener role, the YAML key
that drove the address, and the resolved address. The wrapped error
chain still walks to syscall.EADDRINUSE / EACCES so existing callers
that match those sentinels keep working; the difference is that the
operator now sees which listener failed instead of an unattributed
"binding to TCP address" line.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 43ad10da52 types: add ListenerBindError for typed bind failures
ListenerBindError wraps a TCP listener bind failure with the listener
name and the YAML key that drove its address. Preserves the underlying
*net.OpError via Unwrap so errors.Is(err, syscall.EADDRINUSE) keeps
working through any number of fmt.Errorf("%w") wraps. Lets the
top-level CLI classifier match listener failures with errors.As
instead of string matching.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 7865430419 types: move sub-builder log.Fatal sites into validators
derpConfig, databaseConfig, and dnsToTailcfgDNS used log.Fatal to
reject invalid combinations the moment they were observed. Lift those
checks into modular validators (validateDERPConfig,
validateDatabaseConfig, validateMagicDNSConfig) called from
validateServerConfig so each violation lands in the same configValidator
collector and renders as a structured ConfigError next to all other
config feedback. The sub-builder functions trust validation has run
and no longer crash the process.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 1b4b79901a types: rework validateServerConfig to use ConfigError
Convert the errorText accumulator inside validateServerConfig to the
typed configValidator pattern. Each rule violation now renders as a
multi-line block naming the YAML key, the value the operator wrote,
and a hint pointing at the resolution. The dns.extra_records mutex
that previously crashed via log.Fatal is folded into the same
collector so an operator sees every problem in one go instead of
fixing them one startup attempt at a time.

Test wantErr assertions on validation output switch to substring
match because the rendered errors are now multi-line.

Updates #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 284a4891d7 types: validate listener address collisions
Reject configurations where two configured TCP listeners would bind
the same kernel socket. Covers every pair of listen_addr,
grpc_listen_addr, metrics_listen_addr, and tls_letsencrypt_listen
(when HTTP-01 ACME is in use). Each violation renders as a structured
ConfigError naming both YAML keys, the values the operator wrote, and
a hint pointing at the canonical setup.

Closes the symptom in #3227: the misconfiguration is now rejected at
config-load time with a self-explaining error, instead of failing at
runtime with a misleading "address already in use" log because the
ACME challenge listener and the main HTTP listener competed for the
same port inside the same process.

Fixes #3227
2026-09-23 15:18:34 +02:00
Kristoffer Dalby 6031194c36 types: add listener address helpers
portFromAddr resolves numeric and named ports (":http", ":https")
without touching /etc/services or the resolver. listenersOverlap
follows kernel rules: same port plus a wildcard host on either side,
or same port plus identical specific host, both count as collision;
different specific hosts on the same port do not. Set an explicit
viper default for tls_letsencrypt_listen so a minimal config still
resolves to ":http".

Updates #3227
2026-09-23 15:18:34 +02:00