219 lines
No EOL
15 KiB
Markdown
219 lines
No EOL
15 KiB
Markdown
# Handles and Short IDs
|
||
|
||
This is the syntax-level reference for two related but separate naming schemes. [federation.md](federation.md)'s
|
||
"Unique Handles" section covers *why* Toolshed hands out handles at all and what each kind (user,
|
||
group, tag/property/category) means conceptually; this document covers the parsing rules those
|
||
handles have to follow once they're written down or embedded somewhere - legal characters and
|
||
escaping. It also covers short ids end to end: a separate scheme for packing small integer id
|
||
chains into a compact token.
|
||
|
||
## Handle syntax
|
||
|
||
### Reserved characters
|
||
|
||
A username ends up embedded, unescaped, in several composite formats beyond its own handle, so it
|
||
can't contain any character that already means something else in one of those: `@` (the
|
||
user/domain separator in a user handle), `#` (the group-handle prefix, and the origin/type
|
||
separator in a classification handle, see federation.md's Tags, Properties, and Categories
|
||
section), `:` (the kind/id delimiter in a User-Qualified ID (see below), the type/name delimiter in
|
||
a classification handle, the domain/token delimiter in a Domain-Qualified Short ID (see below), and
|
||
the delimiter in a signed request's `Authorization` header), `+` (reserved as the URL-embedding
|
||
escape for `#`, see below), `~` (the short-id token prefix, see Short IDs below, and a
|
||
collision-disambiguation suffix delimiter on a tag's origin), and `/` (the path-segment delimiter
|
||
every handle and id ultimately sits next to once embedded in a URL). This has to be enforced by an
|
||
explicit validator rather than left to a framework default, which doesn't draw the line in the same
|
||
place.
|
||
|
||
### Embedding a `#`-bearing handle in a URL
|
||
|
||
A literal `#` can't appear unescaped in a URL path segment: per RFC 3986, `#` starts the URI's
|
||
fragment component, so any URL-parsing client (a browser, a QR scanner, a link preview) treats
|
||
everything from the first unescaped `#` onward as a fragment and never sends it to the server at
|
||
all, before a request is even made, not merely a server-side quirk to work around. The usual fix is
|
||
to percent-encode it (`%23`), but that's exactly the encode/decode step a self-contained item URL
|
||
(a physical label, a link shared outside the app) is designed to avoid for anything that sits
|
||
directly in a path segment (`@` needs no such treatment). Instead, whenever a handle containing a
|
||
`#` (a group handle, or a tag/property/category
|
||
handle) has to appear as a raw URL path segment, substitute `+` for `#` in that rendering only:
|
||
`#groupname@domain` becomes `+groupname@domain` in a URL, and `origin#type:name` becomes
|
||
`origin+type:name`. This is a URL-embedding convention, not a second handle format: the canonical
|
||
handle, the one used in the API, in signed requests, in the database, and everywhere else a handle
|
||
is written or displayed, is unchanged and is still written `#groupname@domain`. Reversing the
|
||
substitution when parsing a path segment back into a handle is unambiguous only because `+` is
|
||
otherwise forbidden in every field a handle is built from (see Reserved characters above); if a
|
||
group name or tag name could itself contain a literal `+`, it would be indistinguishable from an
|
||
escaped `#` once decoded.
|
||
|
||
### User-Qualified ID
|
||
|
||
Owned things (an item, a storage location, ...) can be referred to outside their owner's own
|
||
account: anything whose id is only unique within one owner's own numbering needs an owner-qualified
|
||
form to resolve globally.
|
||
|
||
A **User-Qualified ID** is an owner handle (a user handle, or a group handle for a group-owned
|
||
thing) with a kind letter and a local id appended, separated by a single `:`:
|
||
`<owner-handle>:<kind><local-id>`, e.g. `alice@example.com:i42`. This is the same `origin#type:name`
|
||
pattern classification handles already use, built on `:` instead of `#` so it avoids the URL
|
||
fragment-escaping problem (see Embedding a `#`-bearing handle in a URL above) — a User-Qualified ID
|
||
isn't meant to sit directly in a URL path segment.
|
||
|
||
Its payload is a plain decimal local id with nothing self-describing baked in, unlike a Short ID's
|
||
bit-packed payload, so the owner is spelled out as a full handle rather than just a domain:
|
||
`example.com:i42` would be ambiguous between every user on that domain with local item id `42`;
|
||
`alice@example.com:i42` isn't. An owner handle already carries its domain, so a User-Qualified ID is
|
||
cross-domain-resolvable as written, with no separate domain-qualified wrapper needed.
|
||
|
||
The kind letter is a small, closed registry. A group handle already looks visibly different from a
|
||
user handle (`#` prefix), so unlike the short id kind registry below, this one doesn't need separate
|
||
letters for a user-owned vs. group-owned kind — the owner half already says which it is.
|
||
|
||
| letter | kind | notes |
|
||
|---|---|---|
|
||
| `i` | item | covers both a user-owned item and a group-owned item |
|
||
| `s` | storage location | covers both a user-owned and a group-owned storage location |
|
||
|
||
`workflow` has no letter assigned; the registry can grow without breaking anything already printed.
|
||
Kinds with no owner (`category`, `group`, `file`) don't fit this scheme: `category` and `group` each
|
||
already have their own dedicated handle form. `file` has neither an owner to qualify by nor a
|
||
dedicated handle of its own; a Domain-Qualified Short ID (below) is the fallback for it.
|
||
|
||
## Short IDs
|
||
|
||
A general encoding for turning a small, fixed-shape list of integers into a compact, URL-safe
|
||
token, with no server-side lookup table involved: the code *is* the data, nothing is stored
|
||
server-side to make it resolvable. Not item-specific; anything currently addressed by a short chain
|
||
of small integers is a candidate.
|
||
|
||
### Shape: a kind tag, then a fixed list of integers
|
||
|
||
Every short id starts with a small, fixed-width **kind** tag saying which schema the rest of the
|
||
bits should be read against, followed by exactly the integer fields that kind's schema calls for,
|
||
in a fixed order. `kind` is a small, closed, slow-growing set, so it doesn't need to be
|
||
self-delimiting the way the integer fields do: 2 bits directly name kinds 0-2, and the all-ones
|
||
value (3) is an escape meaning "the real kind follows as the next field, offset by this direct
|
||
range" - so kind 3 is encoded as escape + chunked-int `0`, kind 4 as escape + `1`, and so on. This
|
||
costs nothing for a kind that already fits in the direct range, and keeps the tag itself extensible
|
||
forever without ever having to widen it out from under codes that were already printed. A narrow
|
||
tag only pays off if kind usage is actually skewed the way id values are (a few kinds dominate),
|
||
which is why the registry below is ordered by expected frequency, cheapest (most-used) kind first.
|
||
Kind ids are scoped **per arity**, not globally: the same id can (and does) name a different kind
|
||
depending on how many fields follow it, since the arity is always known before the kind id needs
|
||
disambiguating:
|
||
|
||
| kind id | arity | name | fields | notes |
|
||
|---|---|---|---|---|
|
||
| 0 | 2 | `item` | `owner_identity_id`, `item_local_id` | dominant case - a personally-owned item, the primary physical-label use case |
|
||
| 0 | 1 | `category` | `category_id` | label-adjacent (tagging); global, no owner - shares id 0 with `item` since the two never need the same arity |
|
||
| 1 | 2 | `group_item` | `owner_group_id`, `item_local_id` | same use case as `item`, but for a group-owned item |
|
||
| 1 | 1 | `group` | `group_id` | shared even less often; global, no owner |
|
||
| 2 | 2 | `storage_location` | `owner_identity_id`, `storage_location_id` | also label-printed |
|
||
| 2 | 1 | `file` | `file_id` | least often shared standalone; global, deduplicated by content hash |
|
||
| 3 | 2 | `workflow` | `owner_identity_id`, `workflow_id` | shared in-app, not printed - needs the escape range, since every other slot in the direct range (0-2) is already double-booked across the two arities in use |
|
||
| 4 | 2 | `group_storage_location` | `owner_group_id`, `storage_location_id` | same use case as `storage_location`, but for a group-owned location (see `group_item` above for the same owner/owner_group split); the direct range is fully double-booked, so this is the second kind that needs the escape range |
|
||
|
||
A short id is inherently scoped to the backend that minted it (an "owner" field is a row that only
|
||
exists in, and only means anything to, that one backend's database), not a portable replacement for
|
||
a `user@domain.tld` handle, which stays the form to use anywhere cross-domain resolution actually
|
||
matters. Resolving a short id still goes through the same friend/signature checks as everything
|
||
else, unchanged; nothing about how the code looks grants any authority of its own (see Guessability
|
||
below).
|
||
|
||
### Packing one integer: dynamic bit depth
|
||
|
||
Each integer field is made self-delimiting with **continuation chunking** (UTF-8/LEB128-style):
|
||
split the value into fixed-size chunks (4 data bits each, most-significant chunk first), each
|
||
preceded by one continuation bit meaning "another chunk follows" (`1`) or "this is the last chunk"
|
||
(`0`). A value like `42` (`0b101010`) needs two 4-bit chunks, costing 10 bits total (2 × (1
|
||
continuation + 4 data)); `7` fits in one chunk, costing 5 bits. This was chosen over an
|
||
Elias-gamma-style unary/delimiter scheme (encode the value's bit-length in unary, then that many
|
||
literal bits): unary is cheaper for single-digit values but its prefix grows every time the value's
|
||
bit-length grows, so it never wins once ids pass single digits, which is the common case here (auto
|
||
increment database ids realistically sitting in the tens through low-hundred-thousands over an
|
||
installation's life). A 4-bit chunk width is a reasonable fixed default across that whole range;
|
||
per-field tuning was checked against both a uniform and a skewed (geometric) distribution and never
|
||
won by more than a fraction of a character, not enough to justify a tuning knob.
|
||
|
||
### From bits to text: base64 without the byte layover
|
||
|
||
Standard base64 assumes byte-aligned (8-bit) input, grouping 3 bytes into 4 output characters and
|
||
padding to a byte boundary before encoding. Since there's no byte layer here to begin with, the
|
||
bit-packed stream is instead packed directly into 6-bit groups and mapped straight onto the
|
||
URL-safe base64 alphabet (RFC 4648 §5: `-` and `_` in place of `+` and `/`), with the final
|
||
character's unused low bits padded with zeros. That padding is safe by construction: it's always
|
||
0-5 zero bits (just enough to reach the next multiple of 6), and the decoder stops reading fields
|
||
the moment a chunk decodes to 0 with nothing left after it - by the last-field-nonzero rule above,
|
||
that can only be the padding, never a genuine field, so it's discarded rather than counted. No `=`
|
||
padding characters are needed either; those exist in classic base64 purely to communicate
|
||
trailing-byte padding, and there is no byte layer here to need that.
|
||
|
||
### The leading `~`
|
||
|
||
Every token is prefixed with a literal `~`, so a short id in a URL looks like `~DyU`. Its only job
|
||
is to mark "everything after me decodes as one of these": URL-safe base64 never produces a `~`
|
||
itself, so the prefix can never be confused with the payload, and none of Toolshed's other
|
||
path-segment formats (bare usernames, `user@domain` handles, slugs, plain numeric ids) start with
|
||
`~` either. `~` is one of RFC 3986's `unreserved` characters (§2.3, the same class as letters,
|
||
digits, `-`, `.`, and `_`), a stronger guarantee than merely being legal in a path segment: it's
|
||
never a target for percent-encoding and never carries special meaning in any URI component, so a
|
||
short id can be handed to any part of the stack without first checking which encoding rules apply
|
||
there.
|
||
|
||
|
||
### Domain-Qualified Short ID
|
||
|
||
A short id token is deliberately opaque and scoped to whichever backend minted it: handed a bare
|
||
`~DyU`, nothing in the token itself says which backend's numbering it should be read against.
|
||
That's fine as long as the token stays inside a context that already knows the answer (the app the
|
||
user is currently looking at), but not once a token needs to travel outside that context, e.g.
|
||
pasted into a message to a friend on a different domain, or logged somewhere not tied to one
|
||
backend.
|
||
|
||
The fix is the same one every other cross-domain handle in this project already uses: pair the
|
||
opaque part with the domain that's authoritative for it. A **Domain-Qualified Short ID** is a short
|
||
id token prefixed with a domain and a literal `:`, e.g. `toolsheddomain.tld:~DyU`. The domain half
|
||
is exactly a handle's domain half (see federation.md's Unique Handles section): not a location,
|
||
just a statement of which backend to resolve the token against. Decoding still means handing the
|
||
`~token` half to that backend's decoder, same as ever; it's just no longer ambiguous which
|
||
backend's decoder to hand it to.
|
||
|
||
This is deliberately not the same thing as a User-Qualified ID (`owner-handle:i42`) or an Item URL.
|
||
It isn't a URL and isn't meant to be openable by something that doesn't already know what a
|
||
Toolshed short id is: there's no scheme, no path, and the payload after `:~` is bit-packed base64,
|
||
opaque to a human. It belongs to the same "context already makes clear it's Toolshed data" class as
|
||
a compact User-Qualified ID, meant for use inside the app (or between things that already speak the
|
||
short-id format, including a physical label meant for the app's own scanner) - not a link or code
|
||
meant to be opened by something with no idea what a Toolshed short id is. What it adds over a bare
|
||
`~token` is that it no longer depends on "whichever backend I currently happen to be talking to" —
|
||
the domain travels with it, so it keeps resolving to the same entity once copied elsewhere. Unlike
|
||
a User-Qualified ID, it isn't limited to owned kinds: every kind in the short id registry can be
|
||
domain-qualified the same way, since the domain only ever names which backend's token namespace
|
||
applies, not which kind the token decodes to or whether that kind has an owner.
|
||
|
||
### Worked examples
|
||
|
||
Encoding `kind = item` (0), `owner_identity_id = 7`, `item_local_id = 42`:
|
||
|
||
- `kind`: 2 fixed bits → `00`
|
||
- `owner_identity_id = 7`: fits in one 4-bit chunk → 5 bits (`00111`)
|
||
- `item_local_id = 42`: needs two 4-bit chunks → 10 bits (`1001001010`)
|
||
|
||
Total: 17 meaningful bits, padded to the next multiple of 6 (18) with one zero bit, yielding 3
|
||
base64 characters: **`~DyU`**.
|
||
|
||
The same worked-out form for one example of every registered kind - `Bits` is the same
|
||
space-separated segmentation: kind tag, escape offset if present, each field, then padding. Kinds
|
||
are grouped by shared id below to make the per-arity reuse visible:
|
||
|
||
| Kind | Fields | Serialized | Bits | Token |
|
||
|---|---|---|---|---|
|
||
| `item` | `owner_identity_id: 7`, `item_local_id: 42` | `[0, 7, 42]` | `00 00111 1001001010 0` | `~DyU` |
|
||
| `category` | `category_id: 5` | `[0, 5]` | `00 00101 00000` | `~Cg` |
|
||
| `group_item` | `owner_group_id: 5`, `item_local_id: 42` | `[1, 5, 42]` | `01 00101 1001001010 0` | `~SyU` |
|
||
| `group` | `group_id: 11` | `[1, 11]` | `01 01011 00000` | `~Vg` |
|
||
| `storage_location` | `owner_identity_id: 3`, `storage_location_id: 1000` | `[2, 3, 1000]` | `10 00011 100111111001000 00` | `~hz8g` |
|
||
| `file` | `file_id: 123` | `[2, 123]` | `10 1011101011` | `~rr` |
|
||
| `workflow` | `owner_identity_id: 2`, `workflow_id: 9` | `[3, 2, 9]` | `11 00000 00010 01001 0` | `~wiS` |
|
||
| `group_storage_location` | `owner_group_id: 5`, `storage_location_id: 1000` | `[4, 5, 1000]` | `11 00001 00101 100111111001000 000` | `~wln5A` |
|
||
|
||
`workflow`, `group`, and `file` are kinds 3-5, so their `Bits` column shows the escape tag (`11`)
|
||
followed by its own offset segment payload fields. |