code
core.code is the shared primitive behind code-create and
code-update (docs/pages/4-codes/explanation/code_create.md,
docs/pages/4-codes/explanation/code_update.md): a neutral leaf, like core.assign or
core.dissolve, with no api.*()/CLI of its own. It carries every piece
of format/cascade/rewrite logic a hierarchical code needs, generic to any
organization's own convention, not hardcoded to COD-AB's (see
docs/adr/0101).
CodeFormat¶
@dataclass(frozen=True)
class CodeFormat:
root_code: str
delimiter: str
min_width: int | tuple[int, ...] | Literal["auto"]
No field has a default value. resolve_code_format(root_code, delimiter,
min_width) is a thin validating constructor, not a defaults-filling one:
non-empty root, single-character delimiter, and min_width parsed by
parse_min_width() from one positive width (3), one per level, coarsest
first ("2,2,4"), or "auto". fmt.width(level) returns a level's width
(None under auto); fmt.check_level_count(count) raises when a
per-level list doesn't have one width per numbered level (see
docs/adr/0123).
root_code is opaque everywhere, never shape-checked, so a disputed-
territory prefix (XKO, or any user-assigned string) works identically to
an ISO3 one.
Cascade: ranking, not reformatting¶
assign_new_codes(conn, table, *, id_column, parent_column, sort_columns,
code_column, fmt, level, existing_codes=None) is the one function both tools use
to actually mint new codes. It never reformats a raw source value in
place: rows are ranked per parent_column group by sort_columns
(ROW_NUMBER() OVER (PARTITION BY parent_column ORDER BY sort_columns)),
starting from next_available_integer(existing_codes, parent_code, fmt)
for that parent, then formatted as
parent_code || delimiter || lpad(tail, width, '0'). A raw source value
is never trustworthy enough to zero-pad and reuse: it may be non-numeric,
gappy, or duplicated across siblings (a real GADM GID_1 looks like
AFG.1_1, not a clean rankable integer).
next_available_integer(existing_codes, parent_code, fmt) scans
existing_codes for anything starting with f"{parent_code}{delimiter}",
skips a tail that still contains the delimiter (a nested, deeper code
under the same textual prefix) or isn't numeric, and returns
max(found) + 1, or 1 if nothing matches. It only ever looks at the
existing_codes list it's given, never a persisted registry.
code-update seeds the next available number with every OLD code at that
level, retained or retired, so a code is never reissued within or across
consecutive releases; a code retired two or more releases back isn't
tracked (see docs/adr/0126). With an empty delimiter, parse_code()
splits a code by root_code and each level's width instead, so
next_available_integer() accepts only a tail of exactly that width.
Overflow: width grows, it never repads¶
lpad truncates an over-width string (unlike Python's zfill, which
never shortens), so assign_new_codes() widens its own target width to
GREATEST(width, LENGTH(tail)) before padding, width being the level's
own. Under auto, width is the widest tail at that level, retained
codes included, so every new code at the level shares one length. A
parent's 1000th child (at width 3, capacity 10**3 - 1 = 999) gets a
4-digit tail; every child ranked below it keeps its own already-assigned
3-digit code untouched, no whole-parent repad. Without a delimiter a
longer tail can't be split, so assign_new_codes() raises ValueError
instead when a fixed-width level overflows. code-create and
code-update each independently detect and report this condition in
their own outputs stage (own issues report / changelog overflow
outcome, see their own explanation docs), core.code itself has no
reporting concept, only the underlying width behavior.
Format detection¶
detect_code_format(conn, table, code_column) (used only by
code-update, against OLD's own already-coded finest-level column) infers
a CodeFormat straight from a sample of existing values (up to 10,000
distinct, non-null), rather than requiring a caller to state it:
delimiter: the single non-alphanumeric character common to every sampled code. Zero or more than one candidate raisesValueError.root_code: the shared first delimiter-split component. Not constant across the sample raisesValueError, since that meanscode_columnisn't actually this dataset's own root-anchored hierarchy column.min_width: the mode (most common), not the min or max, of the component widths at each level's position, one width when every level agrees, else one per level. This specifically avoids an overflow-widened tail at one parent (see above) skewing the detected width for every other, non-overflowed parent.
Any field that can't be confidently inferred raises ValueError rather
than falling back to a hardcoded literal; code-update always accepts
explicit --root-code/--delimiter/--min-width overrides.
Rewrite: cascading a parent's new prefix¶
rewrite_child_code(old_code, new_parent_code, fmt) reattaches
old_code's own final (tail) component onto new_parent_code, leaving
the tail integer and its sibling ranking completely untouched. This is
the one function that lets code-update cascade a coarser unit's new code
down through every unchanged/renamed descendant without re-ranking them
(docs/adr/0103): only the unit whose own identity actually changed gets
a fresh cascade through assign_new_codes(); everything nested under it
that didn't change keeps its own relative position, just under a new
prefix.
Pure string operations¶
parse_code(code, fmt)/build_code(components, fmt) split/join on
fmt.delimiter. parent_prefix(code, fmt) drops a code's own last
component (raises ValueError, "no parent", for a root-only, single-
component code). last_component(code, fmt) returns a code's own final,
unpadded component. None of these validate a code's shape beyond simple
splitting; a malformed code just produces a malformed result rather than
raising, except at the two explicit validation points above.