schema-join
Inputs¶
schema-joinMUST read one input file and one join file, reprojecting both to EPSG:4326 the same way every other tool does, viacore.assign.load_input()/load_overlay().schema-joinMAY take the samename_field/code_fieldpairschema-maptakes (each containing a{n}placeholder); both MUST be given together, or both omitted. When given, the join layer's hierarchy columns MUST be every column in a{n}-numbered family undername_field's orcode_field's prefix, for every leveldetect_levels()finds on the join layer, raisingValueErrorunder the same missing-level rules asschema-fill.- When
name_field/code_fieldare omitted,schema-joinMUST instead structurally auto-detect the join layer's hierarchy columns (every level's identity columns fromcore.schema_map's cardinality/containment matcher, no naming convention assumed), raisingValueErrorif no level is detected. - Only the join layer's hierarchy columns are ever copied; any other join layer column MUST be ignored.
Assignment¶
schema-joinMUST assign each input feature to the single join feature it shares the most area with (core.assign.assign_many(), per-feature plurality, ties broken by lowest join fid), measured inEQUAL_AREA_CRS.- An input feature overlapping no join feature MUST stay in the output, with every copied join column NULL.
Joining¶
For each join layer hierarchy column:
- absent from the input layer:
schema-joinMUST add it, filled from the input feature's assigned join feature; - present on the input layer and equal (
IS NOT DISTINCT FROM) on every assigned input feature:schema-joinMUST skip it, leaving the input layer's column as-is; - present on the input layer and different on any assigned input feature:
schema-joinMUST leave the input layer's column untouched and add the join feature's values under the next free numbered sibling name (adm2_name1, thenadm2_name2ifadm2_name1is taken on either layer), logging a warning with the differing row count.
schema-join MUST NOT raise on a conflicting value, and MUST NOT
overwrite any input value (see docs/adr/0109).
Outputs¶
schema-joinMUST NOT modify geometry, and so performs no topology hard gate at all.- The output MUST keep every input row, using
name_field/code_field, oradm{n}_name/adm{n}_codewhen omitted. Columns MUST keep input order, each added numbered sibling right after the last existing column of its family, and every column absent from the input layer after all input columns, in template order (seedocs/adr/0119). Rows MUST be sorted by the deepest level's own code column, as inschema-map. schema-joinMUST write an issues file in the shared issues-table column schema, with one row per:no-overlap: an input feature overlapping no join feature;low-overlap: an input feature whose assigned join feature covers less thanmin_overlapof its own area, witharea_m2set to the input feature's area outside that join feature andreasonstating the covered share;value-mismatch: an input feature and column where the input feature's value and its join feature's value are both non-NULL and differ, withreasonnaming the column and both values.unit_aMUST hold the input feature's 1-based row number in the output file, not its input fid, since rows are re-sorted by code.schema-joinMUST NOT write an empty issues file, and MUST remove a stale one at the issues path instead.
Configuration (api.schema_join.join() / CLI)¶
schema-joinMUST process exactly one input file and one join file per call; either MAY be anhttp:///https://URL to a.parquetfile.- The output path MUST default to the input path with a
_joinstem suffix, and the issues path to the output path with an_issuesstem suffix (issues_path/--issues-outputto override). schema-joinMUST raiseFileExistsErrorif either the output or the issues path already exists and overwriting wasn't requested.min_overlap/--min-overlapMUST default to0.5and MUST raiseValueErroroutside(0, 1].step, if given, MUST be one ofinputs,assign,join,outputs; any other value MUST raiseValueError.
Examples¶
Example 1: basic run, structural auto-detection, output name chosen automatically¶
topo-tools schema-join admin3.parquet admin2.parquet
Example 2: chain levels coarsest-first¶
topo-tools schema-join admin2.parquet admin1.parquet admin2_join.parquet
topo-tools schema-join admin3.parquet admin2_join.parquet admin3_join.parquet
Example 3: custom target naming¶
topo-tools schema-join admin3.parquet admin2.parquet --name-field adm{n}_name --code-field adm{n}_pcode
Example 4: explicit issues path and a stricter overlap threshold¶
topo-tools schema-join admin3.gpkg admin2.gpkg admin3_join.gpkg --issues-output review.gpkg --min-overlap 0.9