Inter-Rater Reliability of Automated Modulign Standard Classifiers
Abstract
This paper reports the first empirical inter-rater reliability study conducted on automated classifiers
implementing the Modulign Standard v3.0 — a Dimensional Address Grammar for Observable
Reality (DAG-OR). Two independently designed heuristic classifiers were applied to a corpus of N
= 234 live observation streams registered in the canonical Modulign Observation Registry.
Classifier 1 (ZengineCamBot) employs a sequential keyword scan with first-match resolution;
Classifier 2 (ModulignReasonBot) employs a weighted token accumulation algorithm with domain
vote aggregation and URL structural analysis as a separate evidence channel. The two algorithms
are architecturally independent: same input, different classification logic, independent output.
Overall domain-level agreement was 44.4% (n = 104/234). Agreement was strongly domaindependent: transport (TRN) achieved 100% agreement (n = 8/8), while natural environment (NAT)
and urban (URB) domains each achieved approximately 42–43% agreement, with systematic
disagreement concentrated at the NAT/ENV boundary. The dominant disagreement pattern — 43
NAT→ENV and 40 URB→ENV mismatches — identifies a structural boundary ambiguity in the
current domain taxonomy rather than random classification error. These results constitute the first
empirical validation data for any Modulign Standard implementation, establish a reproducibility
baseline for the Classification Decision Protocol (CDP), and identify a specific amendment target
for v3.1 of the Standard.
Cite it
Gonzalez, V. (2026). Inter-Rater Reliability of Automated Modulign Standard Classifiers. Zenodo. https://doi.org/10.5281/zenodo.19643322