Research / innovation / publication

Measure the behavior.
Keep the failure.

The research program follows mechanisms across product development without turning implementation progress into scientific conclusions. Negative and mixed results remain visible because they often identify the next experiment more clearly than a forced winner.

01 / CURRENT EXPERIMENT

Beyond accuracy in governed agent decisions.

The latest accepted controlled study asks whether actionable corrective feedback improves required authorized movement while preserving restraint, carried state and output quality.

192model generations
96canonical blind evaluations
24assisted human-calibration cases
56 / 56governance errors A / B
Cumulative controlled program

Three completed A/B rounds, kept separate but countable.

The completed experimental series now totals 576 model generations, 288 canonical blind machine grades, and 72 assisted H1 calibration cases across three independently frozen rounds. Those totals describe research volume; they are not pooled into one synthetic efficacy estimate.

Aggregate lift was not demonstrated.

Full-48 machine quality B−A was −0.03125 with paired bootstrap 95% CI [−0.2604, +0.1875]. The tested treatment did not establish an aggregate quality advantage or aggregate governance improvement.

But a partial remediation mechanism was localized.

Across the broader required-signal seam, 4 of 11 cases were repaired and 7 remained unresolved. Three cases also lost a blocked-use boundary without authority. The current mechanism thesis is therefore narrower: detection validity, repair-target specificity, collateral preservation and independent acceptance are separable.

02 / METHODS

Research infrastructure is part of the result.

The program uses matched boundaries, blind treatment mapping, versioned graders, explicit missing/replacement handling, paired uncertainty, human/machine discordance review and claim ceilings that survive a positive-looking result.

Blind evaluation

Freeze before grade

Treatment identity is separated from the grading surface and replacements/missing observations are recorded rather than silently dropped.

BLIND MAP / FROZEN PACKET / PROVENANCE
Statistics

Paired deltas + uncertainty

Paired bootstrap intervals and exact win/tie/loss or discordance summaries are used where the design supports them; decorative math is not the goal.

PAIRED DESIGN / BOOTSTRAP / SENSITIVITY
Human calibration

72 assisted cases across three rounds

Each completed round used a separately frozen 24-case H1 assisted-calibration body. The cumulative 72 cases remain explicitly classified as user-cosigned, multi-machine-assisted evidence—not unaided independent human validation.

3 × 24 H1 / PROVENANCE / CEILING / JUDGE CALIBRATION
Failure-aware evaluation

Benchmark failures stay in the record

Parser, runtime, scorer, evaluator and protocol failures are separated from product behavior so the instrument itself cannot silently become the result.

TEST INSTRUMENT / PRODUCT RESULT / REPAIR
03 / PAPER PROGRAM

Current manuscript map.

These are research work products and manuscript states, not claims of publication or peer review.

PAPER / A

Beyond Accuracy in Governed Agent Decisions

Required movement, restraint, carry, and the boundary between a valid deterministic signal and an authority-bounded repair.

Active manuscript spine
PAPER / B

When the Benchmark Fails Before the Model

Evaluation infrastructure, construct validity, protocol-preserving repair, and what remains scientifically interpretable after the evaluator breaks.

Publication-ready outline
PAPER / C

Reasoning Is Not Free

Quality-per-compute, latency, token use and human correction burden. Held until exact compute evidence is captured.

Hold / evidence demand
PAPER / D

Does Decision Governance Transfer?

Cross-domain testing across logistics, physical assets, disaster risk and real-estate state without treating shared architecture as proof of generalization.

Future replication
PAPER / E

Judging the Judge

Human/machine calibration, exact-state semantic disagreement and the limits of using model judges as claim multipliers.

Watch / build
04 / PUBLIC LIBRARY

Papers, preprints & research notes.

This section is wired for simple future updates. Public PDFs can be added without redesigning the page; the core research narrative above remains static and crawlable.

Loading public research items…
05 / RESEARCH POSTS

Public notes & social posts.

External research threads, release notes, conference posts and public commentary can be linked here as they are published.

Loading public research posts…
Research contact

Replication beats a flattering summary.

Kingan Logic is open to bounded research collaboration, independent replication and evaluation relationships where the protocol, evidence and contribution can be stated clearly.