ADR 0056 — Trajectory-rationality judge 를 새 측정 표면으로
ADR 0056 — Trajectory-rationality judge 를 새 측정 표면으로
- Status: Accepted
- Archive note (2026-06-02): 본 ADR 본문 및
## Verification블록의 real100 경로(reports/real100/rationality.*)와make real-eval예시는 historical이다. 현행 rationality judge 산출물은reports/real100_v2/rationality.*(make real-eval-v2-rationality-judge); legacy real100 은 archive-only(ADR 0005 / CLAUDE.md ban-list). 본문·Verification 명령은 측정 record 보존을 위해 변경하지 않는다. - Implemented: #987 (2026-05-18) —
eval/judges/rationality_judge.py3-axis trajectory surface. (#1297 (2026-05-22) 가 committed artifact 와의 모순을 정정해 일시적으로 2-axis measured + 1-axisanswer_reasoningpending 으로 relabel; #1326 (2026-05-22) 가 stub full-trace 캡처 wiring 을 고쳐 3-axis 모두 measured 로 복원 —answer_reasoningeffective_n=166) - Date: 2026-05-18
- Authors: Hyunsoo Kim
- Related: ADR 0006 (real-data LLM-judge), ADR 0054 (conditional-on-substantive-answer scorer semantics), ADR 0055 (claim_validator)
- Augments: Phase 3 audit (
docs/audits/eval-framework-phase3-audit.md, PR #961) item 3 supply (“trajectory-rationality rubric ✗ absent”) - Issue: #969
Context
Phase 3 audit (PR #961 c50a3e7) 의 4-item 표 중 item 3 진단 = ✗ absent — grep '(planner|retrieval|verifier).*rationale|process.*judg|trajectory.*judg|rubric' 0건. 기존 3 judge gate (Gate 1 real-data answer-quality / Gate 2 synthetic answer-quality / Gate 3 RAGAS answer-quality) 모두 answer correctness 만 채점 — process rationality (planner 가 합리적으로 decompose 했는가, retrieval 재호출이 evidence-driven 인가, synthesis 가 evidence 와 일관한가) 측정 표면 0차원.
ADR 0006 (real-data only) 가 정한 boundary 는 본 judge 에도 그대로 적용 — 본 모듈은 trace JSON 의 read-only consumer 이며 production code path 0 변경.
Step 2 (PR #968, ADR-free) 의 trace schema v2 synthesis_llm_call 키 (BIDMATE_TRACE_FULL=1 env-gated) 가 answer_reasoning axis 의 input 자료 공급.
Decision
- 신규 모듈
eval/judges/rationality_judge.py도입.- 시그니처
judge_rationality(summary, *, backend="stub", traces_dir=None, cache_dir=None, token_budget=200_000) -> (local_payload, aggregate). - 3 axis (각
[0.0, 1.0]continuous, 4-th decimal precision):planner_decomposition—trace["planner"]의query_type/pipeline/stage_sequence/selected_top_k/retrieval_budget.reasonsubset 기반.retrieval_recalls—trace["planner"]["attempts"][*]["verification_reasons"](retry 사유) 기반.answer_reasoning—trace["synthesis_llm_call"].{user_prompt_text, completion_text}기반.BIDMATE_TRACE_FULL=1미설정 시 None → aggregateeffective_n에서 drop.
- 시그니처
- 시그니처는 Gate 3 RAGAS (
eval/judges/llm_judge.py:judge_ragas) 의 패턴 그대로.(summary, *, backend, cache_dir, token_budget) -> (local, aggregate)— 4 judge surface 모두 같은 contract.- Backends:
stub(default, deterministic SHA-256) +openai_compatible(judge_common.build_openai_client재사용, 동일 env contract).
- 신규 CLI
scripts/run_rationality_judge.py.--eval-summary/--output/--out-aggregate/--out-md/--backend/--traces-dir/--cache-dir/--token-budget.eval/judges/llm_judge.py의 CLI 패턴과 1:1.
- 신규 committable artifacts:
reports/real100/rationality.aggregate.json— per-axis mean + 95 % bootstrap CI + effective_n.reports/real100/rationality.md— axis 표 + bottom-3 case per axis (rationale review).- 둘 다
.gitignoreallowlist 에 등재 (기존eda.md/distinguishing_power.mdsibling).
answer_reasoningNone-skip semantics — ADR 0054 의 substantive-only 의미를 trajectory 측정에도 적용.synthesis_llm_call부재 케이스 (env=off 또는 stub answer backend) 는answer_reasoning = None.- aggregate
effective_n["answer_reasoning"]가 실제 측정 가능했던 케이스 수 보고. - mean 분모 제외 — sample 부재 axis 는
mean = None,ci미발행.
- 측정은 stub backend + n=221, 3-axis 모두 measured. 비용 0, deterministic. committed artifact (
reports/real100/rationality.aggregate.json) 는cases_with_synthesis_llm_call=166/answer_reasoningeffective_n=166 (나머지 55 케이스는 abstention 으로 synthesis 미수행 → 정상 drop). 이를 위한 두 전제: ①answer_reasoning측정은 synthesis 가 도는 run 한정 — eval_summary 의case_results는 primary run 것만 담으므로(run_eval.py), real config 의primary_run이agentic_full_llm(prompt_profile=llm_synthesis) arm 을 가리켜야 한다(또는 judge 를--traces-dir .../full_llm로 override). extractivefullprimary 로는 synthesis trace 가 case_results 에 없어 effective_n=0. ② stub synthesis backend 가BIDMATE_TRACE_FULL=1에서 full I/O 를 캡처 — #1326 이전엔_stub_backend가user_prompt_text/completion_text를 반환하지 않아 effective_n=0 였다(#1297 이 pending 으로 정직하게 표시했던 상태). #1326 이 stub 캡처를 live backend 와 동일하게 고쳐 ② 를 해소. LLM backend 실측정(BIDMATE_RATIONALITY_BACKEND=openai_compatible)은 별 PR (token budget + cost analysis 동반).--expect-full-trace(#1297) 가 effective_n=0 인 full-trace 기대 run 을 incomplete 로 표면화.
Why these specific choices
| 결정 | 근거 |
|---|---|
| Verifier axis 제외 (원래 sketch 의 3-axis 중 하나) | Step 2 (#968) 작성 중 발견 — rag_verifier.py 가 rule-based (LLM call 0건). LLM judge 의 ROI 가 약함 (sufficiency rule 의 재검증일 뿐). answer_reasoning 으로 대체 — 실 LLM call site (synthesis) 의 합리성 측정. |
| 1 LLM call per case (3-axis 합본) vs axis 별 3 call | RAGAS (Gate 3) 와 동일 — 3 LLM call 분할은 비용 3× / 일관성 위협. 합본 prompt 는 모든 axis 가 같은 trace context 를 봄. |
traces_dir parameter 도입 |
per-case trace JSON 이 case.trace_path 의 별도 파일 (eval_summary 에 embed 안 됨). traces directory 위치 override 옵션 — CI 환경에서 base/pr artifact 경로 다를 때 대응. |
| stub backend = SHA-256(trace subset, axis, case_id) | byte-identical cross-platform (Gate 3 stub 의 constant scores 보다 strong — case 별 차이 보존 → 분포 aggregate 가 stub 만으로도 의미 있음). 다른 backend 와 변별력 비교용 floor. |
cache_dir 파라미터 reserve |
현재 stub 은 cache 불필요 (free + deterministic). LLM backend cache 는 follow-up PR — Gate 3 와 contract 호환 유지. |
| Markdown bottom-3 per axis | dashboard 없이도 rationale review 가능. rag_pipeline.md / distinguishing_power.md 의 1-screen surface convention 준수. |
Consequences
- Phase 3 audit item 3 (✗ absent → ✓ present) 폐쇄. 5-step portfolio narrative (“측정 → 함정 발견 → 함정 fix → 측정 표면 audit → 자동 게이트 도입 → process rationality 측정 도입”) 의 step 3 (= 측정 표면 완비) 까지 달성.
- 신규
reports/real100/rationality.{md,aggregate.json}두 산출물 → eval surface 의 1-차원 추가. - judge LLM backend 의 실 측정 (n=221 × 1 LLM call) 은 별 PR (#1377, 2026-05-23) 에서 수행 — Claude Sonnet 4.6, ≈$1.5. stub↔LLM Spearman ρ≈0 (모든 축) 으로 stub 이 품질 proxy 가 아닌 분포 floor 임을 확증. 상세는 “Measurement appendix — LLM backend”.
answer_reasoning의 effective_n 는BIDMATE_TRACE_FULL=1+ synthesis trace 캡처 여부에 의존. #1326 측정에서 effective_n=166 (cases_with_synthesis_llm_call=166) — synthesis-primary run (agentic_full_llm) + stub full-trace 캡처 wiring fix 로 3-axis 모두 measured. 나머지 55 케이스는 abstention 으로 synthesis 미수행(정상 drop, ADR 0054 substantive-only semantics). (이력: #987 최초 measured 주장 → #1297 이 effective_n=0 artifact 와의 모순을 pending 으로 정직 정정 → #1326 이 wiring 을 고쳐 measured 복원. 이전 본문의 “모든 case cover” 표현은 abstention drop 을 누락했던 부정확 — “synthesis 수행 case 전부 cover” 가 정확.)- 향후 PR 에서
Claim:(ADR 0055) 으로 rationality axis 의 변화 보고 가능 —Claim: planner_decomposition=+0.05pp식. 단 본 PR 에서는 baseline 측정만, claim 0건.
Invariance check
- ADR 0001 (
naive_baselinebyte-identical) — 본 judge 는 read-only consumer, production code path 0 변경 → 합성 baseline 영향 없음. - ADR 0003 (answer dict schema_version=2) — 변경 없음.
- ADR 0005 (private real / public fixture smoke 분리) —
reports/real100/rationality.*는 ADR 0005 의 aggregate-only allowlist 패턴 그대로 (eda / distinguishing_power 와 동일). per-case (case id 포함) 는rationality.local.jsongitignored.rationality.md의 bottom-3 행은 발주기관명-인코딩 qid 대신 익명 rank (#1/#2/#3) + slice + score 만 노출 (#1297 sanitize; 이전엔 raw qid 가 committed 되어 본 주장과 모순이었음 —real-data-failure-taxonomy.md의 P-NN 컨벤션과 일관화). - ADR 0006 (LLM-judge real-data only) — rationality_judge 도 real eval surface 에서만 의미 (synthetic 의 trajectory 는 deterministic). 본 PR 의 첫 측정은 real n=221, ADR 0006 boundary 준수.
- ADR 0054 (substantive-only scorer semantics) —
answer_reasoning의 None-skip 이 같은 의미를 trajectory 측정 layer 에 propagate. - ADR 0055 (claim_validator) — rationality axis 는 향후
Claim:검증 대상이 될 수 있음. paired_bootstrap_ci 가 None pair drop 하므로 호환.
Out-of-scope
LLM backend 실측정 (완료 (#1377, 2026-05-23) — 아래 “Measurement appendix — LLM backend” 참조.BIDMATE_RATIONALITY_BACKEND=openai_compatible, n=221 × 1 call = ~$1 추정) — 별 PR.- Verifier-axis rationality (Step 2 audit finding 으로 폐기 —
rag_verifier.pyLLM call 0). - Planner full I/O trace dump (Step 2 surgical scope 외).
- Per-axis weighting / composite “process_health” score — 본 PR 은 raw axis 만, 합성 score 는 portfolio narrative 가 정해진 후 별 ADR.
Verification
# Stub backend determinism + 6-case unit test
EMBEDDING_BACKEND=hashing python3 -m pytest -q tests/test_rationality_judge.py
# Real-eval regen at n=221 with BIDMATE_TRACE_FULL=1.
# answer_reasoning 측정은 synthesis 가 도는 run 한정 — local config 의
# primary_run 이 `agentic_full_llm`(prompt_profile=llm_synthesis) arm 을
# 가리켜야 case_results 에 synthesis trace 가 담긴다. extractive `full`
# primary 로는 effective_n["answer_reasoning"]=0 (synthesis 미수행).
BIDMATE_TRACE_FULL=1 BIDMATE_SYNTHESIS_BACKEND=stub make real-eval
# Score traces (stub backend, 0 cost). --expect-full-trace 로 answer_reasoning
# effective_n==0 (silently-incomplete) 인 run 을 exit 3 으로 표면화.
python3 scripts/run_rationality_judge.py \
--eval-summary reports/real100/eval_summary.json \
--output reports/real100/rationality.local.json \
--out-aggregate reports/real100/rationality.aggregate.json \
--out-md reports/real100/rationality.md \
--backend stub --expect-full-trace
# Aggregate inspection
python3 -c "import json; r=json.load(open('reports/real100/rationality.aggregate.json')); \
print('n:', r['n']); print('effective_n:', r['effective_n']); \
print('means:', {k: round(v, 3) if v else None for k, v in r['axis_means'].items()})"
Measurement appendix — LLM backend (Sonnet 4.6)
2026-05-23 — stub vs LLM 변별력 (비공개 real-100, n=221)
본 ADR Decision 6 의 첫 측정은 stub backend (deterministic SHA-256) 였고, Out-of-scope 가 LLM backend 실측정을 별 PR 로 예약했다. PR #1377 (issue #1377) 이 그 deferred 분을 실측 — 같은 synthesis-primary trace 셋을 stub 과 openai_compatible (Claude Sonnet 4.6) 두 backend 로 채점하여 변별력을 비교한다.
- Surface: 비공개 real-100, n=221 (synthesis 수행 154 → answer_reasoning 측정 가능). prebuilt
data/index/real100(hashing) +agentic_full_llmprimary (prompt_profile=llm_synthesis) +BIDMATE_TRACE_FULL=1 BIDMATE_SYNTHESIS_BACKEND=stub(synthesis 자체는 stub — judge 변별력 측정이 목표, synthesis 품질 측정 아님). - Backends: stub = SHA-256(trace subset) 분포 floor. LLM =
claude-sonnet-4-6(Anthropic OpenAI-compat, temp=0). 비용 ≈$1.5 (judge len//3 input estimate 251k tokens, output ~27k). - Enabler:
eval/judges/judge_common.call_openai_json가response_format={"type":"json_object"}를 하드코딩해 Anthropic compat 엔드포인트가 400 (Input should be 'json_schema') 으로 거부 →BIDMATE_JUDGE_RESPONSE_FORMAT=noneenv (기본json_object, byte-identical) 도입으로 kwarg 생략 가능.get_judge_temperature(gpt-5.x temp=0 거부) 와 동일한 endpoint-quirk 우회 패턴.
| axis | stub mean (std) | LLM mean (std) | LLM 95% CI | Spearman ρ (p) | effective_n |
|---|---|---|---|---|---|
planner_decomposition |
0.493 (0.288) | 0.507 (0.194) | (0.482, 0.533) | −0.090 (0.185) | 221 |
retrieval_recalls |
0.515 (0.285) | 0.468 (0.240) | (0.437, 0.499) | −0.033 (0.623) | 221 |
answer_reasoning |
0.505 (0.290) | 0.318 (0.163) | (0.292, 0.344) | −0.012 (0.883) | 151 |
판정:
-
stub 은 품질 proxy 가 아니다 (의도된 설계 확증). 세 축 모두 stub↔LLM Spearman ρ≈0 ( ρ ≤0.09, 전부 p>0.18) — stub SHA-256 점수는 LLM judge 의 trajectory 품질 판단을 전혀 추적하지 않는다. stub 은 cross-platform byte-identical 분포 floor 일 뿐이며, “변별력 비교용 floor” 라는 Decision-table 의 stub 설계 의도가 실측으로 확인됐다. - LLM backend 는 quality-correlated 변별력을 추가한다. LLM std (0.16–0.24) < stub std (0.29 = uniform[0,1] floor) — 일관된 판단으로 분포가 좁다. slice 별로
follow_up의 degenerate trajectory 를 planner 0.225 / retrieval 0.100 로 페널티 (stub 은 0.387 / 0.569 노이즈),answer_reasoning은 stub synthesis(템플릿 passthrough)를 mean 0.318 로 낮게 채점 — evidence-aware 판단의 신호. 상세 slice/query_type breakdown:docs/audits/rationality-llm-judge-comparison.md.
Caveats: single judge / single seed (judge-variance 미추정), single run. answer_reasoning 의 0.318 은 stub synthesis 완성문을 채점한 값이라 synthesis 품질 지표가 아니다 (judge 변별력 신호일 뿐). committed 산출물은 aggregate/per-slice 만 (ADR 0005 — per-case score + qid 는 rationality.*.local.json gitignored). 결과는 비결정적 (LLM) 이라 CI 게이트 baseline 아님 — 본 appendix 가 측정 record.