파운데이션 모델 엔지니어링

10.1 Human Feedback Loop

Human feedback은 학습 데이터이기 전에 측정 시스템입니다. Product feedback은 사용자가 무엇을 보았고, 어떤 신호를 선택했으며, 어떤 interface와 policy 아래 있었는지를 기록합니다. 노출과 sampling context가 없으면 thumbs-up은 model quality의 unbiased label이 아닙니다.


Offline과 Online Feedback

Offline 프로그램은 통제된 rubric 아래 demonstration, ranking, critique, tie, abstention을 수집합니다. 균형 잡힌 sampling과 반복 annotation이 가능합니다. Online 프로그램은 실제 배포 failure를 찾지만 exposure, position, engagement, survivorship, selection bias가 섞입니다.

Candidate set, 표시 순서, serving policy, model/artifact revision, decoding setting, 사용자가 본 latency/error, sampling 또는 propensity 확률을 기록합니다. 인과적인 product 결정을 할 때는 무작위 A/B(randomized A/B) 또는 interleaving experiment를 사용합니다. 더 자주 노출되거나 유리한 위치에 있던 답이 click을 더 받았다고 본질적으로 더 좋다고 결론 내리면 안 됩니다.

Training Buffer: 0 trajectories
Live Trajectory Evaluation (Threshold: Score ≥ 8.0 & Verified)
Click start to begin asynchronous rollouts...

Admission score는 routing 도구이지 high-reward trajectory만 학습해도 된다는 허가가 아닙니다. 성공하거나 점수가 높은 출력만 선택하면 support가 좁아지고 선택 편향(selection bias) 이 생깁니다. 통제된 negative, counterfactual comparison, exploration data를 보존합니다.


Annotation 계약

과업별 rubric에 correct, partially correct, unsafe, over-refusal, unverifiable, out-of-scope 예시를 둡니다. Response가 동등하면 동점(tie), 증거나 전문성이 부족하면 판단 보류(abstain) 를 허용합니다. 모든 item을 binary preference로 강제하면 noise를 만듭니다.

Human gold set으로 annotator를 calibration하고 domain·difficulty별 평가자 간 일치(inter-annotator agreement) 를 측정하며 critical disagreement를 adjudication합니다. Rubric version, randomized position, annotator qualification, timestamp, confidence, reason code를 저장합니다. 동의한 정책 범위를 넘는 sensitive attribute를 사용하지 않으면서 demographic, language, accessibility, domain slice의 체계적인 차이를 모니터링합니다.

Privacy도 data contract입니다. Consent, retention, access, deletion, reuse scope를 정하고 annotation이나 학습 전에 PII, secret, credential, third-party content를 탐지·redaction합니다. 승인된 transformation이 versioned manifest를 만들기 전까지 raw production log는 training corpus 밖에 둡니다.


Feedback Loop를 만들지 않는 Active Learning

Label이 결정을 바꾸는 곳에 review budget을 씁니다. Policy uncertainty, judge disagreement, high-impact domain, novel cluster, safety-critical action, distribution drift가 대상입니다. 데이터가 hard case만으로 채워지지 않도록 ordinary traffic의 probability sample도 포함합니다.

Sampling probability를 기록하고 population metric을 추정할 때 stratification 또는 propensity-aware analysis를 사용합니다. Active-learning data는 학습에는 훌륭해도 weight 없는 product-quality estimate에는 무효일 수 있습니다.

검증 가능한 과업에는 human preference와 unit test, tool-state check, citation, environment outcome 같은 실행 증거를 결합합니다. 설명이 유창하다는 이유로 reward model이 실패한 deterministic verifier를 뒤집으면 안 됩니다.


Leakage 없는 Badcase Loop

Production failure는 바로 training example이 아니라 먼저 재현 가능한 evaluation case가 되어야 합니다.

  1. prompt, context, tool state, expected behavior, artifact revision, failure taxonomy를 고정합니다.
  2. semantic cluster를 배정하고 private evaluation item을 만듭니다.
  3. evaluation-only로 남을 별도의 형제 holdout(sibling holdout) variant를 만듭니다.
  4. 그 뒤에만 remediated example 또는 다른 sibling을 versioned training candidate에 넣습니다.
  5. 원본, sibling holdout, capability-retention, safety suite를 다시 실행합니다.

이 순서가 exact incident 암기로 “통과”하는 일을 막습니다. Untouched private release set을 유지하고 정답을 training system에 넣지 말고 접근 권한을 순환합니다.


Admission, Training, Release Gate

Data admission과 model release threshold를 분리합니다. 데이터에서는 source yield, tie/abstention, agreement, judge disagreement, PII rejection, slice coverage를 기록합니다. 학습에서는 manifest, split cluster, reference/baseline artifact, tokenizer/template, objective, seed를 고정합니다. Checkpoint와 함께 optimizer/scheduler/scaler/RNG/sampler/data cursor를 저장합니다.

Release에는 fixed decoding과 uncertainty가 있는 paired comparison을 사용합니다. Target quality, base retention, safety/over-refusal, format/tool correctness, latency, cost, error rate를 gate로 두며 critical slice failure를 평균으로 숨기면 안 됩니다. Shadow, canary, ramp-up 단계에는 자동 abort window와 연습한 last-known-good rollback이 필요합니다.


Quizzes

Quiz 1: Production thumbs-up이 unbiased quality label이 아닌 이유는 무엇인가요? 사용자는 특정 position, latency, interface, policy 아래 선택된 model output을 보았고 feedback을 남길지 선택했습니다. Exposure, engagement, selection 과정이 관찰 label에 영향을 줍니다.

Quiz 2: Tie와 abstain을 허용해야 하는 이유는 무엇인가요? 동등한 답에는 정직한 방향이 없고 일부 비교는 annotator의 증거나 전문성을 넘습니다. Winner를 강제하면 불확실성이 체계적인 label noise가 됩니다.

Quiz 3: Badcase 학습 전에 sibling holdout을 만드는 이유는 무엇인가요? 정확한 incident는 암기할 수 있습니다. Evaluation-only sibling은 하나의 기대 답을 재현한 것이 아니라 failure cluster에 일반화했는지 검사합니다.

Quiz 4: Active-learning sample과 함께 random traffic을 보존해야 하는 이유는 무엇인가요? Uncertainty sampling은 어렵고 특이한 사례를 과대표집합니다. Probability sample은 unbiased population estimate를 가능하게 하고 ordinary traffic에도 개선이 전이되는지 보여 줍니다.


References

  1. Ouyang, L., et al. (2022). Training language models to follow instructions with human feedback. arXiv:2203.02155.
  2. Bai, Y., et al. (2022). Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv:2204.05862.
  3. Casper, S., et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. TMLR.