OptionalexperimentThe A/B lever. false = planner runs blind (baseline). true = classify the question's
DateIntent LIVE (classifyDateIntent) and feed it to the planner — exactly the shipped
prod behavior, so the arm measures the real change (classifier included), not an idealized
gold-intent injection. The example's gold intent is the judge's REFERENCE, not the input.
OptionaljudgeJudge model; pinned separately so the measuring instrument stays fixed across arms.
OptionalmaxOptionalmetadataOptionalmodelPlanner model; defaults to the chart-review default (gpt-4.1, temp 0 for reproducibility).
OptionalnumRun each example N times to measure per-example variance (a "lift" must beat the noise).
LangSmith dataset of
{ question, goldIntent, goldPlanExpectation, kind, ... }examples.