Run selected prompts on a schema, then score/rescore against their rubrics with a judge model. Ephemeral โ nothing is browsed here later.
Stacktrace: