Fine-tuning & EvalsEvalsModel selectionFixed budget
Model Eval & Selection for a Fintech Agent
Benchmark Collective · Remote (US)
$9k fixed
The outcome
They want data, not vibes: which model + prompt is most accurate and cheapest for their reconciliation agent's real tasks.
The client
Benchmark Collective — Remote (US). A fine-tuning & evals build we delivered end to end.
The challenge
Run a structured bake-off across models and prompts on the client's actual task set, then recommend a config with an accuracy-vs-cost tradeoff writeup.
What we built
- Build a graded task set from real client workflows
- Benchmark models on accuracy, latency, and cost
- Deliver a clear recommendation and reproducible eval
The stack
Stack: eval frameworks, multi-provider LLM APIs, and a head for cost/latency tradeoffs.
Want something like this?
Tell us what you're building and we'll scope it — most projects start within a week.