Apex Lab
AI Cost Assessment Two weeks, often AWS funded

Cheaper AI model, same quality. We prove it on your own data.

We review your prompts, caching and model choice, then run your real workload across the options and show you which setup holds your quality bar for less. In two weeks, not a quarter.

evaluator report · support-agent-v3 quality bar met
current model
Claude Opus 4.1
$18.40 / 1k req
recommended
Sonnet 4.5 + cache
$4.10 / 1k req
quality delta
−0.4 pts
on 1,200 golden cases
cost per 1k requests sample data
current
$18.40
prompt trim
$14.30
+ caching
$9.60
+ right-size
$4.10

Sounds familiar?

Your AI bill grew faster than your user count.

Every feature defaults to the biggest model, because that is what worked in the demo.

You want a cheaper model but nobody dares to switch.

There is no eval set, so “it feels worse” is the only test you have.

Caching and prompt length are nobody’s job.

Prompts grew by copy-paste, context is resent on every call, and the invoice shows it.

Inside the evaluator

Measured, not felt.

Every claim we make comes out of this. Click through what you’ll actually be looking at in week two: your data, your workload, your numbers.

before · 2,140 tokensresent every call
You are a helpful support assistant. You are…
<full_product_catalog> 1,180 tokens </…>
<tone_guidelines> 320 tokens </tone_guide…>
Never mention competitors. Also never mention…
{user_message}
after · 610 tokens1,530 tokens cached
Support assistant for [product]. Answer from context only.
<cache_block id="catalog" ttl="1h">
<cache_block id="tone" ttl="24h">
retrieve(top_k=3) → only matching catalog rows
{user_message}
tokens per call −71%quality on golden set unchangedeffort half a day

Four levers. We pull all of them.

Most teams touch one and stop. The savings compound when you do them in order and measure each step.

Prompts

Trim, restructure, and move static context out of the hot path.

Caching

Prompt caching, response caching, retrieval reuse. Usually the fastest win.

Model right-sizing

Route each task to the smallest model that passes your quality bar.

Evaluation

Build a golden dataset from your real inputs and outputs, so every change is measured, not felt.

How it works

~2 weeks, 2 sessions · secure, stays in your account · often AWS-funded
01
01

Collect

We capture a sample of your real inputs and outputs, or use your existing golden data, and read your current prompts and pipeline.

02
02

Evaluate

We stand up an evaluator in your own cloud account and run your workload across model and prompt variants, side by side.

03
03

Optimize

You get a ranked list of changes with cost and quality impact, plus the rewritten prompts and caching config to ship them.

The deliverable is a ranked list, not a verdict.

Each line has a measured cost delta, a measured quality delta, and an effort estimate. You decide what ships.

#
Change
Saving
Quality
Effort
01
Cache the catalog and tone blocks
$14.4k/mo
±0.0
0.5 day
02
Route classification to Amazon Nova Lite
$8.9k/mo
+0.3
2 days
03
Retrieve catalog rows instead of sending all
$5.1k/mo
−0.4
3 days
04
Draft on gpt-5.1-mini instead of Sonnet
$3.2k/mo
−5.4
not advised

Sample output. Your numbers come from your own workload.

New model out? You’ll know the same week whether it’s worth switching.

Once the evaluator is set up, it doesn’t stop. When a new model ships, it runs your golden dataset against it and tells you if a cheaper model now matches your quality bar.

today 09:12 · evaluator

A newly released model matches your quality score on 1,200 golden cases at a materially lower cost. Review the diff.

3 weeks ago · evaluator

Tested a cheaper alternative. Quality dropped 2.1 pts on classification cases. Not recommended.

What you walk away with

  • Cost breakdown by feature, model, and prompt
  • Ranked optimization list with estimated savings and quality impact
  • Rewritten prompts and caching configuration, ready to deploy
  • Evaluator running in your account, with your golden dataset
  • 30-day follow-up review

Why Apex Lab

AWS Advanced Tier Services Partner with the AI Services Competency.

50+ engineers who ship production AI systems on Bedrock and beyond.

Assessments are how most of our engagements start. We know how to make two weeks count.

AWS

Often funded through AWS partner programs

If you’re on AWS, this is usually funded through AWS’s partner programs. Ask on the call.

The questions we always get

Yes. The method is provider-agnostic. If you are on AWS, the assessment can often be funded through AWS partner programs. Ask us on the call.

Find out what your AI should actually cost.