A scoped four-week pilot. You bring one workload and a memory budget, we return a measured comparison against your current cache: footprint, recall, and the configuration that produced both. No deck, no benchmark theatre, just the field on the traffic you already run.
A pilot is a scoped, time-boxed engagement, not an open-ended trial. We agree a workload and a memory budget up front, run your traffic through both a standard cache and the field, and hand back a comparison with the configuration that produced it.
The point is a decision. At the end you have measured numbers on your own workload and a clear read on whether the field belongs in what you ship.
Illustrative scope. The exact terms are finalised with each pilot partner under NDA.
The pilot measures your numbers, so it starts from your workload. What you bring is small and specific, so the result is real.
A single, representative slice of what you actually run: the traffic, the context lengths, and the prompts that push memory. Not a toy benchmark, the thing that hits the ceiling.
The fixed footprint you want to hold. We size the field and the window to that number, plus whichever higher tiers your workload calls for — verbatim pins, a loaded lens — so the result is measured against a ceiling you chose.
What a good result looks like for you: the recall you cannot lose, the latency you can accept, the footprint that unblocks the roadmap. We measure against it.
The total field footprint under your budget, shown against your current cache on the same workload, so the number you set is the number you get.
Recall deltas on the field's gist, and on verbatim pins where your workload uses them, measured against your current cache on the same traffic, so a held detail is a measured detail rather than a claim.
The model, hardware, context lengths, fidelity setting, and any lens that produced the numbers, handed over so you can re-run the setup on your own stack.
Where SGF fits your workload and where it does not, in writing. If the field is the wrong tool for what you run, the report says so.
SCOPE: the report carries the model, hardware, context length, fidelity setting, and any lens behind every number. Measurements that have not cleared that bar do not go in it.
We agree the workload, the memory budget, and the success bar under NDA, and confirm the data and access needed to run the comparison.
The field goes in behind your serving layer on the pilot workload. No retraining, one memory interface, sitting beside the stack you already run.
We run your traffic against both the standard cache and the field, capturing footprint, recall, and latency with the full configuration recorded.
You get the measured comparison, the configuration that produced it, and a go or no-go recommendation scoped to what you would ship.
SCOPE: four weeks is the default shape. A larger workload or a deeper integration can extend it, agreed up front rather than discovered along the way.
Not every workload needs a bounded field, and we would rather say so early. These are the conditions where a pilot tends to pay for itself. If they describe what you run, the discovery call is the next step.
The workload is limited by how much the cache holds, not by retrieval or by the model itself. Long sessions, long agent runs, or on-device budgets.
You can bring representative traffic and a memory budget. The pilot measures your numbers, so it needs your workload rather than a stand-in.
There is a serving path the field can sit behind for the engagement. Drop-in, no retraining, but it does need somewhere to run.
A pilot is worth running when a measured result would change what you ship. If the answer would not move a decision, it is not the right time.
A scoped four-week pilot: you bring one workload and a memory budget, we return a measured comparison against your current cache: footprint, recall, and the configuration that produced both.