Latency and Accuracy in System One models
I’ve been a fan of Jev for a couple of weeks as a quick classifier for pre-routing agent requests. Today OpenAI released the beta of its Decisions API, so I ran both through the same classifier test for accuracy and latency.
On a dataset of ~300 prompts:
- Accuracy: OpenAI 97%, Jev 96.5%. Effectively a tie.
- Latency p50: Jev ~95 ms, OpenAI ~153 ms
- Latency p95: Jev ~165 ms, OpenAI ~380 ms
Both are plenty fast for routing. I exopect OpenAI will get faster; their API today is just serving the Luna model, unoptimized. Use either with confidence, and compare the features in-depth to match your needs (e.g. OpenAI supports image input, I am not using that so I did not factor it).
I am sure there is lots of possible variance: I tested from Seattle on a low OpenAI usage tier. Want to run it yourself? The source for test is here. I am not a perf engineer, so suggestions for improvements (or mistakes) are welcome.
Note: the dataset in the public sample is much smaller than the one I tested with, but it is representative. Plug in your own.