Three days of talks and workshops for engineers building with LLMs. San Francisco, October 12-14, 2026.
A survey of what actually shipped in AI engineering over the last twelve months, and what it means for the year ahead.
Retrieval metrics look great in eval and fall apart in production. We instrumented ours end to end and found three failure modes nobody writes about.
Hands on workshop. Build a hybrid dense and sparse retrieval pipeline from an empty repo and measure why it beats either one alone.
We cut inference spend by 71 percent over two quarters. Most of it was not model selection.
Every agent framework hand-waves the execution boundary. Here is what actually holds up when the model decides to run rm -rf.
Your prompt changed and something broke three weeks later. A practical eval harness you can add to CI this afternoon.
Powered by Callboarddata 519 ms