Trendline
Back

Applied AI Engineer

Optimus answers questions about live prediction markets. Doing that well is a data-systems problem before it is a model problem. One answer can draw on order books from four venues, trade history going back years, sports results and news, each arriving on its own clock and with its own definition of what a market even resolves on. You own that layer. Every answer is checkable against a settled price or a final score, so there is nowhere for a plausible-sounding mistake to hide.

What you'll do

  • Own the path from raw venue feeds to something an agent can reason over: four order books on four clocks, with sequence gaps, one-sided quotes, and venues that price the same question differently because they settle it differently.
  • Make retrieval precise at the market level. A question about one strike must not be answered with its more heavily traded neighbour, and that class of failure arrives looking like a perfectly good answer.
  • Set the abstention policy. With a stale feed or venues disagreeing past tolerance, the correct output is frequently no number at all, and the system has to distinguish a quiet market from a broken pipe.
  • Keep answers true while they stream. Prices move during generation, so a figure that was accurate at the first token can be stale by the last one.
  • Run the loop that tells us whether a change actually helped. Offline scores and production quality come apart quickly here, and you are the person who notices.
  • Hold latency and cost budgets under live traffic, on a surface people trade against.

What we're looking for

  • Production Python on a streaming path. You have kept something running while data arrived late, out of order, or not at all.
  • A real opinion about partial failure. One source of four goes stale: you know what the user should see, and it is not a spinner.
  • Enough statistics to separate an effect from an artefact in autocorrelated, non-stationary data, and to say “not significant” when that is the finding.
  • Experience evaluating a non-deterministic system, including why a fixed test set stops predicting quality once real traffic hits it.
  • Something you built that survived contact with real users. That counts for more here than where you studied.

  • What do you understand deeply that few people do?
  • What's the most ambitious thing you've ever attempted?