PredMark
Can I find an edge across a hundred thousand prediction markets? And would I know if I was just getting lucky?
Architect & sole engineer
- ~101K
- Markets watched
- 0
- Trades the LLM decided
- 40+
- Co-pilot tools
- 867
- Commits
A profitable week doesn't tell you much. I wanted to separate useful signals from windfalls before convincing myself I'd found a strategy worth trading.
Too much market for one person
PredMark watches roughly 101,000 markets across Kalshi and Polymarket. It uses news, social activity, and price changes to find candidates for closer investigation.
Claude scores candidates against cited evidence, writes reports, and answers questions through the PreMark assistant's tools. Deterministic gates make the trading decisions. An early recursive-feedback failure also led to a rule: model output doesn't become another model's input without human validation.
A winning trade can still be wrong
The measurement layer separates signal accuracy from execution quality. A trade that made money for the wrong reason gets marked as lucky. One of the screenshots below shows the assistant flagging the highest-P&L strategy as a windfall: it had been directionally wrong 94% of the time.
Frozen strategies stay frozen during evaluation. That makes it harder to keep adjusting a losing idea until its backtest tells me what I want to hear. Every agent message, tool call, and result is stored for review.
Research, paper trading, then calibration
Market scans feed a shared news-and-evidence ranker. Paper fleets evaluate variants before promising strategies move to small live calibration. Backtests use a 70/30 in-sample and out-of-sample split, and every deployment automatically pauses live strategies.
The services use Python, FastAPI, React, and Postgres, with pgvector for semantic search over market titles. The interesting question remains whether the measured edge survives further testing. I don't treat the existence of a trading platform as evidence that it does.
A closer look
On the parts list