A new approach to systematic investing
iClaudius Capital runs a reinforcement-learning system that manages a large US portfolio the way no human team could: applying perfectly consistent judgement across every stock at once, while still weighing each company’s own context on its own terms. A genuinely different way to invest — built to pursue returns that outpace simply holding the market.
For investors evaluating the strategy. Written to be readable without a finance background, with the technical detail a professional would ask for kept a click away.
The case
A handful of deliberate choices set the approach apart — each one made to matter to the person putting up capital, not just to the engineering.
Most models try to forecast whether a stock goes up. This one uses reinforcement learning to make portfolio decisions across the whole universe at once — judging every share relative to the others, the way a disciplined manager weighs a book rather than a single name.
The returns don't come from picking which company will win over years. They come from capturing price movement across the portfolio over horizons ranging from a day to three months — a fundamentally different, and more repeatable, source of return than long-term stock-picking.
The bar for going live is deliberately harsh: a strategy that's brilliant on average but breaks in one bad market never makes it into a real account. Resilience across conditions is the test — measured strictly on data the model was never trained on.
The system isn't limited to rising markets: alongside its core long portfolio it runs a small, capped short sleeve that can earn when prices fall, trained specifically on past downturns. But it's deliberately restrained — the book stays firmly net-long, the short exposure is strictly limited, and the aim is downside resilience, not aggressive betting against stocks.
Trading costs are built into what the model is trained to maximise, so it can't win by racking up trades that fees would quietly eat in a live account. It optimises what actually reaches the portfolio, not a flattering simulation.
Success is measured against simply buying and holding the same basket — so the system only gets credit for genuinely doing something smarter. And where there are known limitations, this site says so plainly rather than hiding them. Candour is part of the pitch.
The philosophy
Any strategy worth backing should be able to say where its returns come from — and why that source of return hasn't already been competed away. Here is the honest version.
Analysts, commentators and most investors form views name by name. This system reads the whole basket against itself every day — hundreds of relative comparisons refreshed daily, a breadth of attention no individual sustains.
Sometimes fundamentals drive returns; sometimes news and mood dominate; sometimes only momentum matters. Humans re-weight their thinking slowly. The model was trained across years of shifting conditions to adjust that balance continuously.
Prices routinely move too far on news and too little on quiet evidence. Harvesting that is less about seeing what others can't and more about acting on it consistently, without fear, boredom or fatigue — which is precisely what an automated, cost-aware daily cycle does.
The edge lives in the differences between shares. In markets where everything moves together and nothing distinguishes one name from another, there is less for the system to work with — and a future genuinely unlike anything in its training history is the risk that never fully goes away. That's why the deployment bar is worst-regime survival, and why hard limits sit outside the model.
Edison tested thousands of filaments before a light bulb that lasted. Dyson built 5,127 prototypes before a vacuum that worked. A learning trading strategy is the same kind of endeavour — overwhelmingly failure, until it isn't.
This system learns by trial and error across years of market history, one episode at a time. More than two million have been run so far, and the vast majority were discarded. What survives is only what cleared the deliberately harsh bar described above — nothing reaches a live account because it happened to look good once.
Generations
Before the current production model, which trades a substantially larger universe, an earlier-generation prototype traded a 12-stock universe on a simulated (paper) account for six months. Its equity curve is shown against buy-and-hold of its own basket — the same fair-benchmark test every generation must face.
Earlier-generation model, 12-stock universe, simulated paper-trading account, Jan–Jul 2026. Shown as development history only — simulated results do not include all real-world frictions, this is not the current strategy, and it is not indicative of future performance. The live, independently verified record is being built on Darwinex (CLUD).
The results
Performance is judged against a fair benchmark — how the system did versus simply holding the same basket. That choice does double duty: it strips out the market's own rise and fall, so what remains is genuine decision-making rather than an accidental bet on the market or a sector going up. A live, independently-verified track record is being built on Darwinex under the ticker CLUD, and is linked directly below; the chart on this page is illustrative only.
The shape above is illustrative. What matters is the live track record, measured the same honest way: returns after real trading costs, set against a simple market benchmark rather than a flattering one.
This is a long-term, iterative project — not a finished product making promises. The strategy is refined continually, and the verified record is what speaks for it.
See the live track recordUnder the hood
For those who want the mechanics. The system has two halves that run every trading day: first it builds one clean picture of the market, then a trained model reads that picture and decides what to trade. The rest of this section walks through both — in plain terms.
The big picture
The whole system reduces to this: data comes in, the model reads it, trades go out — then it all happens again tomorrow. Everything that follows is just these two halves in more detail.
Prices, company fundamentals, and news — gathered and cleaned automatically each morning.
A trained model reads the day's snapshot and weighs every share against the others.
Buy, hold, or trim decisions become real orders through the broker, sized with care.
Part One
What the system looks at, where it comes from, and how it becomes one clean picture the model can read.
§ 01 · The ingredients
Think of what a careful human trader would want on their desk every morning. The system gathers the same three things — just for a whole basket of US shares at once, every single day.
Daily price and trading volume for every share in the basket — the raw pulse of the market.
Earnings, analyst views, and the financial signals that say how a business is actually doing.
Economic and company news, read and scored so the mood around each share becomes a number.
§ 02 · The gathering
The point here isn't the plumbing. It's that the whole collection routine runs itself, on a schedule, and is built to keep an honest, unbroken record even when a source has an off day.
Fresh data is pulled automatically each day. Nobody has to sit and press a button for the system to stay current.
Incoming data is checked for holes and oddities, and gaps are filled sensibly so the record stays continuous.
If a source stalls, the system carries on from the last good record rather than breaking or inventing numbers.
Nothing dramatic. The system keeps the last reliable figures, flags that something was missing, and picks the source back up on the next run. A bad-data day never quietly turns into a bad-trade day. Prices arrive already adjusted for splits and dividends, and the stable large-company universe means corporate upheavals that scramble a price history are rare — and checked for rather than assumed away.
§ 03 · The translation
Raw prices and news aren't much use to a model on their own. They get translated into a consistent set of measurements — the same ones, computed the same way, for every share, every day.
How a share has been moving and how quickly — the measures that describe direction and strength.
Company health boiled down into comparable scores, so a strong balance sheet reads clearly to the model.
The tone of the news around each share, rated on a scale, so mood becomes something measurable.
§ 04 · The handoff
Everything in Part One exists to build one thing: a single tidy table, one row per share, refreshed each day. This is the picture the model reads — and the exact point where the data half hands over to the decision half.
| Share | Trend | Model reads → |
|---|---|---|
| MU | strong ↑ | candidate |
| SCHW | flat → | watch |
| ORCL | rising ↑ | candidate |
| DELL | slipping ↓ | review |
| … | … | … |
Part Two
How the model reads that daily snapshot and turns it into real, sensibly-sized trades — with guardrails.
§ 05 · The decision-maker
The model is the decision-maker. It reads the day's snapshot and, for each share, chooses one of a few simple moves: buy, hold, or trim a position. What makes it different from a rulebook is how it learned to choose.
Instead of following instructions like "buy when X happens," the model was trained on years of market history, learning which combinations of signals tended to pay off — and which didn't.
For each share the model buys, holds, or reduces a position — and, within strict limits, can take a modest short position to protect against or profit from a falling market. The book stays net-long throughout; complexity lives in reading the market well, not in exotic actions.
§ 06 · The judgement
The model doesn't look at each share in isolation. It weighs them against one another, and it leans in harder when a share's signals look unusual — the way a good analyst's eye is drawn to something out of the ordinary.
Attention is shared across the whole basket, so the portfolio stays balanced rather than betting the house on one name.
When a share's signals are extreme or out of character, the model gives it more weight — catching moments that a flat, mechanical rule would miss.
§ 07 · The last mile
A decision on a screen isn't a trade. The system turns each choice into an actual order through the broker — sized sensibly, and deliberately avoiding trades so small the costs would eat them.
Every trade carries a fee. The system won't place a trade so tiny that the fee swallows the benefit — a small but real discipline that stops the strategy quietly leaking money on pointless micro-trades.
§ 08 · The guardrails
The honest answer to "how do I know it won't do something reckless?" A handful of hard limits sit around the model, independent of whatever it decides on any given day.
No single share can grow beyond a set share of the portfolio, and total short exposure is strictly capped with the book kept net-long — so no one bet, long or short, can sink the whole book.
The minimum-trade rule keeps the system from churning through fees on trades too small to matter.
Suspect or missing data is caught before it can drive a trade, so a data glitch never becomes a bad order.
If the machine, connection or broker link drops, the portfolio is a net-long book of liquid shares with only a small, capped short sleeve — so an outage means it largely just holds until the system is back, rather than being forced into hurried trades while dark.
§ 09 · The gatekeeping
The production model is not tweaked, tuned or swapped on a whim. Every candidate faces the same gauntlet, and the incumbent keeps its seat until a challenger beats it on identical terms.
Candidate models are trained in their thousands on historical data, then frozen. A candidate is never adjusted after the fact to make its results look better — what was trained is what gets judged.
Every candidate is scored strictly on market periods it never saw in training, across many separate test windows spanning years — with deliberate gaps between training and testing so nothing leaks across the boundary.
The gate looks at a candidate's worst test window, not its average. Brilliant-but-fragile fails. Only after clearing that bar does a candidate face the incumbent head-to-head, then a period of simulated trading, before real money.
Changes to the production model are rare, deliberate, versioned and reversible — a challenger is promoted only when it beats the current model under the full test regime, never because a chart looked promising, and there is no manual override of the live model's daily decisions.
§ 10 · The shape
Beyond individual buy and sell decisions, the portfolio itself has a deliberate architecture — the constraints that hold regardless of what the model wants on any given day.
A hard cap limits how much of the portfolio any one share can occupy. Conviction is allowed; betting the book on one company is not.
The system is never forced to be fully invested. Holding cash is a legitimate position, and in conditions it reads as unattractive, it will simply do less.
Because costs are inside the objective the model is trained on, turnover is self-limiting — the system trades when the expected benefit clears the real-world cost of the trade, and not otherwise.
The strategy trades highly liquid, large US listed companies on a daily cycle. At its current and foreseeable scale, its orders are a negligible fraction of daily traded volume in those names — capacity is not the binding constraint; discipline is.
Common questions
Yes. The system gathers its own data, makes its own decisions, and places its own orders on a daily cycle. A person watches over it, but the day-to-day trading runs without manual input.
A fixed group of large US companies, all drawn from the S&P 500. It doesn't roam the entire market chasing tips; it works the same defined universe every day, which keeps its behaviour consistent and measurable.
One deliberate exclusion: the "Magnificent 7" mega-caps (the handful of dominant names like Apple, Microsoft, Nvidia and their peers). The reason is about signal quality. These companies generate an enormous volume of news and attention every single day, most of which is noise rather than genuinely price-relevant. Because the model judges every stock relative to the others, letting a few names carry that outsized, noisy information load would distort the comparison. Removing them keeps every stock on a comparable footing — so the model's relative judgements stay meaningful rather than being dominated by whichever mega-cap is in the headlines that day.
It reviews the whole basket once per trading day and adjusts where the model sees reason to. It's not a high-frequency system firing thousands of trades — and the cost rule actively discourages needless churn.
Better placed than a purely long portfolio. Position limits cap any single exposure, and the model can both reduce its long holdings and lean on a small, capped short sleeve trained specifically on past downturns — so it can cushion, or even profit from, a falling market rather than simply ride it down. It stays net-long throughout, so this is protection, not a bet on collapse — and no honest system can promise it never loses.
Consistency. It reads the same signals the same way every day, across the whole basket at once, without fatigue, mood, or the pull of a dramatic headline. It's a disciplined process, not a personality.
For the sceptical reader
The questions a professional allocator actually asks — how the models are trained, on what data, and how the well-known traps are handled. No hand-waving.
It's deliberately not supervised prediction — there's no model trying to answer "will this stock go up?" It's reinforcement learning: the model learns a policy for managing a whole portfolio, and crucially it looks at the entire universe at once rather than judging one name in isolation. How each stock looks relative to the rest of the universe is part of every decision.
It's trained to maximise risk-adjusted return net of realistic trading costs — not raw return. Real-world costs are built into what the model is rewarded for, so it can't win by racking up trades that would be eaten by fees in a live account. It optimises the thing that actually matters to a real portfolio, not a flattering paper proxy.
By trial and error across many years of historical market conditions — the model proposes portfolio decisions, sees how they would have played out, and gradually improves the policy that produced them. This is repeated over an enormous number of runs (see the note on failure above), with only the strategies that clear the validation bar ever reaching a live account.
The specifics of the architecture, the exact inputs, and the training and validation regime are proprietary and deliberately not detailed here — but the principles that keep it honest, especially the out-of-sample discipline and the beat-the-basket benchmark, are described throughout this section.
Everything is walk-forward: the model is only ever scored on data strictly after its training window, across many independent out-of-sample periods spanning several years of market history. It is never judged on data it has seen.
The deployment bar is intentionally severe: a model must survive its worst period, not just look good on average. A strategy that's brilliant on average but blows up in one market regime doesn't make it into production. Average competence isn't enough — resilience across regimes is the test.
The benchmark isn't cash. It's always-buy of the same fixed basket — the identical group of S&P 500 stocks the model chooses from — so the model only earns credit for doing something genuinely smarter than simply holding them all. And the whole pipeline is built with rigorous checks against look-ahead — features are constructed so the model can never accidentally learn from information that wouldn't have been available at the time.
Everything comes from a single institutional-grade market-data provider: daily prices and volume, company fundamentals, analyst ratings and earnings revisions, macro/economic series, and news. Using one high-quality source avoids a whole class of data-mismatch problems that come from stitching vendors together.
News is handled with more care than a simple positive/negative score. Rather than taking headlines at face value, the system also accounts for how much the news agrees or disagrees across the market, and leans on news information more in some conditions than others. Exactly how that's done is part of what makes the approach distinctive, and isn't spelled out here.
Not DCF, and not analyst-style valuation. Fundamentals enter as features: analyst rating levels and their changes, earnings-revision momentum, and cross-sectional ranks — where each stock sits relative to the rest of the universe on each dimension, each day. The model learns how much weight those deserve in each market regime, rather than following weights hard-coded by hand.
The cross-sectional ranking turned out to matter a great deal. In the Feb–May 2026 semiconductor selloff, the names that recovered were separable on analyst ratings and earnings revisions — but not on price momentum alone. A model looking only at price would have missed the distinction the fundamentals made visible.
These are the traps that quietly wreck backtests. Handled one by one, including the honest limitations:
For the sceptics
The questions a professional allocator would ask — answered the way they deserve to be answered, which is directly.
A sustained failure to beat the fair benchmark — buy-and-hold of the same basket — beyond pre-defined tolerances. That judgement is made by rule, against thresholds set in advance, not by mood after a bad month. Temporary underperformance inside historical norms is expected and tolerated; a breach of the pre-set floor triggers de-risking. A system with a harsh bar for going live must have an equally clear bar for stepping back.
A future genuinely unlike anything in the training history. No amount of testing on the past fully removes that risk, and it would be dishonest to claim otherwise. The mitigations are structural: the worst-regime deployment gate, hard limits that sit outside the model's control, a strictly capped short exposure with the book kept net-long, and cash as an always-available position.
Several independent defences, all mechanical: every engineered input is tested by perturbing future data and asserting that today's values don't change; testing is strictly walk-forward with deliberate gaps between training and test windows; a final reserved period of history is looked at exactly once, at the end; the deployment gate judges the worst window, not the average; and the benchmark is buy-and-hold of the same basket, so the system gets no credit for a rising market.
The production loop — data gathering, signal construction, decisions, execution, record-keeping — is fully automated, and the live model's daily decisions cannot be manually overridden. Human judgement operates one level up: choosing research directions, approving the promotion of a new model after it clears the full test regime, and deciding the rules themselves. People design the gauntlet; they don't reach into the machine.
Any specific stretch of returns reflects the regimes that happened to occur, and no particular window should be extrapolated. What is designed to repeat is the process: breadth of daily attention, cost-aware discipline, and survival-first gating. That is also why performance is judged against the fair benchmark over long horizons rather than sold on its best quarter.
Partly, yes — choosing a basket of currently listed companies bakes some hindsight into the universe itself, and pretending otherwise would be dishonest. The mitigation is structural: the benchmark is buy-and-hold of the same basket, so whatever survivorship advantage the universe enjoys, the benchmark enjoys it equally — the system is only ever credited for beating a yardstick with the identical bias. These are also deeply established large companies where the risk of disappearance over the tested period is low, and the universe is fixed rather than quietly reshuffled.
Far less than most systematic strategies, by construction. This is not a standard factor portfolio — there is no published recipe being followed alongside billions of dollars of similar money. The policy was learned, not copied; it applies to one specific basket; and its capacity needs are a rounding error in some of the world's most liquid shares. The honest caveat: crowding in the underlying names (large US equities are widely held) is unavoidable and shows up in the benchmark too — which is exactly why the benchmark is the yardstick.
By pre-set rule, not by feel. Underperformance within the range seen across years of testing is a bad patch; a sustained breach of the pre-defined floor against the benchmark is degradation, and it triggers de-risking. Refresh runs continuously but promotes rarely: challenger models are trained and tested all the time, yet the production model changes only when a challenger beats it through the full gauntlet — the incumbent is never retired for novelty's sake.
More than it might seem. The code is the smaller half; the accumulated record of what was tried and failed — thousands of discarded candidates and the reasons they died — is what steers the next generation, and it doesn't ship with the code. The validation discipline would also force any copy through the same multi-year gauntlet before it could be trusted. And a copy trades against the same costs and the same market, with none of the iteration behind it.
A quick vocabulary
The project
iClaudius Capital is an independent systematic trading operation. The goal is straightforward: run a disciplined, fully-automated strategy well enough, and for long enough, to build a verified public track record that can stand on its own.
That record — accumulated on a recognised platform where every trade is logged and independently measured — is the point. It's what turns a private strategy into something a third party can trust and, in time, back.
How this is run. Every function of the operation is systematised and auditable. Data acquisition runs on an automated daily pipeline with integrity checks; every engineered feature is leak-tested against future information; candidate models must clear a multi-year, multi-fold walk-forward validation gate — judged on their worst period, not their average — before any capital is committed; and live execution is fully automated, with hard risk limits that sit outside the model's control. Trades are logged and measured by an independent third-party platform, not self-reported.
The work is continual. Models are tested, refined, and replaced as the evidence dictates; nothing here is presented as finished or guaranteed.
Questions about the strategy or the track record? Enquiries are welcome by email.