BENEVENTE QUANT AI · ENGLISH RESEARCH VERSION
Esta página documenta a pesquisa em inglês, o idioma do manuscrito submetido. O restante do site está em português.
Separating what the system calculates from what language explains.
The production artifact uses dated data and deterministic code to select and weight assets. Only after the proposal is sealed does a language model explain approved facts, risks and review questions. A separate retrospective experiment relaxed that boundary to test what model discretion would add. It found no detectable return benefit and cannot certify temporal purity.
1 · Why the separation exists
Portfolio weights must sum to one, respect policy limits and remain traceable to the data admitted at the decision date. Free text cannot guarantee those properties. The separation is structural: deterministic code creates eligibility, ranking and weights. Language receives a read-only decision record. A named human accepts or rejects the proposal.
2 · Pipeline
Each January the system executes five stages in a fixed order. Universe, filings, screen and allocation are deterministic. Language enters only after the proposal is sealed, and the last stage is human.
2.1 Dated universe
The universe of record is the exchange's own historical quotation archive, not a vendor feed. This matters more than it sounds: a public adjusted-price provider serves the companies that still exist, so building a universe from one silently deletes every company that was delisted, acquired or wound up. In this panel that is 166 of 514 issuers, 77 of which disappear before 2020, and 58 of which had already fallen more than 60% from their own peak when they stopped trading. Rebuilding from the exchange archive restores them.
Corporate actions are recovered in two ways: confirmed against an archived provider event feed where one exists, and otherwise detected from round, persistent price ratios. The detector's precision and recall against the provider's own records are published with the panel rather than assumed, because for the delisted tail it is the only evidence available.
2.2 Point-in-time fundamentals
Fundamentals come from the regulator's standardised filings, carrying both the reference date and the date the filing was received. Only filings received on or before the decision date can enter. Trailing-twelve-month figures are bridged from the annual statement plus the current and comparative interim periods, using the year-to-date column only, the interim filing also publishes the isolated quarter, and mixing the two silently corrupts every ratio from the second quarter onwards.
2.3 Screen
Assets must clear liquidity, profitability, returns on capital and solvency. Operating companies and financial institutions are screened on different quantities because their standardised statements are not comparable. Solvency for operating companies is derived from the standardised accounts rather than left undefined, which is what previously caused an accidental sector concentration.
2.4 Ranking and allocation
Survivors are ranked by a declared composite of quality, valuation yield and twelve-month momentum. The annual configuration is chosen only with years already closed. Deterministic code allocates inside the ceilings and holds the residual at the domestic overnight rate. MVO is an independent comparator, not the Benevente selector.
2.5 Explanation and human review
The sealed proposal carries the policy, rejection reasons, weights, order sizes and estimated cost. The language model turns those approved facts into a readable note but cannot modify them. A person records acceptance or rejection. No order is transmitted.
3 · Holding, cost and tax
The book is bought in January and held. This is worth stating because the natural vectorised expression of a backtest, compounding the inner product of returns and weights, silently rebalances to fixed weights every session at no cost, which is neither the protocol being evaluated nor something an investor can execute.
Costs are the exchange's published fee plus a slippage term that grows with the position's participation in the asset's average traded volume, so a thinly traded name is not assumed to trade like a liquid one. Taxation follows the domestic regime: gains realised at a review are taxed at the equity rate, the cash sleeve at the fixed-income rate, and positions carried into the next year defer their liability exactly as the law allows.
4 · How the published rule was chosen
This is the part that most backtests get wrong, and it was wrong here first. A grid of candidate rules was evaluated, the best was published, and its rank improved from 57th on the declared training criterion to 1st on the holdout, which is only possible if the holdout was read. Nothing downstream of that choice is out of sample.
The next correction was a nested protocol. For decision year t, the configuration was ranked using only years that had already closed, by Sharpe of the excess return over cash. Year t was then evaluated once. Changing configuration was charged a full-turnover rebalance, because in a real book it means liquidating one portfolio to buy another. Three quantities were published alongside every result: the nested outcome, the full-sample winner, and the difference between them, the hindsight premium, 0.65 points a year over 36 configurations.
That fixed reading the holdout. It did not fix searching. In 2026 the same procedure was run over a wider grid, 256 candidates instead of 36, on identical inputs, code and window, and the result inverted the expectation. The nested outcome fell from 15.31% to 12.68% a year, and the deflated Sharpe fell from 0.957 to 0.777, below significance, because the expected maximum Sharpe under the null rose from 0.375 to 0.746. The selector wandered into noise: for 2018 it chose a low-volatility configuration that returned −7.2% in a year when cash returned 6.4%.
Ten annual observations cannot rank 256 candidates. Widening the search does not improve the system. It degrades the selector. Each profile's configuration is therefore declared and frozen before the period, with a public hash, rather than chosen by search. The nested search remains published as the experiment that established its own limit.
4.1 Does deciding more often help?
A distinct question from the one above, and the one most readers ask first: not whether the equity share should move more often, but whether the basket itself should be reselected more often. Holding the rule, the limits and the panel fixed and varying only the decision dates, annual reselection compounded at 18.58% before costs, quarterly at 16.37% and monthly at 13.92%. Paired by calendar year and measured after tax, quarterly trails the annual decision by 1.55 percentage points a year (p = 0.376) and monthly by 3.27 (p = 0.306). Eleven paired years cannot support the stronger claim. This sample provides no evidence that more frequent decisions help, but it does not reject a benefit in other periods, markets or signal families.
The mechanism is not the one usually assumed. The loss sits in the gross return, which falls 4.66 points from annual to monthly before any deduction. Execution cost rises by only 0.08 points a year, and modelled tax actually falls, because each monthly review realises a smaller slice of the book. Reselecting monthly degrades the selection before it raises the bill. The tax model understates the monthly arm's liability, since it charges tax on each period's own gain rather than on the accumulated gain of the position actually sold, a bias in favour of the arm that lost anyway.
The last figure is a separate study. Perfect monthly timing of the equity-versus-cash split would have been worth roughly 50 percentage points a year, so the prize is real. None of seven predictive rules, fitted walk-forward on observable state, captured any of it, across three training-window lengths.
5 · Everything that was tried
Seven hypotheses were tested against the published protocol. None survived. They are listed together because a reader assessing a method needs the failure inventory as much as the result, and because each one closes a specific objection that would otherwise remain open.
| Hypothesis | How it was tested | Outcome |
|---|---|---|
| A wider search finds a better rule | The same nested procedure over 256 candidates instead of 36, identical inputs and window | −2.63 points a year. Deflated Sharpe 0.957 → 0.777, below significance. |
| Positions should be sized by inverse volatility | Equal, inverse-volatility and geometric-blend sizing against the published confidence weighting, over 8 configurations | The published rule won 8 of 8. Inverse volatility bought 0.95 points of drawdown for 2.11 points of return. |
| Risk can be sized from volatility observed in January | An annual volatility target scaling the equity sleeve at the decision date | Cut exposure in 5 of 13 years, always after a stress and never before one, and left the worst drawdown untouched. |
| The year's regime can be forecast | Five January-observable predictors, expanding window, 8 years | Best predictor 3 of 4 live calls, p = 0.31. No signal. |
| The equity share should move more often | Seven continuous rules, 118 monthly and 521 weekly periods | None beat the static weight. Mean-variance was significantly worse. |
| The basket should be reselected more often | Same rule and limits, cadence varied to quarterly and monthly | −1.55 and −3.27 points a year after tax, p = 0.38 and 0.31. |
| A language model adds return | Four arms over 13 years, named / anonymised / deterministic / monolithic | Three nulls. The model reproduces the deterministic ranking. |
Two of these are reported as "no evidence of benefit" rather than "proven harmful", and the distinction is not cosmetic. Eleven paired annual observations cannot reject a null of no difference, so the honest claim is the weaker one, even where every point estimate leans the same way.
6 · Experiment: four arms, one universe
The same January decision was executed four ways over thirteen annual decisions, holding the eligible universe, the optimiser, the ceilings and the cost model constant. Model: gemini-3.1-flash-lite, temperature zero, structured output validated against a schema, every raw response archived.
| Arm | What differs | CAGR | Beats cash |
|---|---|---|---|
| named | Model sees ticker and sector | 13.00% | 7 of 13 |
| deterministic | Pre-declared factor replaces the model | 12.67% | 8 of 13 |
| anonymised | Dated figures only, assets unidentifiable | 12.62% | 10 of 13 |
| monolithic | Model returns weights directly | 11.81% | 7 of 13 |
| Cash (CDI) | Not applicable | 9.59% | Not applicable |
The anonymised arm is a contamination sensitivity test, not a guarantee. A model may infer a company from its financial attributes, and both named and anonymised prompts can activate memory acquired after the historical date. Only prospective observation after a frozen registration can validate a predictive language component. The deterministic arm separately tests whether the model adds anything beyond the structured screen.
| Hypothesis | Annualised gap | p | Reading |
|---|---|---|---|
| The model adds over the factor | −0.05 pp | 0.99 | Indistinguishable from zero |
| Decoupling beats the monolithic prompt | +0.81 pp | 0.78 | Indistinguishable from zero |
| Recognising the company helps | +0.38 pp | 0.91 | Indistinguishable from zero |
7 · Published strategies
| Declared profile | Equity | Issuers | CAGR 2015 a 2025 | Max drawdown |
|---|---|---|---|---|
| Ultra-conservative, annual selection only | 4% | 12 | 9.83% | −1.69% |
| Ultra-conservative, with the risk overlay | 4% | 12 | 9.73% | −0.81% |
| Conservative, annual selection only | 35% | 12 | 13.22% | −16.71% |
| Conservative, with the risk overlay | 35% | 12 | 12.34% | −9.17% |
| Balanced, annual selection only | 55% | 8 | 16.25% | −26.54% |
| Balanced, with the risk overlay | 55% | 8 | 15.39% | −17.88% |
| Aggressive, annual selection only | 75% | 5 | 20.26% | −35.64% |
| Aggressive, with the risk overlay | 75% | 5 | 19.79% | −28.95% |
| Ibovespa (total-return index) | — | — | 11.74% | −46.82% |
| Cash (Tesouro Selic, net of custody) | — | — | 9.36% | −0.60% |
There is no single Benevente number. Each profile declares an equity budget, an issuer count and a fifth of that budget in the global sleeve, and each pair of rows shows what the intra-year overlay costs and returns. Every profile beat cash in 8 of 11 years, and the risk ordering holds: the one that returned more also fell further. The overlay reduces drawdown by 6.7 to 8.6 points for under one point of annual return, a far better trade than inverse-volatility sizing, which bought 0.95 points of drawdown for 2.11 points of return.
The evaluation window is also the window that developed the rule, so it describes the sample. The 2026 allocation predates the registered version and remains a shadow portfolio under the policy that decided it. Prospective evaluation starts with the January 2027 decision, made under a registered version that is not changed after outcomes are observed.
8 · Reproduction contract
The repository publishes code, configuration grid, dated inputs where licensing permits, tests, raw experiment outputs and SHA-256 manifests. The website explains the method without exposing source code in the interface. The paper and supplementary package remain the scientific record. The frozen protocol cannot be changed after an outcome without creating a new version and a new prospective count.
9 · Limitations
- Eleven annual evaluations for the strategy and thirteen for the model experiment. Every test is underpowered. Effect sizes are reported with p-values rather than as findings.
- Corporate actions for the delisted tail rest on a detector whose precision and recall are published but imperfect. Distributions for issuers no provider still serves are imputed from the cross-section and flagged per ticker.
- The named--anonymised gap cannot certify temporal purity. Identity may be inferred from fundamentals and pretrained memory is unobservable.
- Coverage is no longer Brazil only: a declared sleeve holds a B3-listed fund tracking the S&P 500 in reais. Roughly a third of that fund's return over the window came from currency depreciation rather than the American market, so the sleeve is an unhedged long dollar position and must be read as one. Partial-year monitoring uses exchange closing prices and excludes distributions until corporate events are reconciled against primary records.
- The sector limit of three issuers is a governance constraint, not a performance improvement. Where it binds it cost 0.15 points a year and improved the excess Sharpe in none of the twenty configurations affected.
- Choosing between the two ways of combining the risk overlay with the global sleeve was itself a selection over two candidates, and is declared as such in the frozen registration rather than presented as the only option considered.
Benevente Quant AI · reproducible research. Figures on this page are read from the same files the engine writes. This is not investment advice, an offer, or a promise of return. Historical results describe the sample, not the future.
Contato: [email protected]