BBenevente

BENEVENTE QUANT AI · ENGLISH RESEARCH VERSION

Esta página documenta a pesquisa em inglês, o idioma do manuscrito submetido. O restante do site está em português.

Separating what the system calculates from what language explains.

The production artifact uses dated data and deterministic code to select and weight assets. Only after the proposal is sealed does a language model explain approved facts, risks and review questions. A separate retrospective experiment relaxed that boundary to test what model discretion would add. It found no detectable return benefit and cannot certify temporal purity.

One term, defined onceCDI is the Brazilian interbank overnight rate, the benchmark a domestic investor earns on cash with negligible credit risk. It is the local equivalent of a risk-free short rate, and it is high: it compounded at 9.61% a year over the evaluation window. Any Brazilian equity strategy is therefore competing against a cash return that would be unusual elsewhere, which is why it appears as a benchmark throughout and why the defensive sleeve is held in it rather than in bills.

1 · Why the separation exists

Portfolio weights must sum to one, respect policy limits and remain traceable to the data admitted at the decision date. Free text cannot guarantee those properties. The separation is structural: deterministic code creates eligibility, ranking and weights. Language receives a read-only decision record. A named human accepts or rejects the proposal.

Production contractDated data → deterministic rule → valid weights → explanation → human review. There is no path from generated prose back to selection, limits or weights.

2 · Pipeline

Each January the system executes five stages in a fixed order. Universe, filings, screen and allocation are deterministic. Language enters only after the proposal is sealed, and the last stage is human.

2.1 Dated universe

The universe of record is the exchange's own historical quotation archive, not a vendor feed. This matters more than it sounds: a public adjusted-price provider serves the companies that still exist, so building a universe from one silently deletes every company that was delisted, acquired or wound up. In this panel that is 166 of 514 issuers, 77 of which disappear before 2020, and 58 of which had already fallen more than 60% from their own peak when they stopped trading. Rebuilding from the exchange archive restores them.

Corporate actions are recovered in two ways: confirmed against an archived provider event feed where one exists, and otherwise detected from round, persistent price ratios. The detector's precision and recall against the provider's own records are published with the panel rather than assumed, because for the delisted tail it is the only evidence available.

2.2 Point-in-time fundamentals

Fundamentals come from the regulator's standardised filings, carrying both the reference date and the date the filing was received. Only filings received on or before the decision date can enter. Trailing-twelve-month figures are bridged from the annual statement plus the current and comparative interim periods, using the year-to-date column only, the interim filing also publishes the isolated quarter, and mixing the two silently corrupts every ratio from the second quarter onwards.

2.3 Screen

Assets must clear liquidity, profitability, returns on capital and solvency. Operating companies and financial institutions are screened on different quantities because their standardised statements are not comparable. Solvency for operating companies is derived from the standardised accounts rather than left undefined, which is what previously caused an accidental sector concentration.

2.4 Ranking and allocation

Survivors are ranked by a declared composite of quality, valuation yield and twelve-month momentum. The annual configuration is chosen only with years already closed. Deterministic code allocates inside the ceilings and holds the residual at the domestic overnight rate. MVO is an independent comparator, not the Benevente selector.

2.5 Explanation and human review

The sealed proposal carries the policy, rejection reasons, weights, order sizes and estimated cost. The language model turns those approved facts into a readable note but cannot modify them. A person records acceptance or rejection. No order is transmitted.

3 · Holding, cost and tax

The book is bought in January and held. This is worth stating because the natural vectorised expression of a backtest, compounding the inner product of returns and weights, silently rebalances to fixed weights every session at no cost, which is neither the protocol being evaluated nor something an investor can execute.

Costs are the exchange's published fee plus a slippage term that grows with the position's participation in the asset's average traded volume, so a thinly traded name is not assumed to trade like a liquid one. Taxation follows the domestic regime: gains realised at a review are taxed at the equity rate, the cash sleeve at the fixed-income rate, and positions carried into the next year defer their liability exactly as the law allows.

4 · How the published rule was chosen

This is the part that most backtests get wrong, and it was wrong here first. A grid of candidate rules was evaluated, the best was published, and its rank improved from 57th on the declared training criterion to 1st on the holdout, which is only possible if the holdout was read. Nothing downstream of that choice is out of sample.

The next correction was a nested protocol. For decision year t, the configuration was ranked using only years that had already closed, by Sharpe of the excess return over cash. Year t was then evaluated once. Changing configuration was charged a full-turnover rebalance, because in a real book it means liquidating one portfolio to buy another. Three quantities were published alongside every result: the nested outcome, the full-sample winner, and the difference between them, the hindsight premium, 0.65 points a year over 36 configurations.

That fixed reading the holdout. It did not fix searching. In 2026 the same procedure was run over a wider grid, 256 candidates instead of 36, on identical inputs, code and window, and the result inverted the expectation. The nested outcome fell from 15.31% to 12.68% a year, and the deflated Sharpe fell from 0.957 to 0.777, below significance, because the expected maximum Sharpe under the null rose from 0.375 to 0.746. The selector wandered into noise: for 2018 it chose a low-volatility configuration that returned −7.2% in a year when cash returned 6.4%.

Cost of widening the search−2.63 ppper year, 36 → 256 candidates, identical inputs and window
Deflated Sharpe0.957 → 0.777significant at 95% before widening, not after
Tactical allocation0 of 7predictive rules beat the static weight over 118 months

Ten annual observations cannot rank 256 candidates. Widening the search does not improve the system. It degrades the selector. Each profile's configuration is therefore declared and frozen before the period, with a public hash, rather than chosen by search. The nested search remains published as the experiment that established its own limit.

4.1 Does deciding more often help?

A distinct question from the one above, and the one most readers ask first: not whether the equity share should move more often, but whether the basket itself should be reselected more often. Holding the rule, the limits and the panel fixed and varying only the decision dates, annual reselection compounded at 18.58% before costs, quarterly at 16.37% and monthly at 13.92%. Paired by calendar year and measured after tax, quarterly trails the annual decision by 1.55 percentage points a year (p = 0.376) and monthly by 3.27 (p = 0.306). Eleven paired years cannot support the stronger claim. This sample provides no evidence that more frequent decisions help, but it does not reject a benefit in other periods, markets or signal families.

The mechanism is not the one usually assumed. The loss sits in the gross return, which falls 4.66 points from annual to monthly before any deduction. Execution cost rises by only 0.08 points a year, and modelled tax actually falls, because each monthly review realises a smaller slice of the book. Reselecting monthly degrades the selection before it raises the bill. The tax model understates the monthly arm's liability, since it charges tax on each period's own gain rather than on the accumulated gain of the position actually sold, a bias in favour of the arm that lost anyway.

The last figure is a separate study. Perfect monthly timing of the equity-versus-cash split would have been worth roughly 50 percentage points a year, so the prize is real. None of seven predictive rules, fitted walk-forward on observable state, captured any of it, across three training-window lengths.

5 · Everything that was tried

Seven hypotheses were tested against the published protocol. None survived. They are listed together because a reader assessing a method needs the failure inventory as much as the result, and because each one closes a specific objection that would otherwise remain open.

HypothesisHow it was testedOutcome
A wider search finds a better ruleThe same nested procedure over 256 candidates instead of 36, identical inputs and window−2.63 points a year. Deflated Sharpe 0.957 → 0.777, below significance.
Positions should be sized by inverse volatilityEqual, inverse-volatility and geometric-blend sizing against the published confidence weighting, over 8 configurationsThe published rule won 8 of 8. Inverse volatility bought 0.95 points of drawdown for 2.11 points of return.
Risk can be sized from volatility observed in JanuaryAn annual volatility target scaling the equity sleeve at the decision dateCut exposure in 5 of 13 years, always after a stress and never before one, and left the worst drawdown untouched.
The year's regime can be forecastFive January-observable predictors, expanding window, 8 yearsBest predictor 3 of 4 live calls, p = 0.31. No signal.
The equity share should move more oftenSeven continuous rules, 118 monthly and 521 weekly periodsNone beat the static weight. Mean-variance was significantly worse.
The basket should be reselected more oftenSame rule and limits, cadence varied to quarterly and monthly−1.55 and −3.27 points a year after tax, p = 0.38 and 0.31.
A language model adds returnFour arms over 13 years, named / anonymised / deterministic / monolithicThree nulls. The model reproduces the deterministic ranking.

Two of these are reported as "no evidence of benefit" rather than "proven harmful", and the distinction is not cosmetic. Eleven paired annual observations cannot reject a null of no difference, so the honest claim is the weaker one, even where every point estimate leans the same way.

6 · Experiment: four arms, one universe

The same January decision was executed four ways over thirteen annual decisions, holding the eligible universe, the optimiser, the ceilings and the cost model constant. Model: gemini-3.1-flash-lite, temperature zero, structured output validated against a schema, every raw response archived.

ArmWhat differsCAGRBeats cash
namedModel sees ticker and sector13.00%7 of 13
deterministicPre-declared factor replaces the model12.67%8 of 13
anonymisedDated figures only, assets unidentifiable12.62%10 of 13
monolithicModel returns weights directly11.81%7 of 13
Cash (CDI)Not applicable9.59%Not applicable

The anonymised arm is a contamination sensitivity test, not a guarantee. A model may infer a company from its financial attributes, and both named and anonymised prompts can activate memory acquired after the historical date. Only prospective observation after a frozen registration can validate a predictive language component. The deterministic arm separately tests whether the model adds anything beyond the structured screen.

HypothesisAnnualised gappReading
The model adds over the factor−0.05 pp0.99Indistinguishable from zero
Decoupling beats the monolithic prompt+0.81 pp0.78Indistinguishable from zero
Recognising the company helps+0.38 pp0.91Indistinguishable from zero
The retrospective result is nullNone of the three effects survives testing. The model reproduces the deterministic ranking. Direct model-generated weights failed to sum to one in five of thirteen years and silently omitted eligible names in two. This supports keeping language away from allocation, but it does not prove contamination absent or validate the model. The production system therefore restricts language to explanation. Any future predictive role requires a new prospective arm.

7 · Published strategies

Declared profileEquityIssuersCAGR 2015 a 2025Max drawdown
Ultra-conservative, annual selection only4%129.83%−1.69%
Ultra-conservative, with the risk overlay4%129.73%−0.81%
Conservative, annual selection only35%1213.22%−16.71%
Conservative, with the risk overlay35%1212.34%−9.17%
Balanced, annual selection only55%816.25%−26.54%
Balanced, with the risk overlay55%815.39%−17.88%
Aggressive, annual selection only75%520.26%−35.64%
Aggressive, with the risk overlay75%519.79%−28.95%
Ibovespa (total-return index)——11.74%−46.82%
Cash (Tesouro Selic, net of custody)——9.36%−0.60%

There is no single Benevente number. Each profile declares an equity budget, an issuer count and a fifth of that budget in the global sleeve, and each pair of rows shows what the intra-year overlay costs and returns. Every profile beat cash in 8 of 11 years, and the risk ordering holds: the one that returned more also fell further. The overlay reduces drawdown by 6.7 to 8.6 points for under one point of annual return, a far better trade than inverse-volatility sizing, which bought 0.95 points of drawdown for 2.11 points of return.

The evaluation window is also the window that developed the rule, so it describes the sample. The 2026 allocation predates the registered version and remains a shadow portfolio under the policy that decided it. Prospective evaluation starts with the January 2027 decision, made under a registered version that is not changed after outcomes are observed.

8 · Reproduction contract

The repository publishes code, configuration grid, dated inputs where licensing permits, tests, raw experiment outputs and SHA-256 manifests. The website explains the method without exposing source code in the interface. The paper and supplementary package remain the scientific record. The frozen protocol cannot be changed after an outcome without creating a new version and a new prospective count.

9 · Limitations

Benevente Quant AI · reproducible research. Figures on this page are read from the same files the engine writes. This is not investment advice, an offer, or a promise of return. Historical results describe the sample, not the future.

Contato: [email protected]