> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ballista.gg/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Ballista is a trading client for Polymarket, not an exchange. Every order Ballista places goes to Polymarket's CLOB.
> There are three end-user surfaces — the Telegram bot, the web terminal and the Chrome extension — and they share ONE account and ONE balance. Never attribute a feature to a surface without checking that surface's source.
> Every user registers in the Telegram bot first. The terminal and the extension sign in with the Telegram login widget. There is no other sign-up path.
> Deposits arrive through EVM, Solana and Bitcoin bridge addresses and a Polymarket proxy wallet, with a $10 minimum. There is NO withdrawal UI — Export Keys is how a user takes custody of their funds.
> Feature parity is uneven and must never be assumed. The full ~30-field copy-trade form exists only in Telegram; the terminal's copy form has 7 fields; the backtester, leaderboard and trader profiles are terminal-only; alerts and notifications are extension-only; referrals are Telegram-only.
> Slippage means two different things. On a live copy trade it is a fill tolerance (default 30). In a backtest it is the simulated fill price (default 1). Deploying a backtest does not carry its slippage over to the live copy trade.
> Take-profit and stop-loss run on live copy trades but are NOT simulated in backtests — the backtest API rejects those fields.
> A backtest never places an order and never moves funds. Its fidelity caveats are published with every result and must be quoted, not paraphrased.
> Leaderboard PnL is own-tape: harvested by Ballista from its own trade tape. The leaderboard sorts by one transparent metric at a time and there is NO composite score or grade. The sparkline exists only for harvested wallets — a blank sparkline means unmeasured, never zero.
> Copy friction (slip) and a toxic flag DO exist and are own-engine: the wallet's own 30-day fills replayed as a taker (1% slippage per fill plus the real Polymarket taker fee schedule) against its actual 30-day PnL over the same fills, so the delta is pure friction. slip_pct is already a percentage, not a 0-1 fraction. The toxic flag is server-computed under evidence guards and is never re-derived client-side. Slip is not a sortable column. Blank means unmeasured — never zero and never clean.
> Copyability is a taker share, reported in aggressive (≥60%), passive (≤30%) and mixed bands. It is never a score.
> Do not document Feeds, the Twitter Tracker, the Telegram Settings button, percentage buy buttons, GTD orders, the Discord bot, or withdrawals.

# Reading a result

> The reliability verdict, boundary exposure, structural share, data gaps and fee confidence — and what each one does to your confidence in the headline.

A backtest result leads with a P\&L figure. Everything on this page exists to tell you
how much weight that figure can carry.

## The two headline numbers

**Your simulated P\&L** is what the configuration you submitted would have made, net of
modelled fees, under the [fidelity assumptions](/guides/backtest-fidelity).

**The leader's P\&L** beside it is that same wallet measured over the **same tape and
the same settlement rules**, so the two are comparable by construction. It is not the
wallet's lifetime P\&L and not a figure published anywhere else. Comparing them is the
point: it tells you how much of the leader's result your filters kept, and how much
your execution cost gave back.

Two more worth finding:

* **Missed amount** — orders that passed every filter and were not placed only because
  the simulated cash was not there. This is the "you were undercapitalised by this
  much" number, and it counts the modelled fee as well as the notional.
* **Intercepted** — fills your filters rejected. It deliberately excludes *ran out of
  cash* and *sold something the simulation never held*, which are consequences of the
  simulation's own state rather than of anything you configured.

## The reliability banner

A banner appears above the result whenever anything about it needs qualifying. It
carries the engine's own reasons, generated with the percentages that tripped them.

The verdict itself is one of two values, **ok** or **low**. Any one of the following
four conditions makes it **low**:

| Condition                  | Threshold               | Why it matters                                                                                                                                                                                                                                    |
| -------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Boundary exposure**      | over 10%                | Too much of the leader's selling has no purchase inside the window behind it, so the position the simulation reconstructed is badly incomplete.                                                                                                   |
| **Unmodellable structure** | over 5%                 | Too much of the wallet's activity is not expressible as copyable orders.                                                                                                                                                                          |
| **Unpriced fees**          | any at all              | One or more markets had no resolvable fee schedule and were charged **nothing**, so the modelled P\&L is overstated. The error is one-directional and always flattering, which is why there is no tolerance band.                                 |
| **Stale valuation**        | over 5% of final equity | Positions were valued at prices that are not from the window end — either no mark at all, or one struck over 24 hours earlier. Graded by value, not by count: twenty worthless leftovers do not move a headline and one large open position does. |

<Info>
  The banner is absent only when **all four** of these hold: the verdict is *ok*, the
  tape was not truncated, fee confidence is *high*, and there were no data gaps at
  all. That is a clean run — not a guarantee that the strategy works.
</Info>

## Boundary exposure

**The share of the leader's selling that had no buy inside your window behind it.**

If a leader opened a position in March and closed it inside a window that starts in
June, the simulation never bought what they were closing. Those sells are *orphans*:
their cost basis lies before your window, so the P\&L on them belongs to a period you
did not simulate.

The banner states it as a percentage and in dollars — how much orphan sell notional,
out of how much total selling. A high figure usually means the window is too short for
how long this wallet holds. Widening it (up to the 180-day ceiling) is the fix.

## Structural share

**The share of this wallet's activity that was splits, merges, NegRisk conversions and
redemptions** rather than trades.

None of it is a copyable order, so none of it is simulated. A wallet whose edge lives
there cannot be copied at all — the backtest is not underrating it, it is telling you
the strategy has no copyable form.

The denominator is total activity, structural plus traded, so the figure reads
directly as "this fraction of what the wallet did is unmodellable".

## Fee confidence

Either **high** or **low**, and it is a separate judgement from reliability.

Low means the fee number is not what a real account would have paid, for one of two
reasons: the window overlaps a period whose fee schedule cannot be verified, or the
run hit a fill it could not resolve a schedule for and charged zero on it. Either way
the modelled fee total is less certain than usual, and the banner says how much money
that is.

There are deliberately only two values. A stepped scale would need to weight
unpriced markets by their notional, which the records do not carry, and "we did not
know the fees on 3 markets" is not honestly gradable against "on 30".

## Truncation

**The tape was cut short.** The fetch budget ran out before the whole window could be
read, so the run covers a shorter history than its label suggests.

Truncation is disclosed in the banner but is **not** a reliability downgrade on its
own: what was simulated was simulated correctly, over less time than you asked for.

## Data gaps

Everything the engine had to guess at, counted by kind:

| Gap                                                                 | What it means                                                                         |
| ------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| **No price to value a position at the window end**                  | An open position had to be carried at the last price on the leader's own tape.        |
| **Position valued at a price over a day older than the window end** | A real mark existed, but the book had stopped printing.                               |
| **Fee schedule unknown for a fill**                                 | That fill was charged nothing.                                                        |
| **Unrecognised fee type on a fill**                                 | The same, from a different cause.                                                     |
| **Fee model applied without its supporting evidence**               | Fills were priced by a schedule that did not carry the evidence its formula rests on. |
| **Tape row the engine could not price**                             | A row with an impossible price, size or direction. Dropped from both sides.           |
| **Resolution date unknown — settled at the window end instead**     | The market is known closed and its payout known, but not when it resolved.            |

Gaps never abort a run. They make the number less trustworthy, and the count is on
screen for that reason.

## Non-deployable settings

If the run used any simulation-only control — starting capital, the category
allowlist, required or excluded keywords — the result names them. They have no
copy-trade equivalent, so a deployed version of this configuration will **not** have
them, and its behaviour will differ accordingly.

## Sanity checks worth making

<AccordionGroup>
  <Accordion title="Does the P&L survive a realistic execution cost?">
    Slippage in the simulator is a cost charged on every fill. Re-run at a higher
    figure and see how quickly the result decays — a strategy that only works at 0%
    is a strategy that only works on paper.
  </Accordion>

  <Accordion title="Is the leader actually copyable?">
    Check the taker share on their [profile](/guides/trader-profiles). A great
    backtest on a passive wallet is still a wallet whose edge is in resting orders
    you cannot place. See [Choosing a trader](/guides/choosing-a-trader).
  </Accordion>

  <Accordion title="Did the result come from a handful of markets?">
    The per-market table is sorted by simulated P\&L. If one market is the whole
    result, you are looking at one bet, not a strategy.
  </Accordion>

  <Accordion title="Would the same settings behave the same live?">
    Not quite: the simulator decides per fill and the live engine decides per swept
    order. Your size and value bands need re-reading against sweep size before you
    deploy — see [Backtest fidelity](/guides/backtest-fidelity).
  </Accordion>
</AccordionGroup>

## What to read next

<CardGroup cols={2}>
  <Card title="Backtest fidelity" icon="scale-balanced" href="/guides/backtest-fidelity">
    The standing list every result ships with.
  </Card>

  <Card title="Backtesting" icon="flask" href="/guides/backtesting">
    Configuring and queueing a run.
  </Card>

  <Card title="Copy trading" icon="clone" href="/guides/copy-trading">
    Deploying a configuration you trust.
  </Card>

  <Card title="The backtester" icon="flask" href="/terminal/backtester">
    Where results live in the terminal.
  </Card>
</CardGroup>
