HexMathv1.0.3
Proof & methodology

How we know HexMath works.

Proof you can check, not a number with nothing behind it.

HexMath is an agentic AI solving system: it reads a problem, plans the solution, solves it step by step, verifies the result, and retries or fails closed when it isn't sure. This page explains what “tested on 100,000+ problems” actually means, how the engine is measured, and where it can still go wrong.

The claim

“Tested on 100,000+ problems”

v1.0

Means

100,000+ curated problems, tested against the full pipeline

Drawn from

A corpus of 1.5M+ real math problems

Tested with

Close to 1 million solve-and-verify passes

Measured for

Accuracy, re-checked as the engine changes

Honest by default

HexMath shows the work and tells you when a problem is unclear — never a confident guess.

What the number means

From 1.5 million problems to one engine you can trust.

“Tested on 100,000+ problems” isn't a marketing round-number. It describes how much real math went into building and hardening the engine.

1.5M+

real math problems

HexMath's solving engine was built from a corpus of more than 1.5 million real math questions from public sources — the raw material, before any curation.

100,000+

curated test problems

From that corpus we curated more than 100,000 problems with worked, accepted solutions across algebra, calculus, statistics, and more. This curated set is what “tested on 100,000+ problems” refers to.

~1M

solve-and-verify test runs

We ran close to a million solve-and-verify passes through the full pipeline against those curated problems — measuring accuracy and tightening the engine on real failures, not guesses.

Every one of those curated problems has a known, accepted solution, so each test run is scored automatically — the engine's answer is checked against the right answer, at scale, every time the engine changes. A fix in one area can't quietly break another without it showing up.

How a solve is checked

Every answer is verified, not just generated.

The same loop runs on every problem. That verification step is what separates an agentic AI solving system from a general-purpose AI wrapper that hands back its first guess.

  1. 1

    Reads the problem

    Pulls the math from a photo, paste, or typed input. If the image is unclear or holds more than one problem, HexMath says so instead of guessing.

  2. 2

    Solves it step by step

    Works the problem and returns a full, labeled derivation — every line of reasoning, not just a final number.

  3. 3

    Verifies the result

    Runs a separate verification pass on the answer to catch slips before anything reaches you.

  4. 4

    Retries or fails closed

    If a check fails or confidence is low, HexMath reworks the problem — and when it still can't be sure, it tells you instead of handing back a confident guess.

Where it can still go wrong

No solver is perfect. Here's where ours isn't.

Testing reduces mistakes; it doesn't eliminate them. Being clear about the failure modes is part of being trustworthy.

Messy or handwritten photos

Faint, skewed, or cluttered images can be misread. A tight crop and good light help; when in doubt, HexMath asks for a clearer photo.

Ambiguous notation

Notation that could mean two things (loose parentheses, unclear exponents) can be read the wrong way. Typing or cleaning up the input fixes most of these.

Long multi-step word problems

Problems that hide several steps inside prose are harder, and a wrong reading early on can carry through. The shown work makes those easier to catch.

Edge-case and advanced topics

Unusual phrasing or advanced material is where errors are most likely. This is exactly why every answer shows its steps so you can check them.

HexMath is a study aid and can make mistakes. Double-check important work, especially on graded or high-stakes problems.

See the work for yourself.

The best proof is a problem you already know the answer to. Try one and check every step.

Questions about how HexMath is tested? Contact support or read the Terms.