Methodology

MathToModels is a browser-only learning system. No accounts, no cookies, no server-side profiles. This page documents the statistical models behind the learner dashboard.

← Platform Home

1 · Ability estimate (1PL IRT / Rasch model)

Each question has a difficulty b. Your ability θ per subject is estimated by maximum likelihood under the Rasch model:

P(correct | θ, b) = 1 / (1 + exp(−(θ − b)))

We solve the score equation Σ (cᵢ − Pᵢ) = 0 by Newton–Raphson steps. The Fisher information is I(θ) = Σ Pᵢ (1 − Pᵢ), and the standard error is SE(θ) = 1 / √I(θ). Reported intervals are θ ± 1.96 · SE.

Assumptions: unidimensionality within a subject, local independence, no guessing. Item difficulties are expert-set rather than empirically calibrated.

2 · Wilson confidence intervals

For accuracy we report a 95% Wilson score interval rather than the normal approximation, which is unreliable for small n or extreme proportions.

p̂ ± z·√(p̂(1−p̂)/n + z²/(4n²)) / (1 + z²/n)

3 · Spaced repetition (SM-2)

Each question is a card with an ease factor EF, an interval, and a due date. Correct answers (quality ≥ 3) grow the interval; wrong answers reset it to 1 day and increment lapses. EF is updated by the SM-2 rule.

4 · Calibration and Brier score

Before each answer we ask for a confidence level in {25%, 50%, 75%, 95%}. Calibration bins compare stated confidence with observed accuracy. The Brier score is mean((confidence − outcome)²) — lower is better.

5 · Opt-in community cohort

When you tick the opt-in checkbox on the learner dashboard, your ability estimate θ is sent to a small anonymous backend (a Netlify Function). The request contains only: subject name, θ, correct count, and total count. No identifiers, no cookies, no cross-site tracking.

The backend stores an anonymous array of θ values per subject in Netlify Blobs. A rolling 1-hour key derived from your IP is used only for rate limiting, and it expires automatically. You can opt out at any time from the same checkbox, and any new submission stops immediately.

Percentile shown on the dashboard is computed against this real pool when n ≥ 30; below that threshold a deterministic reference distribution is used and the badge displays a placeholder.

6 · LLM-graded teach-back (bring your own key)

If you set a provider key in the Settings panel, teach-back grading is done by your chosen provider (OpenAI, Anthropic, Groq). The prompt includes only the topic title, section, overview, key terms, and the text you wrote. Nothing is stored on our side.

Your key lives in this browser's localStorage and is sent only to the provider you choose. If no key is set, we fall back to a purely local keyword-coverage heuristic.

7 · Code runner

The Python runner uses Pyodide (a WebAssembly build of CPython). The SQL runner uses sql.js (SQLite compiled to WebAssembly). Both run entirely in your browser. No code, output, or data is uploaded anywhere.

8 · Limitations

9 · What we do not do

← Back to the platform