How Sage IQ measures ability
The short version: nine subtests measure six broad abilities from the Cattell–Horn–Carroll model. Items are generated procedurally and administered adaptively; scores are estimated with item response theory and reported with error bars; the overall estimate is a weighted composite standardised against published anchors and, once your group is large enough, against your group.
What is measured
| Ability | Subtests | Lineage |
|---|---|---|
| Fluid reasoning | Matrix reasoning (16 adaptive items) | Raven-style matrices; Carpenter, Just & Shell rule taxonomy |
| Sequential reasoning | Number series (12, typed answer); letter series (10) | Thurstone series completion; Holzman et al. |
| Verbal reasoning | Analogies, odd-one-out, synonyms (14) | Classic crystallised-ability formats (English) |
| Working memory | Digit span forward & backward; spatial span | Wechsler digit span; Corsi block-tapping (Kessels et al.) |
| Processing speed | Symbol coding (90 s); symbol search (75 s) | Symbol Digit Modalities Test; Wechsler coding & symbol search |
| Spatial ability | 2-D mental rotation with mirror-image foils (14) | Shepard & Metzler; Cooper |
How scores are produced
- Adaptive items. Each adaptive subtest starts with two easier warm-up items, then picks the difficulty that is most informative at your current estimate. Your ability is estimated with the expected-a-posteriori method under a three-parameter logistic model, which stays stable even after runs of all-correct or all-wrong answers.
- Timed tasks. Span and speed tasks are scored with the standard rules of their paper-and-pencil ancestors (two trials per length with a discontinue rule; correct responses within the time limit; correct minus incorrect for symbol search).
- Standardisation. Task scores are converted to a common scale using published adult reference values (labelled "provisional anchors"). Once at least 12 consenting members have completed a battery, within-group norms replace the anchors and group percentiles appear.
- Composite. The general ability estimate is a weighted mean of the six domain scores (weights reflect how strongly each task type loads on general intelligence in the literature), rescaled so that the composite has the same spread as its components under a typical inter-ability correlation of 0.5, and expressed on a mean-100, SD-15 scale.
- Error bars. Every number carries a standard error from the measurement model. Intervals are 90%. Treat differences smaller than the intervals as noise.
What the numbers are not
- Not a clinical IQ. Professional tests are administered one-to-one by trained examiners and normed on thousands of people stratified by age, education and region. This is neither.
- Not age-corrected. Speed and memory decline with age from the late twenties; vocabulary keeps growing into the sixties. A 55-year-old and a 22-year-old with the same speed score are not equally fast for their age.
- Not cheat-proof. Any unsupervised test can be gamed. The value here is honest self-measurement among friends.
- Not stable under repetition. A second attempt typically gains 3–5 points from familiarity alone, and the gain persists for months. Space repeat tests widely and read repeat scores with that in mind.
Full technical documentation
Show the complete methodology document
Quotient: scientific background and measurement model
Battery version 1.0. Describes what the code in apps/assessment/engine/ and apps/assessment/services.py actually does and why. Where the implementation and the project's research notes disagree, the disagreement is listed under "Discrepancies to review" rather than glossed over.
1. Purpose and honest positioning
Quotient is a self-hosted Django app that lets a small private group of friends take a browser-based battery of cognitive tasks and see a profile of six broad abilities plus a weighted composite. It borrows the machinery of professional intelligence batteries (a hierarchical ability model, item response theory, adaptive item selection, standardised scores with error bands) but it is not one of them:
- Not a clinical or diagnostic instrument. Nothing has been validated against an external criterion; results must not inform educational, occupational, medical or legal decisions.
- Unsupervised. Sessions run on the user's own hardware with no proctor. Unproctored ability scores run about 0.20 SD higher than proctored ones, mostly on look-up-able knowledge items (Steger, Schroeders & Gnambs 2020).
- No representative norming sample. Published IQ tests are normed on thousands of people stratified by age, sex, education and region (Canivez 2010; AERA, APA & NCME 2014). A friend group cannot produce norms, so scores are reported against provisional anchors (published convenience-sample means) and, once
NORM_POOL_MIN_N = 12consenting users have completed a battery version, against the pool of those users. The score panel always states which was used. - Estimates with error bars. Every subtest, domain and composite score carries a standard error and is shown with a 90 % interval. The composite is labelled "General ability estimate", never "IQ".
- Generated, not calibrated, items. Difficulty is predicted from item design; empirical re-calibration exists but needs hundreds of responses per item family before it changes anything.
What the app can honestly claim: a repeatable, moderately reliable ordering of a small group on six broad abilities, on a scale that approximates adult reference values, with intervals wide enough to say so.
2. Theoretical model
2.1 g and the CHC broad abilities
All cognitive tests correlate positively (Spearman's positive manifold), with a general factor g above a set of broad abilities (Carroll 1993). The Cattell-Horn-Carroll model names the ones IQ batteries measure: fluid reasoning (Gf), comprehension-knowledge (Gc), working memory (Gwm), processing speed (Gs) and visual-spatial processing (Gv) (Schneider & McGrew 2018). A mixed battery predicts g far better than any single task: a four-subtest Wechsler short form reaches g validity .93, whereas Raven's matrices share only about half their variance with g (Gignac 2015). Quotient therefore measures several domains and combines them.
2.2 The six reported domains
battery.py groups nine subtests into six domains:
| Domain | CHC | Subtests | Weight |
|---|---|---|---|
| Fluid reasoning | Gf | Matrix reasoning | 1.0 |
| Sequential reasoning | Gf / Gq | Number series, Letter series | 0.9 |
| Verbal reasoning | Gc | Verbal (analogy, odd-one-out, synonym) | 0.8 |
| Working memory | Gwm | Digit span (forward + backward), Spatial span | 0.8 |
| Processing speed | Gs | Symbol coding, Symbol search | 0.6 |
| Spatial ability | Gv | Mental rotation | 0.7 |
The "full" battery runs all nine in the order matrices, digit span, number series, coding, verbal, spatial span, letter series, symbol search, rotation (about 45 min, alternating heavy and light tasks). The "quick" battery runs five shortened subtests (matrices 10 items, backward digit span only, number series 8 items, coding, rotation 8 items; about 18 min) and is labelled less precise.
2.3 Why these weights
The weights are not fitted (data-driven weights are unstable below several hundred people). They are a rounded, ordinal rendering of published g-loadings, adjusted for the app's versions of the tasks:
- WAIS-IV hierarchical g-loadings (Canivez & Watkins 2010): Digit Span .70, Vocabulary .69, Block Design .68, Matrix Reasoning .66, Coding .54, Symbol Search .52.
- ICAR (Condon & Revelle 2014; Young & Keith 2020): letter-number series and verbal reasoning about .8, matrices about .7, 3-D rotation .5-.6. Raven-type matrices load about .70 (Gignac 2015).
Fluid reasoning gets 1.0: matrices are the most g-saturated novel task and the least searchable. Sequential (0.9) reflects the high ICAR series loading. Verbal is held to 0.8 because the bank is hand-authored, English-only and search-vulnerable. Working memory is 0.8 rather than higher because the app uses simple spans; the WAIS-IV loading includes backward and sequencing, and latent working-memory capacity correlates .72 with Gf (Kane, Hambrick & Conway 2005) but only .48 with g at the true-score level in the largest meta-analysis (Ackerman, Beier & Boyle 2005). Speed gets the lowest weight (0.6), matching its .5 loadings and device dependence. Spatial gets 0.7: the lowest ICAR loading, but language-free and cheat-resistant. The research notes call equal weights the defensible default and g-loading weights a principled alternative; the app chose the latter. Changing DOMAINS[...]["weight"] changes only the composite.
3. The subtests
Adaptive subtests use seven difficulty levels. level_to_b maps level L linearly from b_min (level 1) to b_max (level 7); ranges and fixed parameters come from adaptive.GENERATORS and the generators:
| Subtest | a | c | b range | b at levels 1 … 7 |
|---|---|---|---|---|
| Matrices | 1.3 | 1/8 | −2.4 … 2.6 | −2.40, −1.57, −0.73, 0.10, 0.93, 1.77, 2.60 |
| Number series | 1.1 | 0 | −2.2 … 2.4 | −2.20, −1.43, −0.67, 0.10, 0.87, 1.63, 2.40 |
| Letter series | 1.1 | 1/6 | −2.4 … 2.2 | −2.40, −1.63, −0.87, −0.10, 0.67, 1.43, 2.20 |
| Verbal | 1.1 | 1/5 | −2.4 … 2.4 | −2.40, −1.60, −0.80, 0.00, 0.80, 1.60, 2.40 |
| Rotation | 1.0 | 1/5 | −2.2 … 2.2 | −2.20, −1.47, −0.73, 0.00, 0.73, 1.47, 2.20 |
Discriminations are a priori, in the range of calibrated figural banks (MaRs-IB mean 1.30; Zorowitz, Chierchia, Blakemore & Daw 2023). c = 1/k for k-option multiple choice (the MaRs-IB convention) and 0 for typed answers. Item spec is pure rendering data; the answer key stays in the server's ItemResponse row.
3.1 Matrix reasoning (Gf)
Lineage. A 3 × 3 grid with the bottom-right cell missing and eight options, after Raven's Progressive Matrices. Rules follow Carpenter, Just & Shell's (1990) analysis of the Advanced Progressive Matrices (constant in a row, quantitative progression, distribution of three values, figure addition/subtraction); generation follows the Sandia matrices (Matzen et al. 2010) and the radicals/incidentals logic of Embretson (1998) and Gierl & Lai (2012).
Generation (matrices.py). Each cell's main element has five attributes (shape 8 values, fill 5, count 1-5, size 3, rotation 0/45/90/135°). Each rule governs one attribute across rows: constant, progression (ordered attributes only, step ±1) or distribution (a Latin-square arrangement so rows and columns each contain all three values). An optional second layer of up to four corner markers follows distribution or XOR (third cell = symmetric difference of the first two). LEVEL_RECIPES fixes rule count, kinds, layers and XOR per level: level 1 is one constant/progression rule; level 4 introduces three rules or XOR; level 7 is four progression/distribution rules on two layers, or three rules plus XOR. Seven distractors perturb one (65 %) or two attributes of the answer, preferring rule-governed ones, so each breaks at least one rule; this resembles MaRs-IB's "paired-difference" strategy, about 20 points easier than "minimal-difference" foils (Zorowitz et al. 2023).
Administration. 16 items (10 quick), 75 s each, keys 1-8. A difficulty index (constant 0.6, progression 1.0, distribution 1.35, +0.9 per extra layer, +1.6 for XOR) nudges b within the level: b = level_to_b(level) + 0.15 × (index − typical index of the level), bounded to ±0.35 so levels stay ordered. Family: matrix.L{level}.{xor|rules}{n}.{1L|2L}.
3.2 Number series and letter series
Lineage. Series completion involves detecting the period, describing the pattern and extrapolating (Simon & Kotovsky 1963; Kotovsky & Simon 1973); over 70 % of difficulty variance is explained by working-memory load, hierarchical relations and arithmetic demand (Holzman, Pellegrino & Glaser 1983). Generator studies confirm that second operations, progressive coefficients and interleaving add most difficulty (Sun et al. 2018; Sun, Liu & Luo 2019). Letter series follow the tradition of Thurstone's PMA letter-series task and the Simon & Kotovsky pattern-induction analysis.
Number series (series.py). Free numeric entry ("4", "4.0" and "4,0" are equivalent), so c = 0 and no distractors are needed, as the research notes recommend. Families by level: arithmetic, geometric (1-2); increasing differences, alternating operators, squares with offset (3); quadratic, interleaved, Fibonacci-type (4); affine recurrences xₙ₊₁ = m·xₙ + k (5-7). At levels ≥ 5 only the last five terms are shown (six for interleaved). Stems are rejected if fewer than five terms remain, any term exceeds 5000, or terms repeat. 12 items (8 quick), 60 s, a = 1.1. Family: nseries.L{level}.{kind}.
Letter series. Six options (c = 1/6), 10 items, 45 s. Generators: constant step, backwards, increasing step, letter pairs and mirror pairs, interleaved series, growing blocks (A B B C C C …), letter-number pairs. Distractors are alphabet neighbours of the answer. Family: lseries.L{level}.{kind}.
3.3 Verbal reasoning (Gc)
Lineage. Analogies follow the encoding-inference-mapping-application components of Sternberg (1977); difficulty rises with the abstractness of the A:B relation (Bejar, Chaffin & Embretson 1991; Duran 1987). Such items load on both vocabulary and reasoning (Ullstadius, Carlstedt & Gustafsson 2008), hence the label "Verbal reasoning" and the note that it depends on vocabulary.
Bank and selection (verbal_bank.py, verbal.py). 84 hand-authored English items, 12 per level: 42 analogies, 22 odd-one-out, 20 synonyms, five options each (c = 1/5); levels are judged a priori (level 1 big/large; level 7 pellucid, obsequious). Options are shuffled per presentation. The engine requests a level; pick_verbal_item draws randomly among items at that level not used in this run and not seen by this user in any earlier session (services._seen_verbal_ids), widening the level window to ±1, ±2, ±3, ±6 before falling back to seen items (flagged seen_before). 14 items, 40 s, a = 1.1. Family: verbal.{item_id}, so calibration is per item. Drawing 14 of 84 items near one ability level exhausts a person's unseen pool within a few sessions.
3.4 Digit span and spatial span (Gwm)
Lineage. Digit span follows Wechsler (2008): one element per second, two trials per length, discontinue after both trials at a length fail, span = longest length with at least one correct trial, plus total correct. Spatial span follows the Corsi standardisation of Kessels et al. (2000): nine blocks in the classic irregular layout, from length 2, same stop rule.
Implementation (span.py). The server generates all sequences from the run's seed; the browser presents them and returns reproductions; the server scores and enforces the discontinue rule. Forward digits 3-9, backward 2-8, spatial 2-9, TRIALS_PER_LENGTH = 2, onset_ms = 1000 (one element per second) with a blank_ms = 250 ms (digits) or 300 ms (blocks) gap inside that second, i.e. a digit is visible for 750 ms. Digit lists (1-9) contain no repeated digit and no three-term ±1 runs; block sequences avoid the last two blocks. Backward trials are scored against the reversed list. Anchors: Woods et al. (2011) for digits, Kessels et al. (2000) and eCorsi (Brunetti, Del Gatto & Delogu 2014) for blocks (section 5.1).
3.5 Symbol coding and symbol search (Gs)
Lineage. Coding is a keyboard Symbol Digit Modalities Test (Smith 1982; 90 s) / Wechsler Coding: a key pairs nine symbols with digits 1-9. Symbol search follows Wechsler Symbol Search: two targets, a row of five, yes/no, scored correct minus incorrect (Wechsler 2008). Substitution tests are sensitive but non-specific (Jaeger 2018), one reason for the low weight.
Implementation (speed.py). Nine of twelve glyphs are sampled per session, giving a fresh alternate form each time. Coding: 240-item stream without immediate repeats, 6 practice items, CODING_SECONDS = 90, score = correct, errors separate, median RT recorded. Symbol search: 160 trials at 50 % target-present, 4 practice, SEARCH_SECONDS = 75, keys F/J or buttons, score = max(0, correct − incorrect). Timing uses performance.now(); responses under MIN_RT_MS = 150 count as errors. Coarse device data (input type, viewport class) is stored because touch input and browser latency differ (Anwyl-Irvine et al. 2021; Pronk et al. 2020).
3.6 Mental rotation (Gv)
Lineage. Response time to rotated block figures grows linearly with angular disparity (Shepard & Metzler 1971), likewise for 2-D polygons (Cooper 1975). Foils must be mirror images: structurally different foils can be rejected without rotating (Voyer & Hou 2006; Jost & Jansen 2024).
Implementation (rotation.py). The target is a random polyomino (4 cells at level 1 up to 9 at level 7) that is chiral without rotational symmetry, so its mirror image is never a rotation of it. Five options: the target rotated by a level-specific angle (90/270° at level 1; 135/225/315° at level 7) and four mirror images at other angles. 14 items (8 quick), 40 s, a = 1.0, c = 1/5, accuracy only. Family: rotation.L{level}.c{n_cells}. Sex differences on 3-D rotation are large (d ≈ 0.67; Voyer, Voyer & Bryden 1995) and smaller with mirror-only foils; the app makes no adjustment.
4. Measurement model
4.1 Item model
Adaptive subtests use the three-parameter logistic model (irt.py): P(correct | θ) = c + (1 − c)/(1 + exp(−a(θ − b))). Item information is a²·(q/p)·((p − c)/(1 − c))², test information is the sum, SE = 1/√I(θ). The maximum information a perfectly targeted item can add is about 0.33 (matrices), 0.30 (number series), 0.22 (letter series), 0.21 (verbal) and 0.17 (rotation); guessing costs information, which is why the six- and eight-option formats matter.
4.2 EAP scoring
θ is the expected a posteriori estimate (Bock & Mislevy 1982) over THETA_GRID, −4.0 to +4.0 in steps of 0.05 (161 points), with a standard normal prior (PRIOR_SD = 1.0); the SE is the posterior SD. Before any response θ = 0, SE = 1. EAP is used because the tests are short: maximum likelihood is undefined after all-correct or all-wrong patterns, which occur routinely early on and can span a whole 8-item quick subtest, whereas EAP always returns finite values. The cost is shrinkage towards 0 with few items, which is why very short subtests cannot yield extreme scores.
4.3 Adaptive level selection
adaptive.next_item chooses the level of the next generated item:
- Warm-up. Items 1-2 are at
WARMUP_LEVELS = (2, 3)(b ≈ −1.6 and −0.8). - Maximum information.
best_difficulty_fortargets b = θ̂ for c = 0, otherwise b = θ̂ − ln((1 + √(1 + 8c))/2): 0.19 (matrices), 0.23 (letter series) or 0.27 logits (verbal, rotation) below θ̂, since guessing makes slightly easier items more informative. - Randomisation. Gaussian noise with
spread = 0.35logits is added before choosing the nearest level, so equal abilities do not see identical level sequences (same purpose as randomesque selection; Kingsbury & Zara 1989; Han 2018). - Anti-stall. After three consecutive presentations of one level, the next shifts ±1.
- Exposure. Verbal items are exposure-controlled (3.3); other generators create a fresh item every time.
Subtests are fixed-length (16/12/10/14/14) rather than SE-terminated, so that session length is predictable; the price is SEs of roughly 0.4-0.65 instead of a fixed 0.30. Timed-out items score incorrect and are counted in n_timeouts.
4.4 Precision and reliability
Reliability is approximated everywhere as 1 − SE² (population variance 1 under the prior), the empirical-reliability convention in the notes (SE 0.30 ↔ .90, 0.39 ↔ .85, 0.45 ↔ .80). Best-case SEs with perfectly targeted items:
| Subtest | Items (quick) | Best-case SE | Reliability (code formula) |
|---|---|---|---|
| Matrices | 16 (10) | 0.40 (0.48) | .86 (.81) |
| Number series | 12 (8) | 0.46 (0.54) | .83 (.77) |
| Letter series | 10 | 0.56 | .76 |
| Verbal | 14 | 0.50 | .80 |
| Rotation | 14 (8) | 0.54 (0.65) | .77 (.70) |
With seven discrete levels, jitter and fixed warm-up, achieved SEs run 10-20 % higher: about 0.45-0.6 in the full battery, up to 0.7 in the quick one. This matches short fixed forms in the literature (ICAR-16 α = .81; MaRs-IB 12-item forms ρ ≈ .71; Condon & Revelle 2014; Zorowitz et al. 2023), and precision collapses at the extremes.
5. Scoring and standardisation
5.1 Layer 1: provisional anchors
Adaptive subtests: z = θ̂ (clamped to ±3.5), SE = posterior SD. This scale is anchored only by the level-to-b mapping: θ = 0 means 50 % success on a level-4 item after removing guessing. Task subtests use ANCHORS in scoring.py:
| Task | Mean | SD | Assumed SE (z) | Stated source |
|---|---|---|---|---|
| Digit forward (span) | 6.5 | 1.0 | 0.55 | Woods et al. 2011: 6.52 (1.00), n = 763 |
| Digit backward (span) | 4.9 | 1.05 | 0.55 | Woods et al. 2011: 4.91 (1.06) |
| Spatial span | 6.0 | 1.1 | 0.6 | Kessels et al. 2000: 6.2 (1.3); eCorsi 18-30 y: 6.11 (0.80) |
| Coding (correct, 90 s) | 62.0 | 12.0 | 0.45 | "keyboard symbol-digit"; nearest published: written SDMT < 30 y 58.2 (9.1) |
| Symbol search (correct − incorrect, 75 s) | 34.0 | 8.0 | 0.5 | none in the research notes |
Digit span's z is the mean of forward and backward z, SE = √(0.55² + 0.55²)/2 ≈ 0.39. Coding is rescaled to 90 s if a shorter limit was used. The anchor SEs are assumed reliabilities of a single span or speed score (literature retest values: forward mean span .67, backward .84, Woods et al. 2011; PEBL digit span .63, Piper et al. 2015; Wechsler speed subtests .74-.90), corresponding to about .74-.83 under the code's formula. Their mismatches with the app's administration are listed in section 8.
5.2 Domain and composite scores
A domain z is the unweighted mean of its subtests' z, SE = √Σ(SEᵢ/n)². The composite is the weighted mean of domain z's divided by the SD that mean would have if domains were unit-variance and correlated r = ASSUMED_DOMAIN_CORRELATION = 0.5 (typical inter-domain correlations .4-.6; ICAR-16 factor correlations .41-.70):
SD_composite = √(Σwᵢ² + 2r·Σᵢ<ⱼ wᵢwⱼ) / Σwᵢ
For the full battery this is √(3.94 + 9.55)/4.8 ≈ 0.765 (quick: ≈ 0.777). Someone averaging z = 1 on every domain lands at composite z ≈ 1.31; dividing by the number of domains instead would shrink everyone towards the mean (Tellegen & Briggs 1967). The composite SE is the weighted propagation of domain SEs divided by the same factor.
Scale. Standard score = 100 + 15z; percentile = Φ(z); 90 % interval = z ± Z90 = 1.645·SE; z clamped to ±3.5 (47.5-152.5). Bands: ≥ 130 Well above average; 120-129 Above average; 110-119 High average; 90-109 Average; 80-89 Low average; 70-79 Below average; < 70 Well below average. The composite is labelled "General ability estimate" because 100/15 is only meaningful relative to a defined population, and the reference here is a patchwork of convenience samples and a-priori item scales; "IQ" would imply a norm that does not exist.
Typical interval. With achieved SEs of 0.45-0.6 on adaptive subtests and assumed 0.45-0.6 on anchored ones, domain SEs are about 0.34-0.55 and the composite SE after the 0.765 division is about 0.23-0.30 (3.5-4.5 points). The 90 % interval on the General ability estimate is therefore about ±6-7 points for the full battery (95 %: ±7-9) and about ±8 for the quick battery; the code's composite reliability is typically .90-.95. Single domains are ±9-15 points at 90 %. These intervals contain measurement error only, not the anchor uncertainty of section 8, which for some subtests is of similar size.
5.3 Layer 2: pool norms
Once NORM_POOL_MIN_N = 12 consenting users (research_optin) have completed a battery version, services.refresh_pool_norms takes the anchor-based domain z's and raw composite of each user's first completed session (repeats excluded as practice-inflated; anonymised records may be added) and stores their mean and SD per domain (≥ 3 values, SD > 0.05) and for the composite. Scores are then re-standardised ((z − mean)/SD, SE/SD) with source = "pool", and a within-group percentile (midpoint convention) is shown, also only at n ≥ 12. Norms are recomputed after every completed session and applied at display time, so earlier results move as the pool grows. Pool norms remove the anchor mismatch but only rank people within this group: with n = 20 the median's 95 % CI spans roughly the 28th-72nd percentile, and a friend group is range-restricted.
5.4 Layer 3: empirical family calibration
calibration.py re-estimates each item family's b by alternating EAP person estimates and Newton steps on the 3PL log-likelihood for 6 iterations; a and c stay fixed. A family is updated only with at least MIN_N_FOR_UPDATE = 20 responses, shrunk towards its a-priori b by a pseudo-count prior PRIOR_PSEUDO_N = 15 (15 imaginary responses at the prior value) and clamped to ±4; proportion correct, point-biserial with θ and an updated flag are reported. Parameters an admin activates (ItemParameter.active) override the generator's a and b at family level, with a level-wide fallback ({subtest}.L{level}), via services._param_lookup (cached 60 s). This is family-level calibration in the sense of Gierl & Lai (2012). Rasch difficulties stabilise with a few hundred examinees (Linacre 1994, per the notes), so with dozens of users this layer moves only the most-used families, slightly.
6. Validity safeguards
scoring.validity_flags and services.finalize_session annotate rather than alter (as Meade & Craig 2012 recommend):
| Flag | Rule | Level |
|---|---|---|
rushed |
> 30 % of an adaptive subtest's responses faster than a per-task threshold (matrices and number series 3 s, letter series and verbal 2 s, rotation 1.5 s; roughly 10 % of typical solution times) | warn |
timeouts |
> 40 % of its items timed out | warn |
speed_errors |
coding errors / attempted > 30 % | warn |
zero_span |
any span part with span 0 | error |
retest_soon |
previous completed session < 90 days ago | warn |
retest_gain |
any repeat administration | info |
interrupted |
session resumed after interruption | info |
The rapid-response rule is a simplified response-time-effort screen (Wise & Kong 2005; Wise 2017). Server-side, speed responses under 150 ms are discounted, duplicate submissions ignored, items served one at a time so the client never holds the key, and unfinished sessions deleted after STALE_SESSION_DAYS = 14.
Practice effects. The retest_gain note quotes 0.2-0.3 SD (3-5 points). Scharfen, Peters & Holling (2018; 122 studies, N = 153,185) found 0.33 SD from first to second session, 0.23 with alternate forms and 0.37 with identical forms, a plateau after the third session, and no meaningful decay with interval (−0.0008 SD per week). Hausknecht et al. (2007) and Calamia, Markon & Tranel (2012) agree; Estevis, Basso & Combs (2012) found WAIS-IV gains of about 7 points identical at 3 and 6 months. Generated items make most Quotient subtests alternate forms, hence the lower figure; the verbal bank is the exception. The 90-day threshold is a familiarity heuristic, not a wash-out interval, because none exists.
Unsupervised testing. Web and lab data agree in means, variances and reliabilities (Germine et al. 2012), and unproctored ICAR items retain validity (Condon & Revelle 2014); the 0.20 SD inflation (Steger et al. 2020) concentrates on searchable items, and countermeasures did not help. The verbal subtest is the exposed one; 40 s per item makes searching costly, not impossible.
Device effects. Browser RT overhead is about 70-90 ms on keyboards and 57-70 ms on phones, systematic rather than random (Anwyl-Irvine et al. 2021; Pronk et al. 2020; Bridges et al. 2020), and speed scores vary by device (Passell et al. 2021). This touches the two speed tasks only. Input type and viewport class are stored per session so norms could later be split by device, but are not yet used.
7. Known limitations and roadmap
Limitations
- No age norms. Abilities peak at different ages (Hartshorne & Germine 2015); anchors are adult (≈ 18-50) values and pool norms mix ages.
- English-only verbal bank of 84 items with judged levels. Items cannot be translated and keep their difficulty (ITC 2017); another language needs its own bank and norms.
- Anchors approximate paper-and-pencil or auditory norms. Visual keyboard digit span runs 0.5-1 digit below auditory; keyboard coding scores above written SDMT; symbol search has no documented anchor; fluid anchors drift 0.3-0.4 points per year (Pietschnig & Voracek 2015).
- The adaptive scale is a priori. For matrices, series, verbal and rotation, z = 0 is defined by the generator's level-to-b mapping until pool norms give it an empirical centre.
- Cheating is possible (look-up, a helper, pausing by closing the tab); flags catch only rushing, timeouts and error rates.
- Small pools and short subtests. Twelve users give very wide percentile uncertainty; 8-16 items give subtest reliabilities of .7-.85; families need hundreds of responses before calibration matters.
- No person-fit statistics (Guttman errors, lz) and no normative rapid-guess threshold.
- Rotation is 2-D and accuracy-only; its g-loading and rotation rate differ from 3-D figures.
Possible upgrades
- Adopt the MaRs-IB (480 calibrated items, CC BY 4.0, N = 1,501; Chierchia et al. 2019; Zorowitz et al. 2023) to give fluid reasoning a real external anchor.
- Import Sandia matrix generation (BSD-3-Clause; Matzen et al. 2010) and the brief forms and CAT keys of Harris et al. (2020).
- Use the 384 free Shepard-Metzler stimuli of Ganis & Kievit (2015) for 3-D rotation with RT-by-angle scoring.
- Age- and device-stratified pool norms once the pool allows; a separately normed German verbal bank (ITC 2017).
- SE-based stopping (SE ≤ 0.30 or a 25-30-item cap, floor 10) and content balancing across rule families.
- ICAR items are not an option: their licence is academic-use-only (Doebler et al. 2026); only their design principles may be followed.
7a. Interruptions, resumption and time limits
Every answer is saved as it is given, so a session can be resumed after a reload, a closed tab or a server restart. Three rules keep resumption from becoming a loophole:
- Adaptive items keep their clock. The server records when an item was presented and computes the remaining time on resume; an answer arriving more than 3 s after the limit is scored as timed out whatever the client claims, and the stored reaction time can never exceed the server-observed elapsed time. An item that expired unanswered (tab closed, outage) is discarded and replaced by a fresh item with a full clock: the person cannot answer the item they already saw, and an outage does not cost them a wrong answer. Two or more replacements in a subtest are flagged on the report.
- Timed tasks restart with new stimuli. A span, coding or search task that is interrupted before its answers were saved starts again with a newly generated key and sequences (re-exposure would inflate the score); the restart is recorded and shown as an information flag.
- Saved answers are never lost to the network. Task answers are parked in the browser until the server confirms the save; submissions are idempotent, so a retry after a lost reply cannot double-record.
7b. Languages
The battery runs in English or German; the language is fixed per session and stored with it. Only the
verbal subtest is language-bound: each language has its own hand-authored bank of 84 items (families
verbal.v### and verbal.de.d###), calibrated separately, with the same a-priori level → difficulty
mapping. All other subtests use language-free stimuli and translated instructions. One norming pool is
shared across languages; verbal scores are therefore comparable across languages only through the
a-priori scale until each language has enough responses for empirical calibration. Members should test
in the language they know best.
8. Known limitations and open items
The following points were identified in a review of the implementation against the research notes. Items marked resolved were fixed in the code; the rest are documented limitations.
- Resolved. The matrix difficulty index now adjusts b within a level (see 3.1).
- Carpenter's "distribution of two values" rule is not implemented; the marker layer covers XOR and distribution of three only.
- Resolved. Span presentation is one element per second (750 ms on, 250 ms blank for digits).
- Resolved. Digit-span anchors use the classic maximum-length means from Woods et al. (2011): forward 6.35 (1.15), backward 4.61 (1.22). They still come from an auditory, computerised administration, so treat them as provisional.
- The coding (62 ± 12 in 90 s) and symbol-search (34 ± 8 in 75 s) anchors are design estimates, not published norms: no normative data exist for this exact keyboard, one-symbol-at-a-time format (written SDMT under 30 is 58.2 ± 9.1; keyboard responding is faster). They are labelled "estimate" in the code and are superseded by pool norms as soon as 12 consenting people have completed a battery.
- Resolved. Number-series stems are rejected when a simpler rule (constant difference, constant ratio, constant second difference, two interleaved arithmetic series) fits every shown term but predicts a different next number.
- Resolved. Digit lists contain no repeated digit.
- Fixed-length subtests instead of SE-based stopping (see 4).
- Resolved. One reliability formula (1 − SE²).
- Resolved. Two-subtest domain scores (sequential, memory, speed) are rescaled by the SD of a mean of correlated z-scores (r = 0.6 within a domain), so they keep unit variance like single-subtest domains.
- Partly resolved. Rapid-guessing flags use per-task thresholds; no person-fit statistic is computed.
- Rotation angle pools: 180° appears at levels 2, 4 and 5; oblique angles (45°, 135°) enter from level 3. Difficulty across levels is driven mainly by shape complexity (4 to 9 cells).
best_difficulty_forassumes a = 1 (difference of ≤ 0.05 logits for a = 1.3; negligible).- Resolved. Lineage wording for letter series.
- Resolved. Retest text quotes 0.23-0.33 SD.
- No age norms: speed and memory decline with age from the late twenties, vocabulary rises into the sixties; a pooled group mixes ages.
- English-only verbal bank of 84 items; exposure control avoids repeats across sessions but the bank is finite.
- Unsupervised: nothing prevents outside help; the value is honest self-measurement among friends.
9. References
From docs/research/psychometrics.md and docs/research/tasks-and-norms.md.
- Ackerman, P. L., Beier, M. E., & Boyle, M. O. (2005). Working memory and intelligence: The same or different constructs? Psychological Bulletin, 131, 30-60.
- AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.
- Anwyl-Irvine, A., et al. (2021). Realistic precision and accuracy of online experiment platforms, web browsers, and devices. Behavior Research Methods, 53, 1407-1425.
- Bejar, I. I., Chaffin, R., & Embretson, S. (1991). Cognitive and Psychometric Analysis of Analogical Problem Solving. Springer.
- Bridges, D., et al. (2020). The timing mega-study. PeerJ, 8, e9414.
- Brunetti, R., Del Gatto, C., & Delogu, F. (2014). eCorsi: implementation and testing of the Corsi block-tapping task for digital tablets. Frontiers in Psychology, 5, 939.
- Calamia, M., Markon, K., & Tranel, D. (2012). Scoring higher the second time around. The Clinical Neuropsychologist, 26, 543-570.
- Canivez, G. L. (2010). Review of the WAIS-IV. Mental Measurements Yearbook.
- Canivez, G. L., & Watkins, M. W. (2010). Investigation of the factor structure of the WAIS-IV. Psychological Assessment, 22, 827-836.
- Carpenter, P. A., Just, M. A., & Shell, P. (1990). What one intelligence test measures. Psychological Review, 97, 404-431.
- Carroll, J. B. (1993). Human Cognitive Abilities. Cambridge University Press.
- Chierchia, G., et al. (2019). The matrix reasoning item bank (MaRs-IB). Royal Society Open Science.
- Condon, D. M., & Revelle, W. (2014). The International Cognitive Ability Resource. Intelligence, 43, 52-64.
- Cooper, L. A. (1975). Mental rotation of random two-dimensional shapes. Cognitive Psychology, 7, 20-43.
- Doebler, P., Revelle, W., Condon, D., et al. (2026). Items, psychometric data and automatic item generators from the ICAR project. PsychArchives (Scientific Use License v1).
- Duran, R. P., et al. (1987). GRE verbal analogy items: examinee reasoning on items. ETS Research Report.
- Embretson, S. E. (1998). A cognitive design system approach to generating valid tests. Psychological Methods, 3, 380-396.
- Estevis, E., Basso, M. R., & Combs, D. (2012). Effects of practice on the WAIS-IV across 3- and 6-month intervals. The Clinical Neuropsychologist, 26, 239-254.
- Ganis, G., & Kievit, R. (2015). A new set of three-dimensional shapes for investigating mental rotation processes. Journal of Open Psychology Data, 3(1), e3.
- Germine, L., et al. (2012). Is the Web as good as the lab? Psychonomic Bulletin & Review, 19, 847-857.
- Gierl, M. J., & Lai, H. (2012). Using automated processes to generate test items. NCME ITEMS module.
- Gignac, G. E. (2015). Raven's is not a pure measure of general intelligence. Intelligence, 52, 71-79.
- Han, K. T. (2018). Components of the item selection algorithm in computerized adaptive testing.
- Harris, A. M., McMillan, J. T., Listyg, B., Matzen, L. E., & Carter, N. (2020). Measuring intelligence with the Sandia Matrices. Personnel Assessment and Decisions, 6(3).
- Hartshorne, J. K., & Germine, L. T. (2015). When does cognitive functioning peak? Psychological Science, 26(4), 433-443.
- Hausknecht, J. P., et al. (2007). Retesting in selection. Journal of Applied Psychology, 92, 373-385.
- Holzman, T. G., Pellegrino, J. W., & Glaser, R. (1983). Cognitive variables in series completion. Journal of Educational Psychology, 75, 603-618.
- International Test Commission (2017). ITC Guidelines for Translating and Adapting Tests (2nd ed.). International Journal of Testing, 18(2).
- Jaeger, J. (2018). Digit Symbol Substitution Test: the case for sensitivity over specificity. Journal of Clinical Psychopharmacology, 38(5), 513-519.
- Jost, L., & Jansen, P. (2024). The influence of the design of mental rotation trials on performance and possible differences between sexes. Quarterly Journal of Experimental Psychology.
- Kane, M. J., Hambrick, D. Z., & Conway, A. R. A. (2005). Working memory capacity and fluid intelligence are strongly related constructs. Psychological Bulletin, 131, 66-71.
- Kessels, R. P. C., et al. (2000). The Corsi Block-Tapping Task: standardization and normative data. Applied Neuropsychology, 7(4), 252-258.
- Kotovsky, K., & Simon, H. A. (1973). Empirical tests of a theory of human acquisition of concepts for sequential patterns. Cognitive Psychology, 4, 399-424.
- Matzen, L. E., et al. (2010). Recreating Raven's. Behavior Research Methods, 42, 525-541.
- Meade, A. W., & Craig, S. B. (2012). Identifying careless responses in survey data. Psychological Methods, 17, 437-455.
- Mueller, S. T., & Piper, B. J. (2014). The PEBL and PEBL Test Battery. Journal of Neuroscience Methods, 222, 250-259; Piper, B. J., et al. (2015), PEBL reliability study.
- Passell, E., et al. (2021). Cognitive test scores vary with choice of personal digital device. Behavior Research Methods, 53, 2544-2557.
- Pietschnig, J., & Voracek, M. (2015). One century of global IQ gains. Perspectives on Psychological Science, 10, 282-306.
- Pronk, T., et al. (2020). Mental chronometry in the pocket? Behavior Research Methods, 52, 1371-1382.
- Scharfen, J., Peters, J. M., & Holling, H. (2018). Retest effects in cognitive ability tests: A meta-analysis. Intelligence, 67, 44-66.
- Schneider, W. J., & McGrew, K. S. (2018). The Cattell-Horn-Carroll model of intelligence. In Contemporary Intellectual Assessment (4th ed.).
- Shepard, R. N., & Metzler, J. (1971). Mental rotation of three-dimensional objects. Science, 171, 701-703.
- Simon, H. A., & Kotovsky, K. (1963). Human acquisition of concepts for sequential patterns. Psychological Review, 70, 534-546.
- Steger, D., Schroeders, U., & Gnambs, T. (2020). A meta-analysis of test scores in proctored and unproctored ability assessments. European Journal of Psychological Assessment, 36, 174-184.
- Sun, L., et al. (2018). Evaluating an automated number series item generator using linear logistic test models. Journal of Intelligence, 6(2), 20.
- Sun, Y., Liu, ..., & Luo, F. (2019). Automatic generation of number series reasoning items of high difficulty. Frontiers in Psychology, 10, 884.
- Tellegen, A., & Briggs, P. F. (1967). Old wine in new skins: Grouping Wechsler subtests into new scales. Journal of Consulting Psychology, 31, 499-506.
- Ullstadius, E., Carlstedt, B., & Gustafsson, J.-E. (2008). The multidimensionality of verbal analogy items. International Journal of Testing.
- Voyer, D., Voyer, S., & Bryden, M. P. (1995). Magnitude of sex differences in spatial abilities: a meta-analysis. Psychological Bulletin, 117, 250-270.
- Voyer, D., & Hou, J. (2006). Type of items and the magnitude of gender differences on the Mental Rotations Test. Canadian Journal of Experimental Psychology, 60(2), 91-100.
- Wechsler, D. (2008). WAIS-IV Administration and Scoring Manual; Technical and Interpretive Manual. Pearson.
- Wise, S. L. (2017). Rapid-guessing behavior: Its identification, interpretation, and implications. Educational Measurement: Issues and Practice, 36(4), 52-61.
- Wise, S. L., & Kong, X. (2005). Response time effort: a new measure of examinee motivation in computer-based tests. Applied Measurement in Education, 18, 163-183.
- Woods, D. L., et al. (2011). Improving digit span assessment of short-term verbal memory. Journal of Clinical and Experimental Neuropsychology, 33(1), 101-111.
- Young, S. R., & Keith, T. Z. (2020). An examination of the convergent validity of the ICAR16 and WAIS-IV. Journal of Psychoeducational Assessment.
- Zorowitz, S., Chierchia, G., Blakemore, S.-J., & Daw, N. D. (2023). An item response theory analysis of the Matrix Reasoning Item Bank (MaRs-IB). Behavior Research Methods.
Cited in the research notes by author and year only: Spearman (1904), positive manifold; Bock & Mislevy (1982), EAP estimation; Kingsbury & Zara (1989), randomesque exposure control; Linacre (1994), Rasch sample sizes; Smith (1982), Symbol Digit Modalities Test; Sternberg (1977), componential analysis of analogies.