Psychometric Assistant
Neuropsychological calculators for clinical practice
Score embedded and stand-alone PVTs against published cut-offs, with base-rate-aware interpretation.
Estimate premorbid ability using ToPF and OPIE-4, compared against WAIS-IV and WMS-IV.
Build APA-formatted report tables with confidence-interval columns and premorbid comparison.
Evaluate test-retest change across five methods, from descriptive SDI to Crawford regression.
Chart scores entered elsewhere in the suite against classification and reliable-change bands.
Convert between and interpret effect sizes: Cohen's d, r, η², odds ratios and more.
Convert between standard, scaled, T and z scores and percentiles, with classification bands.
Score Converter
Full conversion table every score, scannable — the entered score's row is highlighted
A whole score appears only on the row where it actually falls — a blank cell means that metric has no whole score at that position. Columns finer than the anchored one show their nearest whole score on every row; z is exact to 2 dp. Percentile and classification conventions match the converter above. Ranges: Standard 40–160 · T 10–90 · Scaled 1–19 · z −4 to +4 in 0.25 steps.
AACN = American Academy of Clinical Neuropsychology · Ranges shown as Standard Score (SS)
Clinical Outcomes Table
SD mode: * ≥1 SD below, ** ≥1.5 SD, *** ≥2 SD. SEE mode: * below 90% CI, ** below 95% CI, *** below 99% CI lower bound.
| # | Subtest | Raw | Score | CI | Percentile | Classification |
|---|
Input Type
Effect Size Tools
Convert between effect-size metrics or derive them from group data, with a visual against the standard normal.
Group comparison at a target value
Use these published effect sizes to anchor your results. The values below span small to huge magnitudes so you can compare your finding against familiar clinical and epidemiological benchmarks.
| Heavy smokers (30+/day) vs never smokers, lung cancer (Pesch et al., 2012) | 2.60 |
| UK male vs female adult height (UK Biobank; Lui et al., 2021) | 2.04 |
| Smokers (any) vs never smokers, lung cancer (Pesch et al., 2012) | 1.75 |
| Cognitive therapy vs control for PTSD (Watts et al., 2013) | 1.63 |
| Former smokers vs never smokers, lung cancer (Pesch et al., 2012) | 1.10 |
| Exposure therapy vs control for PTSD (Watts et al., 2013) | 1.08 |
| EMDR vs control for PTSD (Watts et al., 2013) | 1.00 |
| Clozapine vs placebo for schizophrenia (Huhn et al., 2019 Lancet) | 0.89 |
| CBT vs control for depression (Cuijpers et al., 2023 World Psychiatry) | 0.80 |
| Methylphenidate vs placebo for ADHD, children (Storebø et al., 2023 Cochrane) | 0.75 |
| CBT vs placebo for anxiety disorders (Hofmann & Smits, 2008) | 0.70 |
| CBT for depression, low-risk-of-bias subset (Cuijpers et al., 2023) | 0.60 |
| Interpersonal Therapy for depression (Cuijpers et al., 2011) | 0.50 |
| Antidepressants vs placebo (Cipriani et al., 2018 Lancet) | 0.30 |
| CBT vs treatment-as-usual for chronic pain (Williams et al., 2020 Cochrane) | 0.20 |
| CBT vs active control for chronic pain (Williams et al., 2020 Cochrane) | 0.10 |
| Sugar on children's hyperactivity (Wolraich et al., 1995 JAMA) | 0.00 |
| No or negligible effect | 0.00 |
Performance Validity
Score embedded and stand-alone validity indicators against their published cut-offs, then weigh them together. No single index is a verdict. Results are cut-off comparisons only.
Formula & weighting table
Each raw score converts to a weighted score (Silverberg, Wertheimer & Fichtenberg, 2007, Table 2); EI = weighted Digit Span + weighted List Recognition (range 0–12). Higher EI = less credible performance. Some weights are unreachable for a given subtest (Digit Span never yields 1 or 4).
| Digit Span (raw) | List Recognition (raw) | Weighted score |
|---|---|---|
| 8–16 | 18–20 | 0 |
| — | 17 | 1 |
| 7 | 15–16 | 2 |
| 6 | 13–14 | 3 |
| — | 11–12 | 4 |
| 5 | 10 | 5 |
| 0–4 | 0–9 | 6 |
Caution. Over-flags in dementia (~48%) and severe impairment — corroborate before concluding.
Formula
ES = ( List Recognition − [ List Recall + Story Recall + Figure Recall ] ) + Digit Span, all raw scores (Novitski et al., 2012). Lower ES = more suspicious of poor effort. Cut-off: ES < 12, applicable only once the gating rule is met. The gate exists because in cognitively intact examinees free recall normally far exceeds the ceiling-limited recognition score, so an ungated ES over-flags; only 17% of the normal standardisation sample fall below the combined < 28 screen, and 15.1% meet it while also scoring ES < 12.
Caution. Unstable outside Alzheimer-type amnesia — confirm with a stand-alone PVT.
Scoring & worked example
RDS = (longest forward span passed on both trials) + (longest backward span passed on both trials) (Greiffenstein, Baker & Gola, 1994; cross-validated by Meyers & Volbrecht, 1998). The classic Forward + Backward variant is what the validation literature uses. Document which variant you administered. Floor rule: failing at least one trial each of 3 forward and 2 backward is recorded as RDS = 3. Worked example from the original paper: passes both trials of 3 forward but fails one 4-forward trial → forward = 3; passes both trials of 2 backward but fails one 3-backward trial → backward = 2; RDS = 3 + 2 = 5. Reference values: mean RDS ~8.8 in non-malingering TBI vs ~6.7 in probable malingerers.
Caution. Prefer ≤ 6 in genuine impairment — and even ≤ 6 over-flags in some groups.
Indices & base rates
The Digit Span age-corrected scaled score is entered as scored on the record form. The age correction is already applied by the Wechsler norms, so no separate age entry is needed here. ACSS ≤ 5 occurs in 3.8% of the WAIS-III standardisation sample and 3.4% of a combined clinical sample spanning TBI, alcohol abuse, Korsakoff's, temporal lobectomy and Alzheimer's disease (Iverson & Tulsky's suspicion guideline). Axelrod et al. found ≤ 7 the best single discriminator of probable malingering from moderate/severe TBI (sens. .75, spec. .69, with .77 on non-litigating mild-TBI cross-validation), and no genuine mild-TBI patient scored below 6. Iverson & Tulsky give three further suspicion indices, each evaluated against base rates: a Vocabulary − Digit Span difference ≥ 5 (7.1% of the standardisation sample, 2.8% of the combined clinical sample), a longest span forward ≤ 4 applicable only under age 55 (base rates 2.5–5.5% in the under-55 bands, climbing to 11% at 85–89, which is why the index is age-limited; the app reads the top-bar patient age and withholds it without one), and a longest span backward ≤ 2 (2.0–6.0% across all bands, 3.4% clinical). The longest spans here are the longest passed on either trial, as the standardisation data tabulate them, not the RDS both-trials span. Both papers derive from the WAIS-III. Document the edition administered.
Caution. Not a stand-alone measure, and it shares its subtest with RDS — the summary counts them as one indicator.
Scoring & administration
Combination score = free recall + (recognition correct − false positives), cut-off < 20; the classic free-recall cut-off is < 9 (Boone et al., 2002, Table 6 — the combination raises sensitivity from .47 to .71 at comparable specificity). Administration: expose the stimulus page for 10 seconds ("there are 15 different things so you will have to learn them very quickly"), remove it and have the examinee draw what they remember; then present the recognition page ("on this page are the 15 things I showed you as well as 15 items that were not on the page — circle the things you remember"). The stimulus and recognition pages are deliberately not reproduced here.
Caution. Specificity falls in dementia and intellectual disability — corroborate before concluding.
How the threshold is derived, and why it moves with age
Base rate by age (the default, and the only reading sourced to the CVLT-3). Table D.13 gives the percentage of the standardisation sample scoring each number of hits or fewer, by age band. There is no published cut score, so the flag is derived rather than stored: the rarest score whose published base rate is still at or below the selected per-test false-positive criterion. 10% is this page's own convention (PVT cut-offs are conventionally set so specificity is .90 or better), so nothing outside the manual enters the number itself.
The age band does real work here. 15 of 16 hits is the 8.6th percentile across all ages but the 26th at ages 80–90, so the derived threshold is ≤ 15 in every band except 80–90, where it is ≤ 14. Reading 15 as “one off perfect, fine” is a 3–6% event in a working-age adult and an unremarkable one in an 85-year-old. With no age entered the manual's own All ages column is used, and the readout says so.
The two published cut-offs are CVLT-II figures, and the page says so wherever one is in force. The trial is structurally identical across editions (16 List A targets, one distractor each, about 10 minutes after Yes/No Recognition), so the transfer is the ordinary one, but it is a transfer. ≤ 14 is the de facto standard: across 17 studies and 4,432 patients it gives sensitivity .50 at specificity .93, and no healthy control in the review scored at or below it (Schwartz et al., 2016). ≤ 15 treats a single error as a failure: in 104 adults with TBI it raised sensitivity to .56 while holding specificity at .92 across seven reference PVTs, identifying about 6% more invalid response sets (Erdodi et al., 2018). Note the two approaches converge: the derived threshold is already ≤ 15 for every patient under 80.
Failing rules invalidity in, but passing rules nothing out. Examinees who failed a reference PVT were about eight times more likely to fail Forced Choice, but roughly half of invalid response sets pass it, because a sophisticated simulator recognises how easy the task is (Schwartz et al., 2016).
Critical items are targets recalled, or recognised on Yes/No, earlier in the test but not chosen on the easier Forced Choice trial. Tables D.14 and D.15 give their base rates the other way round — that many or more — and their threshold is always derived from those tables, whichever basis the hits use. They come from the same administration as the hits, so all three count as one indicator.
Caution. Failure rates climb with genuine impairment; a near-perfect score rules nothing out.
Cut-offs & classification accuracy
Meta-analytic weighted-mean specificity/sensitivity for neurocognitive and psychiatric samples (Martin et al., 2020): Trial 2 / Retention < 45 — spec. .96–.98, sens. .45–.55; Trial 2 / Retention < 49 — spec. .91–.97, sens. .59–.70; Trial 1 < 42 — spec. .91, sens. .67–.69; Trial 1 < 41 — spec. .93, sens. .66. All are Martin et al.'s weighted-mean values for neurocognitive/psychiatric samples; an earlier aggregation of Trial 1 (Denning, 2012) reported .92/.77 averaged across cut-offs 34–44, which is why the meta-analysis examined each cut-off individually. Positive and negative predictive power are derived from these values and the selected base rate by Bayes' theorem, as in Martin et al. Tables 16–17.
Caution. Do not interpret traditional cut-offs in suspected or confirmed dementia.
Failing ≥ 2 independent PVTs supports probable invalidity, within Slick et al.'s (1999) criteria, which also require a substantial external incentive. Classification accuracy in Larrabee's combined sample (6 PVTs + 1 SVT):
| Threshold | Specificity | Sensitivity | Total correct |
|---|---|---|---|
| ≥ 2 of 7 failures | 88.9% | 97.6% | 92.6% |
| ≥ 3 of 7 failures | 96.3% | 87.8% | 92.6% |
| ≥ 4 of 7 failures | 100% | 63.4% | 84.2% |
Why ≥ 2, and when to demand more
Monte Carlo warnings that false positives climb steeply with more tests overestimate the problem: genuine patients' PVT scores are skewed and at-ceiling rather than multivariate-normal, so real failures cluster differently than random normal data predict. Two practical rules follow. Independence is required. Indices sharing subtests are correlated, so failing one raises the odds of failing another: the Effort Index and Effort Scale here are computed from the same RBANS subtests and count as one indicator, and RDS also loads on digit span. The false positives that do occur are the severely impaired. All six non-malingering cases failing ≥ 2 of 7 in Larrabee's sample had severe TBI with prolonged coma, complicated mild TBI, or stroke with a lesion on CT, and most failed only just inside the invalid range (RDS 7, WCST failure-to-maintain-set 2). So weigh the count against clinical and neurological coherence, and lean on high-specificity forced-choice measures, before reading two failures as probable invalidity in a genuinely impaired examinee (Larrabee, 2014).
References — Performance Validity
Axelrod, B. N., Fichtenberg, N. L., Millis, S. R., & Wertheimer, J. C. (2006). Detecting incomplete effort with Digit Span from the Wechsler Adult Intelligence Scale—Third Edition. The Clinical Neuropsychologist, 20(3), 513–523.
Boone, K. B., Salazar, X., Lu, P., Warner-Chacon, K., & Razani, J. (2002). The Rey 15-Item recognition trial: A technique to enhance sensitivity of the Rey 15-Item Memorization Test. Journal of Clinical and Experimental Neuropsychology, 24(5), 561–573.
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test, Third Edition (CVLT-3). Bloomington, MN: NCS Pearson. [Forced Choice Recognition base rates by age band, Appendix D Tables D.13-D.15; the manual publishes no cut-off for this trial.]
Denning, J. H. (2012). The efficiency and accuracy of the Test of Memory Malingering trial 1, errors on the first 10 items of the Test of Memory Malingering, and five embedded measures in predicting invalid test performance. Archives of Clinical Neuropsychology, 27(4), 417–432.
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test, Third Edition (CVLT-3). Bloomington, MN: NCS Pearson.
Erdodi, L. A., Abeare, C. A., Medoff, B., Seke, K. R., Sagar, S., & Kirsch, N. L. (2018). A single error is one too many: The Forced Choice Recognition trial of the CVLT-II as a measure of performance validity in adults with TBI. Archives of Clinical Neuropsychology, 33(7), 845–859.
Erdodi, L. A., Abeare, C. A., Medoff, B., Seke, K. R., Sagar, S., & Kirsch, N. L. (2018). A single error is one too many: The Forced Choice Recognition trial of the CVLT-II as a measure of performance validity in adults with TBI. Archives of Clinical Neuropsychology, 33(7), 845–859. [The CVLT-II ≤ 15 cut-off and its accuracy across seven reference PVTs.]
Greiffenstein, M. F., Baker, W. J., & Gola, T. (1994). Validation of malingered amnesia measures with a large clinical sample. Psychological Assessment, 6(3), 218–224.
Iverson, G. L., & Tulsky, D. S. (2003). Detecting malingering on the WAIS-III: Unusual Digit Span performance patterns in the normal population and in clinical groups. Archives of Clinical Neuropsychology, 18(1), 1–9.
Larrabee, G. J. (2014). False-positive rates associated with the use of multiple performance and symptom validity tests. Archives of Clinical Neuropsychology, 29(4), 364–373.
Martin, P. K., Schroeder, R. W., Olsen, D. H., Maloy, H., Boettcher, A., Ernst, N., & Okut, H. (2020). A systematic review and meta-analysis of the Test of Memory Malingering in adults: Two decades of deception detection. The Clinical Neuropsychologist, 34(1), 88–119.
Meyers, J. E., & Volbrecht, M. (1998). Validation of reliable digits for detection of malingering. Assessment, 5(3), 303–307.
Novitski, J., Steele, S., Karantzoulis, S., & Randolph, C. (2012). The RBANS Effort Scale. Archives of Clinical Neuropsychology, 27(2), 190–195.
Schroeder, R. W., Twumasi-Ankrah, P., Baade, L. E., & Marshall, P. S. (2012). Reliable Digit Span: A systematic review and cross-validation study. Assessment, 19(1), 21–30.
Schwartz, E. S., Erdodi, L., Rodriguez, N., Ghosh, J. J., Curtain, J. R., Flashman, L. A., & Roth, R. M. (2016). CVLT-II Forced Choice Recognition trial as an embedded validity indicator: A systematic review of the evidence. Journal of the International Neuropsychological Society, 22(8), 851–858.
Schwartz, E. S., Erdodi, L., Rodriguez, N., Ghosh, J. J., Curtain, J. R., Flashman, L. A., & Roth, R. M. (2016). CVLT-II Forced Choice Recognition trial as an embedded validity indicator: A systematic review of the evidence. Journal of the International Neuropsychological Society, 22(8), 851–858. [The CVLT-II ≤ 14 cut-off; classification accuracy and the impairment gradient in failure rates.]
Silverberg, N. D., Wertheimer, J. C., & Fichtenberg, N. L. (2007). An effort index for the Repeatable Battery for the Assessment of Neuropsychological Status (RBANS). The Clinical Neuropsychologist, 21(5), 841–854.
Slick, D. J., Sherman, E. M. S., & Iverson, G. L. (1999). Diagnostic criteria for malingered neurocognitive dysfunction: Proposed standards for clinical practice and research. The Clinical Neuropsychologist, 13(4), 545–561.
Tombaugh, T. N. (1996). Test of Memory Malingering (TOMM). North Tonawanda, NY: Multi-Health Systems.
Enter scores on Score Tables, Change Analysis, the SD Index or the premorbid page to chart them here.
How these charts are drawn
Charts read the same data and settings as their source tables, so the two cannot disagree. Each source (Score Tables, Premorbid, Change Analysis, SD Index) is its own pane, and only sources with data appear.
Score Tables charts draw each subtest in its native metric with classification bands and confidence intervals matching the table. Base-rate measures show a cumulative step line. Error measures (marked ↓) reverse the bands to describe performance. Raw-score measures are listed with their confidence interval but not plotted.
Change Analysis charts plot both testings against the reliable-change interval for the selected method. Outcomes state significance only, never a direction.
SD Index charts plot change in SD units against the significance band. Premorbid charts show predicted vs achieved scores with prediction intervals and base rates, read directly from the tables.
The Score Tables pane offers four axis modes: Native metric, Percentile, Standard score and Raw scores. Use All charts to see the full profile, or One at a time for a closer look with arrow-key paging.
Standard Deviation Index
Quantify abnormality of test-retest discrepancy in standard-deviation units. Useful when reliability data are unavailable or for descriptive comparison.
Score Type
Basic Reliable Change Index
Jacobson & Truax (1991). Computes whether observed change exceeds measurement error, using the test's reliability coefficient and standard deviation.
Test data & patient scores
| # | Subtest | SD | r | Date 1 | Date 2 | RCI (z) | p | Outcome |
|---|
Practice Effect-Adjusted Reliable Change Index
Iverson (2001). Adjusts the standard RCI to control for the average improvement (practice effect) observed between assessments in the normative sample.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | Date 1 | Date 2 | RCI (z) | p | Outcome |
|---|
McSweeney Regression-Based (SRB) Reliable Change Index
McSweeney et al. (1993). Predicts each patient's expected retest score from their baseline and the normative sample's regression parameters; the residual is standardised against the standard error of estimate.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | Date 1 | Date 2 | Ŷ₂ | RCI (z) | p | Outcome |
|---|
Crawford Regression-Based Reliable Change Index
Crawford & Garthwaite (2007). Extends the standardised regression-based approach to use a t-distributed test statistic that incorporates the normative sample size (N), correctly accounting for uncertainty in the regression parameters when N is modest. Returns a sample-size-adjusted standard error of prediction.
Test data & patient scores
| # | Subtest | M₁ | SD₁ | M₂ | SD₂ | r | N | Date 1 | Date 2 | Ŷ₂ | t(RB) | p | Outcome |
|---|
Premorbid Estimate
Inputs
Enter whichever predictors are available. Leave unavailable fields blank; the estimate table will update only for models with enough information.
Figure. Premorbid FSIQ estimates with 90% confidence intervals.
Enter the patient's actual WAIS-IV / WMS-IV index scores in the Achieved column to compute ToPF-predicted vs actual discrepancies. Base rates are the published figures from the ToPF-UK manual and are shown only for negative discrepancies (achieved < predicted). The manual derives them from a normal model with SD = SEE rather than from observed standardisation-sample frequencies.
| Index | Predicted | Lower 90% | Upper 90% | Achieved | Difference | Base rate |
|---|---|---|---|---|---|---|
| WAIS-IV | ||||||
| Full Scale IQ | - | - | - | - | - | |
| Verbal Comprehension Index | - | - | - | - | - | |
| Perceptual Reasoning Index | - | - | - | - | - | |
| Working Memory Index | - | - | - | - | - | |
| Processing Speed Index | - | - | - | - | - | |
| WMS-IV | ||||||
| Immediate Memory Index | - | - | - | - | - | |
| Delayed Memory Index | - | - | - | - | - | |
| Visual Working Memory Index | - | - | - | - | - | |
Enter age (16–90), sex, plus Vocabulary and/or Matrix Reasoning raw scores in the Inputs panel above. Rows appear automatically for each model whose required inputs are present. Enter the patient's actual FSIQ / GAI in the Achieved column - the prorated index is calculated per ACS manual procedures, excluding the subtest(s) used as predictors. The three FSIQ rows predict three different prorated criteria and are not expected to agree with each other.
| Model | Predicted | Lower 90% | Upper 90% | Achieved | Difference | Base Rate |
|---|---|---|---|---|---|---|
| Enter age plus Vocabulary and/or Matrix Reasoning to populate the table. | ||||||
Methods & References
A clinical psychometric calculation tool for neuropsychological report writing. All computation is local; no patient data is ever transmitted.
Methods & conventions
What this tool does
Eight working pages: Premorbid Estimate, Score Tables, Change Analysis, Performance Validity, Score Charts, Score Converter, Effect Size Tools and Data. Every calculation runs locally in the browser. No patient data is transmitted off-device, and the app works with no network connection.
The auto-fill normative database holds published parameters for seven instrument families — D-KEFS (original and Advanced), WAIS-IV, WMS-IV, WISC-V, CVLT-3, CVLT-C and the RBANS — with the retest sample size N where the publisher reports one. N is required for the Crawford & Garthwaite method and may need entering by hand where it is unavailable. Clinicians should verify every imported parameter against the current manual, and against local service standards, before interpreting it.
Score conversion and classification
Conversions between standard (M 100, SD 15), T (50, 10), scaled (10, 3) and z scores assume an approximately normal reference distribution. Two descriptor schemes are offered, and the one in force is named in the note beneath every exported table: Wechsler bands follow the WAIS-IV/WMS-IV manual conventions, and AACN labels follow Guilmette et al. (2020). Confidence levels throughout are 90% (z = 1.645) and 95% (z = 1.960); intervals round the estimate and the margin separately, so the printed bounds stay symmetric about the printed value.
Confidence intervals on Score Tables
Confidence intervals and standard errors of measurement. The CI column is the obtained score ± z × SEM, where SEM = SD × √(1 − r), centred on the obtained score rather than on an estimated true score.
The standard deviation is the normative SD of the metric the score is reported in — 15, 10, 3 or 1 — because a coefficient computed on, or corrected to, the normative sample must be paired with that sample's variability. Where a measure's stored statistics are raw, its own standard deviation is used instead, that being the only one in the right units. Four publishers state that rule outright, and the arithmetic confirms it: this pairing reproduces every published standard error of measurement the app is able to check, exactly, at the precision each is printed to — all 300 cells of WAIS-IV Table 4.3, 242 of WISC-V Table 4.4, 240 of WMS-IV Table 3.3, 168 across the D-KEFS SEM tables, 126 of RBANS Update Table 3.7, and all 38 CVLT-3 measures in Tables 3.4 and 3.5.
The reliability is, by default, the retest coefficient held in the normative database — an alternate-form coefficient in the case of the CVLT-3, which publishes no same-form retest — corrected for the normative sample's variability where the publisher reports a corrected value. Retest is the default for two reasons: it keeps a single, stated basis across a table that may mix batteries, and it is the appropriate coefficient for the many timed measures in the database, since split-half and alpha are not valid reliability estimates for speeded tests. The WAIS-IV manual makes that second argument itself for Coding, Symbol Search and Cancellation, describing the split-half coefficient as "not a proper reliability estimate" for a Processing Speed subtest; the values used here for those three are the ones it publishes, in all 38 of the cells its Table 4.1 gives them.
That default is set aside for a measure only where its publisher both reports an internal-consistency coefficient and derives its own published intervals from it. Seven manuals meet that bar:
| Instrument | Coefficient used | Source | Not applied to |
|---|---|---|---|
| CVLT-C | Odd–even split-half, by age | Manual Table 6.5 | Every index but List A Trials 1–5 Total; item scores on a word-list task are not independent. The interval printed in the manual's own worked example reproduces exactly. |
| D-KEFS | Internal consistency, by normative age band | Technical Manual, Tables 2.1–2.24 | Colour–Word Interference, whose only coefficient table is for a composite this app does not hold; Design Fluency, where item interdependence precluded the procedure; and five of the six Trail Making measures, the published table covering the composite alone. |
| D-KEFS Advanced | Split-half, by normative age band | Table 3.4 | Trail Making and Verbal Fluency, which that manual treats as speeded and scores on stability coefficients. |
| WAIS-IV | Split-half or alpha, by normative age band | Table 4.1 | Coding, Symbol Search and Cancellation — speeded. These keep the corrected stability coefficient the same table publishes for them. |
| WISC-V | Split-half, by single year of age | Table 4.1 | Coding, Symbol Search and Cancellation, together with the Cancellation Random and Cancellation Structured process scores — speeded, and likewise on the corrected stability coefficient. |
| WMS-IV | Split-half or alpha, by normative age band, Adult and Older Adult batteries separately | Table 3.1 | Verbal Paired Associates II Word Recall, a free-recall score with no consistent item count, which takes a stability coefficient. The recognition memory measures are absent altogether: their published reliability is a decision-consistency percentage, not a correlation, and cannot enter a standard error of measurement. |
| RBANS Update | Internal consistency, by normative age band | Table 3.6 | Figure Copy, Semantic Fluency, Coding, Story Recall and Figure Recall, which that table itself marks as estimated from test–retest and which therefore keep a stability coefficient taken from the same table. The four subtests reported as raw scores appear nowhere in it, the manual publishing reliability for its eight scaled subtests only, so no interval is shown for them. |
Which coefficient is right is a question for each manual rather than a policy of this tool, and the manuals genuinely disagree — the two D-KEFS manuals reach opposite conclusions about the same two test names. Each is followed as written, and the Data page names the basis actually in force for every measure in the database.
Age. Where a coefficient is tabulated by age, the interval uses the band for the patient's age, and the age used is named in the note beneath the table. Entering an age is optional. If none is entered, or the age falls outside a measure's normed range, the publisher's all-ages figure is used instead: the published average where a manual prints one, and otherwise the total-sample retest coefficient — which for the D-KEFS is that manual's own second regime rather than a substitute for a missing number. Both paths are therefore the publisher's own figures.
Reliable-change analysis is unaffected by any of the above and always uses the retest coefficient, which measures a different thing.
One consequence is worth bearing in mind when comparing output against a test manual. For the measures that remain on the retest default, where a manual derives its published intervals from internal-consistency reliability — almost always the higher of the two coefficients — the intervals shown here run wider than the manual's. They are therefore the more conservative, and answer the question how much would this score be expected to move on retesting rather than how precisely was it measured on the day. Not every publisher offers that comparison: the CVLT-3 manual declines to report internal-consistency reliability at all, on the grounds that item scores on a word-list task are not independent — recalling one word alters the probability of recalling the others, both within a trial and on later ones — and reports alternate-form coefficients in their place.
Measures with no normative-sample coefficient, and the reliability control
Every interval here multiplies a normative standard deviation by a reliability, and that is only a valid standard error of measurement when the two describe the same group. Most manuals supply a coefficient computed on, or corrected to, their normative sample. Three do not — the D-KEFS, D-KEFS Advanced and CVLT-C manuals report only the correlation observed in their own retest studies, a few dozen people each, and pair it with the normative standard deviation regardless.
The D-KEFS manual states that outright, fixing the standard deviation unit at 3 for all its scaled scores and deriving its test–retest standard errors of measurement "from the total sample of cases". Its Table 2.8 shows the arithmetic: the three Design Fluency all-ages values of 1.94, 1.97 and 2.47 are exactly 3 × √(1 − r) on the uncorrected coefficient. Those measures are therefore scored the way their own manuals score them, and the intervals shown reproduce the published ones. The Data page labels each such measure retest, uncorrected, so which rows rest on that footing can be read off rather than inferred.
The statistical objection to the pairing is nonetheless real, so Score Tables offers a reliability control with two settings. Published, the default, uses each manual's own coefficient and reproduces its printed interval. Corrected applies the standard range-restriction correction of Allen and Yen, rxx = 1 − (s²retest ÷ s²norm)(1 − r), to those measures alone, so that the coefficient describes the same population as the standard deviation it multiplies. A published coefficient is never overwritten in either setting.
The correction is not a guess: across the 267 database entries carrying both an observed and a publisher-corrected coefficient, it reproduces the publisher's own value to a median error of .003. But the resulting figures are not printed in the manuals concerned, which is why the default is Published and why the note beneath a corrected table says so. In practice the control moves 46 of the measures reachable from Score Tables, every one of them D-KEFS or D-KEFS Advanced: at the 95% level 9 intervals widen and 4 narrow, 33 are unchanged after rounding, and the largest single change is 2 scaled-score points. Reliable-change analysis is not affected by the control. Where a corrected reading is taken, the Data page follows it rather than continuing to show the published one.
Change analysis
Five methods, in ascending order of what they model. The Standard Deviation Index is descriptive only: SD Δ = (X₂ − X₁) ÷ SD, with no reliability correction. Simple Reliable Change (Jacobson & Truax, 1991) tests the observed change against measurement error. Practice Effect-Adjusted change (Iverson, 2001) subtracts the mean retest gain observed in the normative sample first. McSweeney Regression-Based change (McSweeney et al., 1993) predicts the retest score from baseline and standardises the residual against the standard error of estimate. Crawford & Garthwaite Regression-Based change (2007) does the same but with a standard error of prediction that accounts for the normative sample size and for the distance of the baseline score from the normative mean.
All p-values are two-tailed. The first four methods use the standard normal distribution; Crawford & Garthwaite uses the Student t distribution with N − 2 degrees of freedom, so a small normative sample raises the threshold — at N = 25 the 95% critical value is 2.069 against 1.960 for z.
These calculations use the retest coefficient paired with the standard deviation of the same retest sample, so that both terms describe one population. Where a publisher reports a coefficient corrected to the normative sample's variability, that value is offered as an option but is not the default, because it describes a differently distributed population from the standard deviation it would be multiplied by — and in the two regression methods the coefficient is a fitted slope, so substituting it changes the predicted score rather than only the interval. Reliability type varies by instrument and is stated in the note beneath each generated table: CVLT-3 coefficients are alternate-form, RBANS Form A coefficients are same-form retest.
Outcomes are reported as significance only — "Reliable change" or "No reliable change" — never as improvement or decline. The database holds many measures on which a higher score is the worse result (intrusions, perseverations, errors, false positives), and carries no score-direction flag, so reading a clinical direction off the sign of the statistic would assert the wrong conclusion for all of them. The signed statistic is displayed alongside, so the direction stays visible without the app interpreting it.
Performance validity
Six measures, each scored against its published cut-off: two RBANS-embedded indices — the Effort Index (Silverberg et al., 2007) and the Effort Scale (Novitski et al., 2012) — two WAIS-embedded ones — Reliable Digit Span (Greiffenstein et al., 1994; cut-offs per Schroeder et al., 2012) and the Digit Span indices (Iverson & Tulsky, 2003; Axelrod et al., 2006) — and two stand-alone tests, the Rey 15-Item with recognition trial (Boone et al., 2002) and the TOMM (Tombaugh, 1996; cut-off accuracy per Martin et al., 2020). Every result is a cut-off comparison only, reported with the published sensitivity and specificity at the chosen cut-off; the app issues no verdict on the protocol, because no single indicator supports one.
Aggregation follows Larrabee (2014): failing two or more independent validity indicators supports probable invalidity. Independence is the load-bearing word — measures derived from the same administration of the same instrument share error and cannot be counted twice, so the two RBANS indices count as one indicator between them, as do the two Digit Span measures. The Summary tab reports Larrabee's published classification accuracy at each failure count. Positive and negative predictive power appear on the TOMM tab, derived by Bayes' theorem from the meta-analytic sensitivity and specificity and a base rate the clinician selects (Martin et al., 2020, Tables 16–17) — a selectable rate rather than a fixed one, because the base rate of invalid performance differs by setting and the choice belongs to the clinician, not the app.
Embedded indices are computed from subtests that also measure genuine ability, so each carries a caution naming the populations in which it over-flags — dementia and severe impairment chief among them — and the caution takes emphasis only once that measure has actually flagged. Cut-offs validated in one edition of an instrument are labelled with that edition, and the clinician should document which edition was administered.
Premorbid estimation
Premorbid estimates combine ToPF-based and demographic equations with OPIE-4 prorated models, and produce predicted-versus-achieved discrepancy output with confidence intervals, base-rate lookups and APA-formatted export tables. Predicted-versus-achieved significance flagging uses a three-tier scheme at z = 1.645 (*), 1.960 (**) and 2.576 (***); this is separate from the 90%/95% confidence-interval selector.
OPIE-4 is provided for illustration only in a UK context and its output should not be quoted as a concrete premorbid estimate. The regression terms reproduce Holdnack et al. (2013), Table eA5.8, but the published equations also carry US education, ethnicity and region terms that are not applied here, which fixes every prediction at the US reference category (12th-grade high-school graduate, not African-American, not resident in the western US). Those terms are omitted rather than mapped because the education dummies encode how unusual a given attainment level is within the US population the model was fitted on, not years of schooling, and that does not transfer: the US reference category corresponds to A-levels if matched by years but to GCSE/O-level if matched by population position, and UK school-leaving age was raised to 16 only in 1972, so leaving school without qualifications was normative for older cohorts in a way it was not in the US sample. Expect estimates to run high for patients who left school early and low for graduates, by an amount this tool cannot quantify. For a UK demographic estimate, use the Crawford & Allan (1997) model.
Base rates. The ToPF predicted-difference base rates are the published figures from the ToPF-UK manual (Wechsler, 2011) and are used as published. The manual derives them from a normal model with SD equal to the model's standard error of estimate, rather than tabulating observed standardisation-sample frequencies, and they are labelled as such wherever they appear: every published cell equals the normal-curve proportion below that discrepancy, exactly, at the printed precision. A predicted-difference table is necessarily built this way, a standardisation sample of about a thousand cases yielding no observed frequency at every discrepancy point. The OPIE-4 discrepancy base rates of ACS Table eA5.12, by contrast, are empirical, sitting on a count grid, and are likewise used as published. The two published tables therefore answer the same question by different methods, and differ by roughly 10% relatively across the decisive −5 to −20 band on models of almost identical standard error of estimate: a discrepancy of −15 gives 3.78% on the ToPF table against about 4.3% on the OPIE-4 one. That difference is a property of the two sources, not an adjustment made here; neither table is modified.
Delis, D. C., Kaplan, E., & Kramer, J. H. (2001). Delis–Kaplan Executive Function System (D-KEFS): Technical manual. San Antonio, TX: The Psychological Corporation. [Internal-consistency coefficients and standard errors of measurement by age band, chapter 2 and Tables 2.1–2.26.]
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (1994). California Verbal Learning Test – Children's Version (CVLT-C): Manual. San Antonio, TX: The Psychological Corporation. [Split-half reliability and standard errors of measurement: Table 6.5. Standardised score equivalents: Tables A.1 and A.2.]
Delis, D. C., Kramer, J. H., Kaplan, E., & Ober, B. A. (2017). California Verbal Learning Test – Third Edition (CVLT-3): Manual. Bloomington, MN: Pearson. [Alternate-form reliability and standard errors of measurement: Tables 3.4 and 3.5.]
Holdnack, J. A., Drozdick, L., Weiss, L. G., & Iverson, G. L. (2013). WAIS-IV, WMS-IV, and ACS: Advanced clinical interpretation. Oxford: Academic Press. [OPIE-4 prorated regression coefficients: Table eA5.8. OPIE-4 discrepancy base rates: Table eA5.12.]
Randolph, C. (2012). Repeatable Battery for the Assessment of Neuropsychological Status Update (RBANS Update): Manual. Bloomington, MN: Pearson. [Reliability by age band: Table 3.6. Standard errors of measurement: Table 3.7. Test–retest and alternate-form data: Tables 3.8–3.9.]
Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Administration and scoring manual. London: Pearson Assessment. [Longest-span base rates: Tables C.4 and C.5.]
Wechsler, D. (2010). Wechsler Adult Intelligence Scale – Fourth UK Edition (WAIS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.3. Test–retest stability parameters: Table 4.5.]
Wechsler, D. (2010). Wechsler Memory Scale – Fourth UK Edition (WMS–IVUK): Technical and interpretive manual. London: Pearson Assessment. [Reliability coefficients: Table 3.1. Standard errors of measurement: Table 3.3.]
Wechsler, D. (2011). Test of Premorbid Functioning (ToPF-UK): Manual. London: Pearson Assessment. [ToPF-predicted vs obtained discrepancy base rates for the WAIS-IV and WMS-IV indices, derived by the publisher from a normal model on the standard error of estimate.]
Wechsler, D. (2014). Wechsler Intelligence Scale for Children – Fifth Edition (WISC-V): Technical and interpretive manual. Bloomington, MN: NCS Pearson. [Reliability coefficients: Table 4.1. Standard errors of measurement: Table 4.4. Test–retest stability: Table 4.7.]
Allen, M. J., & Yen, W. M. (1979). Introduction to measurement theory. Monterey, CA: Brooks/Cole. [Correction of a reliability coefficient for range restriction in the sample it was observed in.]
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum.
Crawford, J. R., & Allan, K. M. (1997). Estimating premorbid WAIS–R IQ with demographic variables: Regression equations derived from a UK sample. The Clinical Neuropsychologist, 11(2), 192–197.
Crawford, J. R., & Garthwaite, P. H. (2007). Using regression equations built from summary data in the neuropsychological assessment of the individual case. Neuropsychology, 21(5), 611–620.
Crawford, J. R., Millar, J., & Milne, A. B. (2001). Estimating premorbid IQ from demographic variables: A comparison of a regression equation vs. clinical judgement. British Journal of Clinical Psychology, 40(1), 97–105.
Guilmette, T. J., Sweet, J. J., Hebben, N., Koltai, D., Mahone, E. M., Spiegler, B. J., Stucky, K., Westerveld, M., & Conference Participants. (2020). American Academy of Clinical Neuropsychology consensus conference statement on uniform labeling of performance test scores. The Clinical Neuropsychologist, 34(3), 437–453.
Iverson, G. L. (2001). Interpreting change on the WAIS-III/WMS-III in clinical samples. Archives of Clinical Neuropsychology, 16(2), 183–191.
Jacobson, N. S., & Truax, P. (1991). Clinical significance: A statistical approach to defining meaningful change in psychotherapy research. Journal of Consulting and Clinical Psychology, 59(1), 12–19.
McSweeney, A. J., Naugle, R. I., Chelune, G. J., & Lüders, H. (1993). "T scores for change": An illustration of a regression approach to depicting change in clinical neuropsychology. The Clinical Neuropsychologist, 7(3), 300–312.
Sawilowsky, S. S. (2009). New effect size rules of thumb. Journal of Modern Applied Statistical Methods, 8(2), 597–599.
Axelrod, B. N., Fichtenberg, N. L., Millis, S. R., & Wertheimer, J. C. (2006). Detecting incomplete effort with Digit Span from the Wechsler Adult Intelligence Scale—Third Edition. The Clinical Neuropsychologist, 20(3), 513–523.
Boone, K. B., Salazar, X., Lu, P., Warner-Chacon, K., & Razani, J. (2002). The Rey 15-Item recognition trial: A technique to enhance sensitivity of the Rey 15-Item Memorization Test. Journal of Clinical and Experimental Neuropsychology, 24(5), 561–573.
Denning, J. H. (2012). The efficiency and accuracy of the Test of Memory Malingering trial 1, errors on the first 10 items of the Test of Memory Malingering, and five embedded measures in predicting invalid test performance. Archives of Clinical Neuropsychology, 27(4), 417–432.
Greiffenstein, M. F., Baker, W. J., & Gola, T. (1994). Validation of malingered amnesia measures with a large clinical sample. Psychological Assessment, 6(3), 218–224. [Reliable Digit Span: scoring, floor rule and traditional cut-off.]
Iverson, G. L., & Tulsky, D. S. (2003). Detecting malingering on the WAIS-III: Unusual Digit Span performance patterns in the normal population and in clinical groups. Archives of Clinical Neuropsychology, 18(1), 1–9.
Larrabee, G. J. (2014). False-positive rates associated with the use of multiple performance and symptom validity tests. Archives of Clinical Neuropsychology, 29(4), 364–373. [Aggregation of multiple PVTs; classification accuracy by number of failures.]
Martin, P. K., Schroeder, R. W., Olsen, D. H., Maloy, H., Boettcher, A., Ernst, N., & Okut, H. (2020). A systematic review and meta-analysis of the Test of Memory Malingering in adults: Two decades of deception detection. The Clinical Neuropsychologist, 34(1), 88–119. [TOMM cut-offs and the sensitivity/specificity values behind the predictive-power display, Tables 16–17.]
Meyers, J. E., & Volbrecht, M. (1998). Validation of reliable digits for detection of malingering. Assessment, 5(3), 303–307.
Novitski, J., Steele, S., Karantzoulis, S., & Randolph, C. (2012). The RBANS Effort Scale. Archives of Clinical Neuropsychology, 27(2), 190–195. [Effort Scale formula, screening gate and cut-off.]
Schroeder, R. W., Twumasi-Ankrah, P., Baade, L. E., & Marshall, P. S. (2012). Reliable Digit Span: A systematic review and cross-validation study. Assessment, 19(1), 21–30. [Basis for preferring the conservative RDS ≤ 6 cut-off.]
Silverberg, N. D., Wertheimer, J. C., & Fichtenberg, N. L. (2007). An effort index for the Repeatable Battery for the Assessment of Neuropsychological Status (RBANS). The Clinical Neuropsychologist, 21(5), 841–854. [Effort Index weighting table and cut-offs.]
Slick, D. J., Sherman, E. M. S., & Iverson, G. L. (1999). Diagnostic criteria for malingered neurocognitive dysfunction: Proposed standards for clinical practice and research. The Clinical Neuropsychologist, 13(4), 545–561. [The criteria the two-failure rule sits inside; they also require a substantial external incentive.]
Tombaugh, T. N. (1996). Test of Memory Malingering (TOMM). North Tonawanda, NY: Multi-Health Systems.
Cipriani, A., Furukawa, T. A., Salanti, G., Chaimani, A., Atkinson, L. Z., Ogawa, Y., Leucht, S., Ruhe, H. G., Turner, E. H., Higgins, J. P. T., Egger, M., Takeshima, N., Hayasaka, Y., Imai, H., Shinohara, K., Tajika, A., Ioannidis, J. P. A., & Geddes, J. R. (2018). Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: A systematic review and network meta-analysis. The Lancet, 391(10128), 1357–1366.
Cuijpers, P., Geraedts, A. S., van Oppen, P., Andersson, G., Markowitz, J. C., & van Straten, A. (2011). Interpersonal psychotherapy for depression: A meta-analysis. American Journal of Psychiatry, 168(6), 581–592.
Cuijpers, P., Miguel, C., Harrer, M., Plessen, C. Y., Ciharova, M., Ebert, D., & Karyotaki, E. (2023). Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression: A comprehensive meta-analysis including 409 trials with 52,702 patients. World Psychiatry, 22(1), 105–115.
Hofmann, S. G., & Smits, J. A. J. (2008). Cognitive-behavioral therapy for adult anxiety disorders: A meta-analysis of randomized placebo-controlled trials. The Journal of Clinical Psychiatry, 69(4), 621–632.
Huhn, M., Nikolakopoulou, A., Schneider-Thoma, J., Krause, M., Samara, M., Peter, N., Arndt, T., Bäckers, L., Rothe, P., Cipriani, A., Davis, J., Salanti, G., & Leucht, S. (2019). Comparative efficacy and tolerability of 32 oral antipsychotics for the acute treatment of adults with multi-episode schizophrenia: A systematic review and network meta-analysis. The Lancet, 394(10202), 939–951.
Pesch, B., Kendzia, B., Gustavsson, P., Jöckel, K.-H., Johnen, G., Pohlabeln, H., Olsson, A., Ahrens, W., Gross, I. M., Brüske, I., Wichmann, H.-E., Merletti, F., Richiardi, L., Simonato, L., Fortes, C., Siemiatycki, J., Parent, M.-E., Consonni, D., Landi, M. T., … Brüning, T. (2012). Cigarette smoking and lung cancer — Relative risk estimates for the major histological types from a pooled analysis of case–control studies. International Journal of Cancer, 131(5), 1210–1219.
Storebø, O. J., Storm, M. R. O., Pereira Ribeiro, J., Skoog, M., Groth, C., Callesen, H. E., Schaug, J. P., Darling Rasmussen, P., Huus, C.-M. L., Zwi, M., Kirubakaran, R., Simonsen, E., & Gluud, C. (2023). Methylphenidate for children and adolescents with attention deficit hyperactivity disorder (ADHD). Cochrane Database of Systematic Reviews, 3, CD009885.
Watts, B. V., Schnurr, P. P., Mayo, L., Young-Xu, Y., Weeks, W. B., & Friedman, M. J. (2013). Meta-analysis of the efficacy of treatments for posttraumatic stress disorder. The Journal of Clinical Psychiatry, 74(6), e541–e550.
Williams, A. C. de C., Fisher, E., Hearn, L., & Eccleston, C. (2020). Psychological therapies for the management of chronic pain (excluding headache) in adults. Cochrane Database of Systematic Reviews, 8, CD007407.
Wolraich, M. L., Wilson, D. B., & White, J. W. (1995). The effect of sugar on behavior or cognition in children: A meta-analysis. JAMA, 274(20), 1617–1621.
The Effect Size Tools page also cites a UK Biobank male-versus-female adult height contrast (d = 2.04) attributed to Lui et al. (2021). That source could not be verified and is not listed here; treat the value as an illustrative anchor only.
Privacy & use
What happens to the data you enter, and what you are agreeing to when you use these calculators.
Where your data goes
Nothing you enter is uploaded
Scores, tables, patient age and everything in the Working Report stay in this browser. The application makes no network requests of its own: it contains no analytics, no tracking, no error reporting and no server to send anything to. Every calculation runs on this device, and the whole suite works with no network connection at all.
The one external request is to Google Fonts, which loads the two typefaces the interface is set in. It carries no patient data. Beyond that, the only thing stored anywhere is stored here, on this computer.
Patient data lasts only as long as the tab
Entered scores and the Working Report are held in session storage, which survives a reload of the same tab and nothing else. Closing the tab, the window or the browser clears them. Nothing patient-identifiable is written to disk, and an older version of this app that did keep reports on disk has its leftovers deleted the first time this version loads.
Session storage is per-tab, so two tabs are two independent patient sessions rather than one shared report. Within a tab, use New patient in the top bar between patients: it clears every table and the age together, which is why it is one control rather than two.
Two things do persist here deliberately, and neither is patient data: your view preferences, such as drawer width and whether the report drawer starts minimised, and any custom test norms you add on the Data page, which would be worth little if they vanished with the tab.
Responsibility for use
This is a calculator, not a clinical judgement
The suite computes and formats. It does not decide anything, and it cannot tell whether the method you have chosen suits the question you are asking. Clinical interpretation rests entirely with you.
You are responsible for the inputs and the method
Every result is only as sound as the norms, reliabilities and scores entered to produce it. Verify imported parameters against the current published manual before interpreting them, and satisfy yourself that the method is the right one for the case. Where a figure is a compromise the app says so rather than hiding it: the reliability basis is named in the note under every exported table, and the Data page labels each coefficient with what it actually rests on.
You are responsible for being qualified to use it
These instruments are restricted measures. By using this tool you confirm that you hold the training, registration and professional competence required to administer, score and interpret them, and that you are working within your scope of practice and your local service standards. The tool does not check this and cannot.
Read the note that travels with the table
Every exported table carries an APA note stating the method used, the reliability basis in force and the age band applied. That note is the record of how the numbers were produced, and it is written to be read by whoever receives the report. Specific cautions are flagged where they apply: OPIE-4 is labelled illustrative-only for UK use because the published equations carry US education, ethnicity and region terms that are not applied here, and the discrepancy base rates on the premorbid page are a parametric normal model rather than observed frequencies.
Reports produced here are your responsibility
Output from this tool may end up in clinical records and in medico-legal reports that are scrutinised. Check the numbers before they leave. No warranty is given that any calculation is fit for a particular purpose, and responsibility for anything written on the basis of them is yours.