STarT Back Screening Tool Calculator

Free STarT Back Screening Tool calculator. Score all 9 items in about a minute, get the Q5-9 psychosocial subscale, and see the low, medium or high risk subgroup with the matched Keele care pathway.

STarT Back Screening Tool Calculator

Thinking about the last 2 weeks, answer all 9 items. The calculator returns the total score, the Q5–9 psychosocial subscale, the risk subgroup, and the matched care pathway Keele defines for that subgroup.

0 of 9 answered

Attribution. The Keele STarT Back Screening Tool © Keele University 01/08/07, funded by Arthritis Research UK. Hill JC, Dunn KM, Lewis M, et al. Arthritis Rheum. 2008;59(5):632–641. Keele materials may not be reproduced without permission — keele.health.

Not a diagnosis and not a red-flag screen. The STarT Back Tool stratifies prognostic risk of persistent back-related disability. It does not identify serious spinal pathology, which must be excluded separately, and it is not a substitute for assessment by a licensed clinician.

Responses are calculated entirely in your browser. Nothing you enter is transmitted or stored.

Topics Covered in this page

STarT Back Screening Tool Calculator: Score, Stratify and Match Care

Most outcome measures tell you how a patient is doing. The STarT Back Screening Tool is different: it exists to tell you what to do next. Nine questions, answered in about a minute, sort a patient with low back pain into one of three prognostic subgroups — and each subgroup comes with a specific, manualised care pathway developed at Keele University.

Score it below. Then read the parts that most STarT Back pages leave out: the item that is mis-scored more often than any other, and the honest state of the evidence — which is stronger for the tool's prognostic accuracy than for the claim that stratified care improves outcomes everywhere it has been tried.

What the STarT Back Tool does

The Subgrouping for Targeted Treatment (STarT) Back Screening Tool is a nine-item questionnaire developed by Jonathan Hill and colleagues at Keele University and published in Arthritis & Rheumatism in 2008. It stratifies patients with non-specific low back pain by their risk of persistent disabling back pain, using a mix of physical and psychosocial prognostic indicators.

It is a prognostic and allocation instrument, not a severity measure and not a diagnostic test. It was developed and validated in UK primary care, and that context matters when you transport it — see the evidence section.

The tool has been translated into more than 45 languages, has a version for children and adolescents, a six-item short form, and a LOINC panel code (91349-1) for interoperability.

The 9 items

The instruction stem is: "Thinking about the last 2 weeks tick your response to the following questions:"

Items 1 to 8 are first-person statements, answered Agree or Disagree. Item 9 is a question with five options.

#ItemConstructScoring1My back pain has spread down my leg(s) at some time in the last 2 weeksReferred leg painAgree = 12I have had pain in the shoulder or neck at some time in the last 2 weeksComorbid painAgree = 13I have only walked short distances because of my back painWalking limitationAgree = 14In the last 2 weeks, I have dressed more slowly than usual because of back painDressing limitationAgree = 15It's not really safe for a person with a condition like mine to be physically activeFear-avoidanceAgree = 16Worrying thoughts have been going through my mind a lot of the timeAnxietyAgree = 17I feel that my back pain is terrible and it's never going to get any betterCatastrophisingAgree = 18In general I have not enjoyed all the things I used to enjoyLow moodAgree = 19Overall, how bothersome has your back pain been in the last 2 weeks?BothersomenessSee below

Item 9 is the most mis-scored item on the tool. The five options score Not at all = 0 · Slightly = 0 · Moderately = 0 · Very much = 1 · Extremely = 1. "Moderately" scores zero. The Keele form prints this as an explicit 0 0 0 1 1 row beneath the options. At least one widely mirrored US copy transcribed "Moderately" as scoring 1 — which is arithmetically inconsistent with the tool's 0–9 range — and another well-known physiotherapy site states the item is scored 0–4. Both errors change the risk category for real patients. The calculator above scores it correctly.

One further caution: item 8 is negatively worded ("I have not enjoyed…"). A US copy circulating in second-person question form ("Have you enjoyed…?") inverts its meaning. Use the Keele statement form.

How to score and stratify

Two numbers come out of the nine items.

The algorithm, exactly as implemented in Keele's official calculator:

IF   total <= 3                       -> LOW RISK

ELSE IF total >= 4 AND subscale <= 3  -> MEDIUM RISK

ELSE IF total >= 4 AND subscale >= 4  -> HIGH RISK

Check the total first, then the subscale. Note that all nine items are required — the STarT Back Tool has no pro-rating rule, so an incomplete form cannot be stratified.

What each risk subgroup means

These are not labels. Each maps to a distinct package of care defined in Keele's implementation manual and tested in the IMPaCT trial.

LOW RISK — TOTAL ≤ 3

16.7% had a poor outcome at 6 months in the development cohort.

Matched pathway: a single 30-minute primary-care consultation. Comprehensive assessment including physical examination; individualised education and reassurance about diagnosis, prognosis and treatment; advice on medication, activity and work; written materials and a short educational video. The patient is then discharged with advice to re-consult if necessary. Routine onward referral to physiotherapy is explicitly not part of the low-risk pathway.

The main clinical risk in this group is over-treatment, and the economics of stratified care depend substantially on not putting low-risk patients into extended episodes of care.

MEDIUM RISK — TOTAL ≥ 4, SUBSCALE ≤ 3

53.2% had a poor outcome at 6 months. Relative risk of persistent disability versus low risk: 2.19 (95% CI 1.10–4.38).

Matched pathway: standardised physiotherapy, up to six sessions — one assessment plus up to five follow-ups. An individualised plan negotiated with the patient, aimed at reducing symptoms and disability and promoting self-management, using advice, explanation, reassurance, education, manual therapy and exercise.

The manual explicitly does not recommend bed rest, traction, massage or electrotherapy. Acupuncture is optional at the discretion of the physiotherapist and patient.

HIGH RISK — TOTAL ≥ 4, SUBSCALE ≥ 4

78.4% had a poor outcome at 6 months. Relative risk versus low risk: 7.30 (95% CI 4.11–12.98).

Matched pathway: psychologically informed physiotherapy — a 60-minute assessment plus 45-minute treatment sessions, up to six in total. Everything in the medium-risk package, plus: building rapport; validating and normalising the patient's experience; a comprehensive biopsychosocial assessment; addressing knowledge gaps and correcting misunderstandings; and creating opportunities for the patient to respond differently to difficult internal experiences. It specifically targets the four psychological prognostic indicators — catastrophising, low mood, anxiety and pain-related fear — using simple cognitive-behavioural techniques.

Keele's manual stresses that clinicians delivering this pathway need ongoing clinical supervision from appropriately skilled personnel. Allocating a high-risk patient to a clinician without that support is the most common implementation failure, and it is why the tool sometimes produces no benefit at all.

Does stratified care actually work? The honest answer

This is where most STarT Back content stops at the 2011 Lancet headline. The full picture is more useful.

THE POSITIVE EVIDENCE

Hill et al., Lancet 2011 (the STarT Back RCT). 851 patients across ten English general practices, randomised 2:1 to stratified care or current best practice. Roland–Morris Disability Questionnaire improvement at 4 months: 4.7 versus 3.0 points, adjusted difference 1.81 (95% CI 1.06–2.57), effect size 0.32. At 12 months: 4.3 versus 3.3, difference 1.06 (95% CI 0.25–1.86), effect size 0.19. Economically the strategy was dominant — +0.039 QALYs at lower cost (£240.01 versus £274.40 per patient, a £34.39 average saving).

Foster et al., IMPaCT Back, Annals of Family Medicine 2014. A population-based implementation study across 64 family physicians and 922 patients. Overall RMDQ improved 0.7 points; the high-risk group improved 2.3. Time off work halved (4 versus 8 days) and sick certification fell about 30% (9% versus 15%), with lower health care costs and no worsening of outcomes.

THE REPLICATIONS THAT FAILED

MATCH (US), Cherkin et al., JGIM 2018. A pragmatic cluster RCT across six primary care clinics in Washington State, 1,701 participants, with intensive implementation support — six clinician training sessions, five days of PT training, the tool embedded in the EHR. Result: no statistically significant differences in either primary patient outcome overall or within any risk subgroup at 6 months, and no change in care processes. Clinicians used the tool for about half of patients but, in the authors' words, "did not change the treatments they recommended."

TARGET (US), Delitto et al., eClinicalMedicine 2021. 76 primary care clinics across four US health systems, ~9,900 patients screened. Usual care plus psychologically informed physical therapy did not reduce the transition from acute to chronic low back pain: 47% versus 51% chronic at 6 months. Only 39% of intervention patients actually received the PIPT referral.

Denmark, Morsø et al., European Journal of Pain 2021. 334 patients randomised (half the intended sample). RMDQ improvement 5.9 versus 5.5 at 3 months, 6.1 versus 6.5 at 12 months — non-significant. Stratified care did produce fewer treatment sessions (3.5 versus 4.5) and lower prescription and imaging costs, but no overall cost-effectiveness advantage.

WHAT TO TAKE FROM THIS

The tool's prognostic performance replicates reasonably well. The claim that stratified care improves outcomes replicated in England and failed in the US and Denmark. The failures were largely implementation and fidelity failures rather than clean refutations — but that cuts both ways: it means the benefit depends on a delivery system that reliably changes clinician behaviour, and it will not appear simply because the questionnaire is in the chart.

The MATCH authors offered three explanations worth heeding: no audit-and-feedback on adherence (unlike the UK trials), matched treatment options that were more numerous and harder to access, and a substantially more disabled US baseline population (RMDQ 11.8 versus 8.4 in the English studies).

Practically: if you deploy the STarT Back Tool, deploy the pathways and the audit with it. Screening alone is inert.

Psychometrics

PropertyValueSourceTest–retest (quadratic weighted kappa)0.73 overall; 0.69 psychosocial subscaleHill et al. 2008Internal consistency (Cronbach's α)0.79 full tool; 0.74 psychosocial itemsMultipleICC0.90Bruyère et al. 2014Discriminative validity (AUC)0.73 (referred leg pain) to 0.92 (disability)Hill et al. 2008, developmentPoor outcome at 6 months by subgroup16.7% low / 53.2% medium / 78.4% highHill et al. 2008Inter-rater agreement with expert cliniciansWeighted kappa 0.28; 47% agreementHill et al. 2010, Clin J PainFloor / ceiling10.8% scored 0, 5.4% scored 9 (adequate); subscale floor 22.2% (poor)—

That inter-rater figure deserves its own sentence. The tool and independent expert clinicians disagree about half the time. The developers themselves recommend using it "as an adjunct to their own decision-making rather than a replacement." A calculator that returns "High risk" is offering you a prior, not a verdict.

STarT Back or the Örebro questionnaire?

The Örebro Musculoskeletal Pain Screening Questionnaire (ÖMPSQ) is the main alternative. They were built for different jobs: the ÖMPSQ was designed as a prognostic tool, the SBST as a treatment-allocation tool.

Outcome predictedSTarT Back pooled AUCÖMPSQ pooled AUCPain0.59 — non-informative0.69 — poorDisability0.74 — acceptable0.75 — acceptableAbsenteeismnot pooled0.83 — excellent

Karran et al., BMC Medicine 2017 — systematic review and meta-analysis of 18 prospective cohorts in recent-onset low back pain.

The headline finding is the one to carry into practice: the STarT Back Tool predicts disability, not pain. A pooled AUC of 0.59 for future pain is essentially non-informative. If your question is "will this patient still hurt in six months," the SBST is the wrong instrument.

Head-to-head agreement between the two is only moderate (Cohen's kappa 0.42, 70.2% agreement in 315 Swedish primary-care patients), and the SBST classified 53.7% as high risk versus 36.5% for the short ÖMPSQ — it flags substantially more people. Agreement was particularly poor in women over 50. On the other hand the SBST is far less error-prone to score in practice: 13 miscalculations versus 54 for the ÖMPSQ-short across those 315 patients. For work and absenteeism outcomes, the ÖMPSQ is the better instrument.

Limitations

Using the STarT Back Tool in a clinical workflow

Licensing

The Keele STarT Back Screening Tool is © Keele University 01/08/07, funded by Arthritis Research UK. The tool, its translations, the implementation manual and Keele's own calculator are published openly and free of charge, and Keele materials may not be reproduced without permission. Enquiries: health.iau@keele.ac.uk. Retain the copyright line on any reproduction.

Frequently asked questions

How is the STarT Back Screening Tool scored?

Items 1 to 8 score 1 for Agree and 0 for Disagree. Item 9 scores 1 only for "Very much" or "Extremely" — "Moderately" scores zero. The total runs 0 to 9, and the psychosocial subscale is the sum of items 5 to 9, running 0 to 5.

What are the STarT Back risk categories?

A total of 3 or less is low risk. A total of 4 or more with a psychosocial subscale of 3 or less is medium risk. A total of 4 or more with a subscale of 4 or 5 is high risk.

Which items make up the STarT Back psychosocial subscale?

Items 5, 6, 7, 8 and 9 — fear-avoidance, anxiety, catastrophising, low mood and bothersomeness. Some secondary sources list this incorrectly; Keele's own form labels the box "Sub Score (Q5-9)".

Does Moderately score a point on STarT Back item 9?

No. Only "Very much" and "Extremely" score 1. The Keele form prints the score row as 0, 0, 0, 1, 1 beneath the five options. This is the most commonly mis-scored item on the tool.

What treatment matches each STarT Back risk group?

Low risk: a single 30-minute primary-care consultation with education, reassurance and advice, then discharge. Medium risk: standardised physiotherapy, up to six sessions. High risk: psychologically informed physiotherapy — a 60-minute assessment plus 45-minute sessions targeting catastrophising, low mood, anxiety and pain-related fear, with clinical supervision for the treating clinician.

Does the STarT Back Tool predict pain?

No. A meta-analysis of 18 cohorts pooled its accuracy for future pain at AUC 0.59, which is non-informative. It predicts disability (AUC 0.74), not pain intensity.

Is the STarT Back Tool validated outside the UK?

It has been translated into more than 45 languages and its prognostic performance replicates reasonably well. But the claim that stratified care improves outcomes has not replicated everywhere: two large US trials (MATCH, TARGET) and a Danish trial found no significant benefit, largely because matched pathways were not reliably delivered.

Can the STarT Back Tool replace clinical judgement?

No. Agreement between the tool and independent expert clinicians is poor (weighted kappa 0.28, 47% agreement), and the developers recommend using it as an adjunct to clinical reasoning rather than a replacement. It also does not screen for red flags.

References

Did you like our content?

Why settle for long hours of paperwork and bad UI when Spry exists?

Modernize your systems today for a more efficient clinic, better cash flow and happier staff.
Schedule a free demo today