Clinical assessment. 12 minute read.

Outcome measures in physical therapy: how to choose them, read the change and use them in your notes

Outcome measures in physical therapy are standardized questionnaires and tests that you repeat across an episode of care to show whether a patient is really changing. In practice that means one patient-reported measure for the problem and one performance test for the task that matters most, both scored at the first visit and again at planned reviews and discharge. Judge each change against the measure's minimal detectable change (MDC) and minimal clinically important difference (MCID), write both the score and the comparison in your note, and let the result shape the home exercise program (HEP).

The sections below go through what regulators ask for, how to pick a measure and how to read a change score. Near the end there is a table of published change values for common measures, then a worked re-evaluation note.

Why do outcome measures matter in physical therapy?

"Knee feels better" and "walking improved" are hard to act on. A score of 30 on the Lower Extremity Functional Scale (LEFS) at visit 1 and 42 at visit 6 is not. Moore and colleagues' core set guideline for adults with neurologic conditions puts the case plainly: they let clinicians track changes in a patient's status over time and put numbers on both observations and patient-reported function. The same guideline says measures help with communication and with comparing results across patients and services.

They also help the patient. The Moore guideline recommends discussing the purpose of each measure and its results with the patient, and deciding together how the results should shape the plan of care. A patient who sees their chair stand count go from 9 to 11 has a reason to keep doing the HEP.

Use has not always matched the theory. Jette and colleagues surveyed 1,000 randomly selected American Physical Therapy Association (APTA) members, and 48% of those who answered used standardized outcome measures. More than 90% of users felt the measures helped communication with patients and helped direct the plan of care. The barriers they reported were practical: the time patients needed to fill the forms in, the time clinicians needed to analyze them, and patients finding the forms hard to complete alone. The survey dates from 2009, so current practice may differ.

What professional bodies and regulators expect

In the USA, the APTA Standards of Practice (last updated 2024) say the physical therapist examination includes "selecting, and performing or ordering, appropriate diagnostic and or physiologic procedures, tests, and measures including outcome measures". The APTA documentation guidelines ask for plan-of-care goals "stated in measurable terms". The discharge summary should record the patient's current physical and functional status and how far the goals were met.

Medicare once required functional data on claims. From January 1, 2013, functional reporting asked therapy providers to report nonpayable G-codes and severity modifiers at the start of therapy, at least once every 10 treatment days, at each evaluation or re-evaluation, and at discharge. The severity modifier reflected the patient's level of functional impairment as judged by the therapist. The Centers for Medicare and Medicaid Services (CMS) discontinued the requirement for dates of service on and after January 1, 2019. Your plan of care still needs measurable goals under the APTA guidelines, and if you bill other payers, check what they ask for in your notes.

The Health and Care Professions Council (HCPC) sets the UK standards of proficiency for physiotherapists. The current version, in force since 1 September 2023, says physiotherapists must "evaluate care plans or intervention plans using recognised and appropriate outcome measures, in conjunction with the service user where possible, and revise the plans as necessary" (standard 11.5). Standard 11.4 covers taking part in quality management, including the use of appropriate outcome measures.

In Australia and New Zealand, the Physiotherapy practice thresholds from the two national boards describe the threshold competence needed for first registration and for continuing practice. They include the ability to "incorporate relevant diagnostic tests, assessment tools and outcome measures during the physiotherapy assessment". The thresholds also ask physiotherapists to use specific and relevant measures to evaluate a client's response to physiotherapy and recognize when that response is not as expected. Another item asks them to measure outcomes, analyze clients' responses to physiotherapy and plan changes to improve the results.

Patient-reported vs performance-based outcome measures

A patient-reported outcome measure (PROM) asks the patient. The LEFS, the Oswestry Disability Index (ODI), the Neck Disability Index (NDI), the QuickDASH and the Knee injury and Osteoarthritis Outcome Score (KOOS) are all questionnaires about pain and function in daily life. A pain rating on a 0 to 10 scale is a PROM too; our pain scales page covers the numeric and visual analog versions and their published change values.

The Patient-Specific Functional Scale (PSFS) is a PROM built around the patient's own priorities. The patient names activities they find hard and rates each one on an 11-point scale, where 0 means unable to perform and 10 means able to perform at the level before the problem started. Because the activities come from the patient, the PSFS doubles as a goal list.

A performance-based measure asks the patient to do something while you time or count it. The Timed Up and Go test, the 30-second chair stand test and the six-minute walk test are the common ones in outpatient and community practice. Each of those pages explains how to run the test, with published norms and the limits of the cutoffs.

The two types answer different questions. A PROM tells you how the patient experiences their function across a normal week. A performance test tells you what they can do today, under set conditions, in front of you. Using one of each gives a fuller picture than two of the same kind.

How to choose an outcome measure

Start from the problem and the patient, then run the candidate measure through a few questions.

  1. Does it measure what you are treating? A knee questionnaire suits knee osteoarthritis; it tells you little about a patient whose main limit is balance. This is validity: the measure reflects the thing you care about in this kind of patient.
  2. Does it give the same score when nothing has changed? This is reliability. Binkley and colleagues found excellent test-retest reliability for the LEFS (R = .94). Cleland and colleagues found only fair to moderate test-retest reliability for the NDI in mechanical neck pain (ICC .50). A less reliable measure needs a bigger change before you can trust it.
  3. Can it pick up the change you expect? A measure that barely moves when patients improve will not help you decide anything.
  4. Will it fit into a normal session? The Shirley Ryan AbilityLab database lists about 5 minutes for the LEFS and under 4 minutes for the PSFS. Time, for the patient and for you, was among the barriers reported most often in the Jette survey.
  5. Does it have a published MDC or MCID from patients like yours? Without one, you have a number but no way to judge its change.

Condition-specific clinical practice guidelines often point you to a short list. The Moore guideline is one example: for adults with neurologic conditions it recommends a core set of the Berg Balance Scale, the Activities-specific Balance Confidence Scale, the Functional Gait Assessment, the 10 meter walk test, the 6-minute walk test and the 5 times sit-to-stand, alongside the patient's own goals.

MDC and MCID explained

These two numbers tell you whether a change in score means anything. Haley and Fragala-Pinkham describe two separate questions: is the change bigger than the error you expect from the measure itself, and does it make a difference in the patient's life?

The minimal detectable change (MDC) answers the first question. Every measure has some noise: the patient has a good day, reads a question differently or tries harder on the retest. Binkley and colleagues put the error around a single LEFS score at plus or minus 5.3 points, and the MDC at 9 points, both with 90% confidence. A change smaller than the MDC may be measurement error rather than real change.

The minimal clinically important difference (MCID) answers the second. It is the smallest change that patients experience as meaningful. Researchers often find it by comparing each patient's change in score with their own rating of how much better they feel; Cleland and colleagues used a global rating of change scale this way for the NDI. A change can be real but too small to matter to the patient, so check both numbers.

MCID values are not fixed properties of a test. They shift with the patient group and the setting, and with the method used to calculate them. Wright and colleagues calculated the change linked with a major improvement in four performance tests for hip osteoarthritis, using three different methods, and concluded that the methods "provided very different results". So quote the value from the study closest to your patient, and name it in your note.

Common outcome measures and their published change values

The table lists values reported in the studies named. They come from specific groups of patients and may not transfer to yours.

Measure What it covers Scoring Published change value Where it comes from
Lower Extremity Functional Scale (LEFS) Lower limb function 20 items, 0 to 80, higher is better MDC 9 points and MCID 9 points; 9 to 16 points in a later clinic sample, depending on how large a change patients rated Binkley 1999; Abbott and Schmitt 2014
Patient-Specific Functional Scale (PSFS) Activities the patient chooses Each activity 0 to 10, 10 is the level before the problem 1.3 (small change), 2.3 (medium), 2.7 (large) Abbott and Schmitt 2014, 1,708 patients with musculoskeletal disorders in 5 physical therapy clinics
Oswestry Disability Index (ODI) Low back pain disability 0 to 100, higher is more disability 10 points, or a 30% improvement from baseline Ostelo 2008, international consensus for low back pain
Neck Disability Index (NDI) Neck pain disability 10 items scored 0 to 5; 0 to 50, or 0 to 100 as a percentage; higher is more disability 19 percentage points (0 to 100 scale) Cleland 2008, 137 outpatients with neck pain
QuickDASH Arm, shoulder and hand disability 11 items, 0 to 100, higher is more disability 8.0 points Gummesson 2006 (scoring); Mintken 2009, 101 patients with shoulder pain
Knee injury and Osteoarthritis Outcome Score (KOOS) Knee: 5 separately scored subscales 0 to 100, 100 is no knee problems 8 to 10 points suggested as the smallest change patients notice, based on early data after knee ligament (ACL) reconstruction Roos and Lohmander 2003
Numeric pain rating scale Pain intensity 0 to 10 2 points Ostelo 2008, low back pain
Timed Up and Go (TUG) Basic mobility Seconds 0.8 to 1.4 s faster, for a major improvement (not the minimal change), by three methods Wright 2011, 65 patients with hip osteoarthritis
30-second chair stand test Getting up from a chair, leg strength Full stands in 30 s 2.0 to 2.6 stands for a major improvement (hip osteoarthritis); minimal change 2.5 stands and 0.7 stands in two knee osteoarthritis studies Wright 2011; Mostafaee 2024; Ramalho 2025
Six-minute walk test (6MWT) Walking capacity Meters Minimal important difference 30 m; the second walk averaged 26 m farther than the first (learning effect) Holland 2014 and Singh 2014, chronic respiratory disease

A few notes on reading the table. Cleland and colleagues pointed out that the NDI change needed to be sure it went beyond measurement error was twice the figure reported before them. The chair stand values show how far studies can differ. In knee osteoarthritis alone, one study (Mostafaee and colleagues) reported 2.5 stands; another (Ramalho and colleagues) found 0.7 by the method they preferred and 1.5 by another.

For the six-minute walk test, the learning effect is the number to know first, because it is almost as large as the 30 m change that matters. The European Respiratory Society and American Thoracic Society technical standard says 2 tests must be done when you use the test to measure change over time, at least 30 minutes apart, and the better distance recorded. The 30 m value comes from people with lung disease and may not apply to other conditions. Our six-minute walk test page covers that standard in detail.

Our LEFS guide has the full scoring rules and more of the research behind them. The PSFS guide does the same for the Patient-Specific Functional Scale.

Choosing by body region

Lower limb. The LEFS covers most problems from the hip down to the foot, and the KOOS goes deeper on the knee. Pair either with the chair stand test or the TUG. Our knee osteoarthritis and hip osteoarthritis programs show the kind of exercise these measures often track.

Spine. The ODI suits low back pain, and the NDI suits neck pain. A PSFS alongside either one brings in the activities the patient cares about, such as sitting through a work meeting or turning to reverse the car.

Upper limb. The QuickDASH covers the whole arm, from the shoulder to the hand, including a stiff frozen shoulder. Add the PSFS for the specific tasks, like reaching a top shelf or fastening a bra.

Balance and falls. The TUG and the chair stand test are quick screens in older adults, but neither should be used on its own to judge falls risk, as our tool pages explain. For neurologic conditions, start from the core set in the Moore guideline.

How often should you measure?

Measure at the first visit, before treatment starts. That score is your baseline, and the APTA Standards of Practice place outcome measures within the examination.

Measure again whenever you formally review progress. The APTA documentation guidelines describe re-examination as repeated or new examination elements used to evaluate progress and to modify or redirect treatment. The Moore guideline recommends testing at admission and discharge, and in between when feasible.

Measure at discharge. The discharge summary should state current function and how far the goals were reached, and a score beside each goal answers that directly.

How often to review in between depends on the patient and the setting. There is little value in retesting before a change larger than the MDC is likely, because a smaller shift may be noise. And keep the conditions identical each time: the same version of the form, the same chair and course, the same instructions and the same walking aid. For the six-minute walk test, allow for the learning effect described above.

How to use outcome measures in SOAP notes and the HEP review

Treat a measure like any other finding: write it so someone else could repeat it. Record the name and version of the measure, the score, the conditions of the test, the baseline score, and the MDC or MCID you are comparing against with its source. For the rest of the note, see the SOAP notes guide; the free SOAP note template has a place for the scores.

Where the scores sit in the note matters less than doing it the same way each time, and clinics differ here. One common habit is to record questionnaire scores under subjective, since they are the patient's own report, and test results under objective. Other clinics list every outcome measure under objective, and so does our SOAP note template. Whichever you choose, the assessment section is where you explain what the change means.

Write the goals in the same units. "LEFS 45 or higher in 6 weeks" is a goal you can check at the review. "Improve function" is not. The PSFS makes this easy, because each activity the patient named becomes a goal with a number beside it.

Then let the scores steer the HEP. The home exercise program guide suggests deciding at each review whether to progress, keep, change or drop each exercise, and outcome scores give that decision some footing. If a score has stalled below the MDC at two reviews, find out how much of the program the patient has really been doing before you add more. Why patients don't do their home exercises has tips for that conversation.

Worked example: a re-evaluation note

Mr. T and his scores are made up. The note shows how to record and read the measures, not what to prescribe. Patients reading this: your own physio picks the measures and adjusts the program for you.

Mr. T, 68, hip osteoarthritis, re-evaluation at visit 6.

S (subjective)

  • PSFS: climbing the stairs at home 4/10 (2/10 at visit 1); walking the dog for 20 minutes 5/10 (3/10). Average 4.5 (2.5).
  • LEFS 41/80 (33/80 at visit 1).
  • Doing the home exercises most days; missed a few in a busy week at work.

O (objective)

  • 30-second chair stand: 11 stands (9 at visit 1). Same 43 cm (17 inch) chair against the wall, arms crossed, same instructions.
  • TUG: 10.1 s (11.3 s at visit 1), usual shoes, no walking aid.

A (assessment)

  • PSFS average up 2.0: above Abbott and Schmitt's 1.3 for a small change, just short of their 2.3 for a medium one.
  • Chair stand up 2 and TUG 1.2 s faster: both reach the lowest of the values Wright and colleagues linked with major improvement in hip osteoarthritis (2.0 stands and 0.8 s), but not the highest (2.6 stands and 1.4 s).
  • LEFS up 8: below the 9-point MDC from Binkley and colleagues, so not yet clearly beyond measurement error. Recheck at discharge.
  • Stairs remain the main limit.

P (plan)

  • HEP reviewed: sit to stand progressed to a lower seat; step-up added to work toward the stairs goal.
  • Goals unchanged: stairs at home 7/10 or higher on the PSFS, LEFS 45 or higher.
  • Repeat all four measures at discharge with the same set-up.

Read it the way a colleague would. Every number has a baseline beside it, the assessment says which changes are real and which are not yet, and the plan ties the new exercise to the goal the patient chose.

In PocketPhysio, adding the step-up to Mr. T's program is about a minute's work once you know the library: pick the exercise, set the sets and reps, and type a cue of your own. He gets a video of each exercise with a voice guide to follow. You can send the program by SMS, by email or as a link, or through the patient app, Pocket Physio Care; WhatsApp also works. At discharge you can see what he was given and progress it from there.

When a score is not the whole story

An outcome score tells you about function. It is not a screen for serious problems. If a score gets worse or a new symptom appears, reassess the patient and go through your red flag questions again before changing the exercises.

Performance tests are exercise, so screen before you run them. Our chair stand, TUG and six-minute walk pages list when not to run each test and when to stop. Stop any test straight away for chest pain or pressure, dizziness or feeling faint, a racing or irregular heartbeat, marked shortness of breath, sweating or a pale or gray look, leg cramps or sharp joint pain, or if the patient staggers or loses their balance. Help the patient sit or lie down as needed, and act on chest pain as the rule below says. Record why you stopped: a test you had to stop gives no score you can compare with later ones.

The chest pain rule is the same in the clinic and at home, so teach it to patients for their home exercises. If chest pain, pressure or tightness does not ease quickly when you rest, spreads to your arm, neck, jaw, stomach or back, or comes with sweating, feeling sick, feeling light-headed or being short of breath, call emergency services straight away, as this can be a heart attack. If it eases quickly and none of these happen, get medical advice the same day, and do not exercise again until you have been checked. If you faint while exercising, call emergency services, even if you feel fine again quickly.

The short version

Outcome measures in physical therapy turn "better" into a number you can check. Use one patient-reported measure and one performance test that fit the problem, and score them at the start, at planned reviews and at discharge, the same way each time. Compare each change with the MDC to see if it is real and with the MCID to see if it matters, naming the study you used. Then put the scores and your reading of them in the note, and let them decide the next step in the HEP.

References

  1. American Physical Therapy Association. Standards of Practice for Physical Therapy. HOD S07-24-08-12, last updated 24 September 2024. https://www.apta.org/siteassets/pdfs/policies/standards-of-practice-pt.pdf
  2. American Physical Therapy Association. Guidelines: Physical Therapy Documentation of Patient/Client Management. BOD G03-05-16-41, last updated 19 May 2014. https://www.apta.org/siteassets/pdfs/policies/guidelines-documentation-patient-client-management.pdf
  3. Centers for Medicare and Medicaid Services. Functional Reporting. https://www.cms.gov/medicare/billing/therapyservices/functional-reporting
  4. Health and Care Professions Council. Standards of proficiency: physiotherapists. Effective 1 September 2023. https://www.hcpc-uk.org/standards/standards-of-proficiency/physiotherapists/
  5. Physiotherapy Board of Australia and Physiotherapy Board of New Zealand. Physiotherapy practice thresholds in Australia and Aotearoa New Zealand. Amendments approved October 2023. https://cdn.physiocouncil.com.au/assets/volumes/downloads/Physiotherapy-Board-Physiotherapy-practice-thresholds-in-Australia-and-Aotearoa-New-Zealand.PDF
  6. Jette DU, Halbert J, Iverson C, Miceli E, Shah P. Use of standardized outcome measures in physical therapist practice: perceptions and applications. Physical Therapy. 2009;89(2):125-135. doi:10.2522/ptj.20080234
  7. Moore JL, Potter K, Blankshain K, Kaplan SL, O'Dwyer LC, Sullivan JE. A core set of outcome measures for adults with neurologic conditions undergoing rehabilitation: a clinical practice guideline. Journal of Neurologic Physical Therapy. 2018;42(3):174-220. doi:10.1097/NPT.0000000000000229
  8. Haley SM, Fragala-Pinkham MA. Interpreting change scores of tests and measures used in physical therapy. Physical Therapy. 2006;86(5):735-743. doi:10.1093/ptj/86.5.735
  9. Binkley JM, Stratford PW, Lott SA, Riddle DL. The Lower Extremity Functional Scale (LEFS): scale development, measurement properties, and clinical application. Physical Therapy. 1999;79(4):371-383. doi:10.1093/ptj/79.4.371
  10. Abbott JH, Schmitt J. Minimum important differences for the patient-specific functional scale, 4 region-specific outcome measures, and the numeric pain rating scale. Journal of Orthopaedic and Sports Physical Therapy. 2014;44(8):560-564. doi:10.2519/jospt.2014.5248
  11. Ostelo RW, Deyo RA, Stratford P, et al. Interpreting change scores for pain and functional status in low back pain: towards international consensus regarding minimal important change. Spine. 2008;33(1):90-94. doi:10.1097/BRS.0b013e31815e3a10
  12. Cleland JA, Childs JD, Whitman JM. Psychometric properties of the Neck Disability Index and Numeric Pain Rating Scale in patients with mechanical neck pain. Archives of Physical Medicine and Rehabilitation. 2008;89(1):69-74. doi:10.1016/j.apmr.2007.08.126
  13. Beaton DE, Wright JG, Katz JN; Upper Extremity Collaborative Group. Development of the QuickDASH: comparison of three item-reduction approaches. Journal of Bone and Joint Surgery (American). 2005;87(5):1038-1046. doi:10.2106/JBJS.D.02060
  14. Gummesson C, Ward MM, Atroshi I. The shortened disabilities of the arm, shoulder and hand questionnaire (QuickDASH): validity and reliability based on responses within the full-length DASH. BMC Musculoskeletal Disorders. 2006;7:44. doi:10.1186/1471-2474-7-44
  15. Mintken PE, Glynn P, Cleland JA. Psychometric properties of the shortened Disabilities of the Arm, Shoulder, and Hand Questionnaire (QuickDASH) and Numeric Pain Rating Scale in patients with shoulder pain. Journal of Shoulder and Elbow Surgery. 2009;18(6):920-926. doi:10.1016/j.jse.2008.12.015
  16. Roos EM, Lohmander LS. The Knee injury and Osteoarthritis Outcome Score (KOOS): from joint injury to osteoarthritis. Health and Quality of Life Outcomes. 2003;1:64. doi:10.1186/1477-7525-1-64
  17. Wright AA, Cook CE, Baxter GD, Dockerty JD, Abbott JH. A comparison of 3 methodological approaches to defining major clinically important improvement of 4 performance measures in patients with hip osteoarthritis. Journal of Orthopaedic and Sports Physical Therapy. 2011;41(5):319-327. doi:10.2519/jospt.2011.3515
  18. Mostafaee N, Rashidi F, Negahban H, Ebrahimzadeh MH. Responsiveness and minimal important changes of the OARSI core set of performance-based measures in patients with knee osteoarthritis following physiotherapy intervention. Physiotherapy Theory and Practice. 2024;40(5):1028-1039. doi:10.1080/09593985.2022.2143253
  19. Ramalho RB, Chaves TC, Terluin B, Selistre LFA. Minimal important changes of common outcome measures of physical function in individuals with knee osteoarthritis: a prospective clinical study. Archives of Physical Medicine and Rehabilitation. 2025;106(12):1829-1836. doi:10.1016/j.apmr.2025.04.016
  20. Singh SJ, Puhan MA, Andrianopoulos V, et al. An official systematic review of the European Respiratory Society/American Thoracic Society: measurement properties of field walking tests in chronic respiratory disease. European Respiratory Journal. 2014;44(6):1447-1478. doi:10.1183/09031936.00150414
  21. Holland AE, Spruit MA, Troosters T, et al. An official European Respiratory Society/American Thoracic Society technical standard: field walking tests in chronic respiratory disease. European Respiratory Journal. 2014;44(6):1428-1446. doi:10.1183/09031936.00150314
  22. Shirley Ryan AbilityLab Rehabilitation Measures Database. Lower Extremity Functional Scale. https://www.sralab.org/rehabilitation-measures/lower-extremity-functional-scale
  23. Shirley Ryan AbilityLab Rehabilitation Measures Database. Patient Specific Functional Scale. https://www.sralab.org/rehabilitation-measures/patient-specific-functional-scale
  24. Shirley Ryan AbilityLab Rehabilitation Measures Database. Neck Disability Index. https://www.sralab.org/rehabilitation-measures/neck-disability-index

Written and checked by the PocketPhysio editorial team. Last updated 2026-09-28.