The PHQ-9 and GAD-7 are the two most widely administered symptom measures in outpatient behavioral health, and they're often used together because depression and anxiety co-occur so frequently. Both are brief, free to use, and well validated, but their clinical value depends almost entirely on how you interpret the scores and what you do with them afterward. This guide covers scoring, cutoffs, change interpretation, documentation, and the limits you should hold in mind.
Key Takeaways
Both instruments are validated screeners and severity measures, not diagnostic tools. In the original validation study, a PHQ-9 score of 10 or higher identified major depression with 88% sensitivity and 88% specificity, which is strong screening performance and still not a diagnosis.
Repeated administration is where the clinical value lives. A single baseline score tells you where someone is, while serial scores tell you whether your treatment is working and give you defensible data for the golden thread.
Scoring and documenting a standardized instrument is a billable service under CPT 96127 when the documentation requirements are met, so routine measurement can support both clinical quality and revenue.
What Each Instrument Measures
The PHQ-9 assesses the nine DSM symptom criteria for major depressive disorder over the previous two weeks. Each item is rated on a 4-point scale from 0 (not at all) to 3 (nearly every day), producing a total score from 0 to 27. It covers mood, anhedonia, sleep, energy, appetite, self-perception, concentration, psychomotor changes, and thoughts of death or self-harm.
The GAD-7 assesses generalized anxiety symptoms over the same two-week window. It uses seven items on the same 4-point scale, producing a total score from 0 to 21. It covers worry, restlessness, irritability, difficulty relaxing, and related somatic and cognitive features.
Both instruments include a final functional impairment question asking how difficult the reported symptoms have made work, home life, and relationships. That item isn't included in the total score, and it's routinely overlooked. It's often the single most clinically useful line on the form, because two clients with identical scores can have very different levels of impairment, and impairment is what drives treatment intensity decisions.
Both were developed with Pfizer funding, and the copyright holder states that no permission is required to reproduce, translate, display, or distribute them, which is why they're standard across primary care and behavioral health.
Scoring Bands and Cutoffs
The severity bands are straightforward, and worth knowing without looking them up.
For the PHQ-9: scores of 5, 10, 15, and 20 mark the thresholds for mild, moderate, moderately severe, and severe depression. For the GAD-7: scores of 5, 10, and 15 mark mild, moderate, and severe anxiety.
A score of 10 is the conventional screening cutoff for both. In its original validation, at a cutoff of 10 the GAD-7 showed 89% sensitivity and 82% specificity for detecting generalized anxiety disorder. Some validation work supports a slightly lower GAD-7 threshold of 8 or 9 depending on population, and performance varies meaningfully across settings. In university student samples, for example, both instruments have shown lower specificity and higher false positive rates, which is why they're recommended as initial screens with positive cases followed by fuller assessment.
The practical implication for licensed clinicians: treat 10 as a threshold for closer attention rather than a verdict. In a specialty behavioral health caseload, where base rates are far higher than in primary care, a score above cutoff confirms very little you didn't already suspect. The number's job is to quantify severity and track it, not to make the call.
Interpreting Change Over Time
This is where these instruments earn their place in your workflow, and where most practices underuse them.
Commonly used thresholds for clinically meaningful change are a 5-point reduction on the PHQ-9 and a 4-point reduction on the GAD-7. A change smaller than that may reflect measurement noise rather than real improvement. A client moving from 18 to 16 hasn't demonstrably improved. A client moving from 18 to 11 has, and that's a documentable treatment response.
A few interpretive cautions to hold:
Regression to the mean affects second administrations. Clients often score highest at intake, when distress prompted the referral, so some drop at session four reflects natural fluctuation as much as treatment.
Score increases aren't automatically treatment failure. Symptom reporting can rise as insight improves, avoidance decreases, or trauma processing begins. Document your clinical interpretation rather than the number alone.
Plateaus carry information. A client stable at moderate severity across eight sessions is data supporting a change in approach, a referral, or a medication consultation.
Reviews of measurement-based care consistently find that routine outcome monitoring with feedback to the clinician improves outcomes, with the largest benefit for clients who aren't responding. The mechanism isn't the score itself. It's that the score surfaces non-response early enough to do something about it.
A workable cadence for most outpatient practices is baseline at intake, then every four to six sessions, then at discharge. Weekly administration adds burden without proportional information for most presentations, though it can be appropriate in intensive or higher-acuity settings.
Handling PHQ-9 Item 9
The ninth PHQ-9 item asks about thoughts of being better off dead or of self-harm, and it requires a defined protocol rather than case-by-case judgment.
Any endorsement above zero warrants direct clinical follow-up in that session. The item is a screening prompt, not a risk assessment, and it doesn't distinguish passive ideation from intent or planning. A client endorsing "several days" may be describing intermittent passive thoughts, or may be understating something considerably more acute.
What this means practically:
Review scored measures before or at the start of the session, not afterward, so an endorsement can be addressed while the client is still with you.
Follow up with a structured risk assessment appropriate to your setting rather than relying on the item alone.
Document the endorsement, your follow-up assessment, your clinical determination, and any resulting safety planning or change in level of care. An endorsed item with no documented response is a significant liability finding.
Have a defined protocol for measures completed between sessions or through a portal, since an endorsement sitting unreviewed for days is a real risk.
If your practice administers these measures asynchronously, build a review step into your workflow before anything else about the process gets standardized.
Known Limitations
Both instruments are strong for what they are, and you should know where they weaken.
Somatic item overlap. Sleep disturbance, fatigue, appetite change, and concentration difficulty are inflated by medical illness, chronic pain, pregnancy, and medication effects, which can push PHQ-9 scores up independent of mood.
Cultural and linguistic variation. Validation differs substantially across populations and translations, and optimal cutoffs aren't universal. Somatic expression of distress is more prominent in some cultural contexts, which shifts item performance.
Limited scope. The PHQ-9 doesn't screen for bipolar disorder, and treating an elevated score as unipolar depression without assessing for hypomanic history is a recognized clinical error. The GAD-7 was developed for generalized anxiety and performs less precisely for panic, social anxiety, and PTSD.
Self-report constraints. Scores reflect what the client is willing and able to report, which is affected by alexithymia, minimization, secondary gain, and cognitive load.
Not a substitute for clinical formulation. A number doesn't tell you why, and your case conceptualization still carries the interpretive work.
None of this argues against using them. It argues for reading the score alongside your mental status exam, your clinical observation, and the client's own account.
Documenting and Billing Standardized Measures
Documentation is what converts a completed form into a defensible clinical record and a billable service.
CPT 96127 covers a brief emotional or behavioral assessment using a standardized instrument, including scoring and documentation. The PHQ-9 and GAD-7 both qualify, and each instrument counts as a separate unit, so a client completing both in one visit can generate two units. Most payers cap the code at four units per patient per date of service, and non-standardized questionnaires or clinical interview alone don't qualify. Payer rules vary, so verify yours before building it into your billing routine. Our overview of CPT codes for psychotherapy covers how this fits alongside your session codes.
To support the claim and the clinical record, document:
The specific instrument administered, named in full rather than abbreviated.
The raw score and the corresponding severity band.
Your clinical interpretation, including how the score compares to prior administrations.
The action taken, whether that's a treatment plan adjustment, a referral, a change in session frequency, or continued monitoring.
Any item-level findings requiring follow-up, particularly risk endorsements.
That last element is what ties measurement into the golden thread. A score documented with no interpretation and no resulting clinical decision reads to a reviewer as a form that was collected rather than a service that was delivered. When you connect the score to a treatment plan goal, as in a goal targeting reduction in anxiety symptoms with the GAD-7 as the measure, you've created exactly the kind of objective progress evidence payers look for. Our guidance on evaluating client progress and writing measurable treatment goals covers how to build that link deliberately.
Bringing Measurement Into Your Documentation Workflow
The administrative reality is that scores get collected more often than they get integrated. Forms accumulate in the chart while progress notes continue describing session content, and the measurement data never reaches the place where it would strengthen the record.
Berries is an AI scribe built exclusively for mental health clinicians, and it's designed to keep clinical data connected across the record rather than siloed. It generates progress notes from live sessions, dictated summaries, or typed input, produces treatment plans aligned to your documented clinical content, and maintains golden thread continuity between diagnoses, goals, interventions, and session-to-session progress. It supports SOAP, DAP, BIRP, and other standard formats, suggests ICD-10 codes, and works with any EMR.
Berries is HIPAA and PHIPA compliant and SOC 2 certified, doesn't store session recordings, and doesn't use client data to train its models. Your first 20 sessions are free with no credit card required. Discounts are available for students, trainees, and early career clinicians.
Frequently Asked Questions
Can the PHQ-9 or GAD-7 be used to make a diagnosis?
No. Both are screening and severity measures. A diagnosis requires clinical interview, differential consideration, assessment of duration and functional impairment, and ruling out medical and substance-related causes. An elevated score supports a diagnostic hypothesis, nothing more.
How often should I administer them?
Baseline at intake, then every four to six sessions, then at discharge works for most outpatient caseloads. Higher acuity, medication management, or measurement-based care programs may warrant more frequent administration. Match the cadence to the clinical question you're trying to answer.
Are the PHQ-9 and GAD-7 free to use?
Yes. The copyright holder explicitly permits reproduction, translation, display, and distribution without seeking permission, which is why they appear in EHRs, primary care workflows, and research protocols worldwide.
What should I do if a client's score goes up?
Investigate rather than assume deterioration. Rising scores can reflect improved insight, reduced avoidance, an external stressor, the start of trauma processing, or genuine worsening. Document your clinical interpretation and any resulting adjustment. A rise that persists across multiple administrations warrants a more substantial review of the treatment approach.
Can I bill for administering both instruments in the same session?
Typically yes, as each standardized instrument counts as a separate unit of CPT 96127, subject to payer unit limits and documentation requirements. Confirm your specific payer policies, since caps, modifier requirements, and diagnosis pairings vary.
Should clients complete these before or during session?
Before is generally better, since it preserves session time and lets you review results in advance. The tradeoff is that any risk endorsement needs a defined review protocol so it isn't sitting unseen. If you can't guarantee timely review, in-session administration is the safer default.
What alternatives should I consider for anxiety presentations other than GAD?
The GAD-7 was validated primarily for generalized anxiety, so panic, social anxiety, OCD, and PTSD presentations are better served by disorder-specific measures used alongside it. Using the GAD-7 as a general distress indicator is reasonable, but don't treat it as a precise measure of a condition it wasn't built for.
This article is for educational purposes and professional development only. It does not constitute clinical supervision or replace professional judgment in therapeutic practice.
If you or someone you work with is experiencing thoughts of suicide or self-harm, the 988 Suicide and Crisis Lifeline is available 24/7 by call or text in the US.
