Elench · Clearpath research · October 2026

Log delivery first, compare students to themselves

The most useful analytics Clearpath can offer special educators is not an "accommodation effect" score. It is a reliable record of whether each accommodation was actually provided and used on each assessment, set beside that student's own scores over time. The only comparison that can hint at whether an accommodation helps a particular student is that student's scores on comparable assessments with and without it. Comparisons to the class average, pooled results across students who share an accommodation, and any score gap between a tailored (modified) version and the main version do not answer that question. Two constraints narrow the space further. First, IEP-mandated accommodations must be implemented as written, so a "without" condition can never be manufactured for them. For those, the honest output is fidelity plus trend, not an effect estimate. True with-versus-without comparisons are legitimate only for supports a teacher is trialing outside the IEP, for comparing one bundle of accommodations against another, for optional supports the student sometimes declines, and (with heavy caveats) for scores from before the accommodation entered the plan. Second, teachers are badly overloaded and read progress graphs poorly, so the feature must capture data as a by-product of grading, gate every comparison behind a minimum number of scores, and state results in plain, non-causal words that are safe for parents to read. Done this way, Clearpath would fill a real gap. No product found in this research computes or displays an accommodation's relationship to outcomes. And research shows that teachers' unaided judgments about who benefits are no better than chance.

Only a student's own scores can say an accommodation helped

Two definitions underpin every comparison. An accommodation changes how a student accesses or responds to content (presentation, response, timing, setting) without changing what is measured. A modification changes content or expectations, "resulting in scores that differ in meaning from scores from an unmodified assessment" (NCEO Brief 33). Clearpath's three tiers map onto this cleanly. Main and accommodated scores sit on one scale. Tailored versions tied to an IEP objective measure something else and belong on that objective's progress track. Some of Clearpath's accommodations can cross the line depending on what an item targets: read-aloud on a decoding item, a picture cue that gives the answer away, or a worked example of the exact skill being tested. Whether read-aloud is an accommodation, for instance, depends on whether the test measures decoding (IRIS Kettler interview).

Group research is a weak guide for any one student. The field's validity test is "differential boost": a fair accommodation should raise scores for students with disabilities more than for peers. Sireci, Scarpati and Li's review of 40 studies found results inconsistent "due to the wide variety of accommodations, the ways in which they were implemented, and the heterogeneity of students." Extended time was the most consistent winner (National Academies, ch. 7). Read-aloud shows how much the average hides. Buzick and Stone found an average gain of 0.56 SD on reading tests but only 0.13 SD on math (ETS). Elbaum's study of 625 secondary students went further: peers without disabilities gained more from read-aloud on math (ES 0.44) than students with LD (0.20). At the individual level, 15.5% of students with LD benefited, 76.8% showed no difference, and 7.7% did worse with read-aloud (Elbaum 2007). Elbaum concluded that assigning read-aloud to all students with LD would be "inappropriate" without prior individual evidence. Clearpath's newer accommodation types (one question at a time, picture cues, answer orally first, worked example) have no comparable research base at all. That makes the student's own data the only evidence available for them.

Teacher intuition does not fill the gap. Teachers were no more successful than chance at predicting who would benefit from read-aloud in math (BRT Tech Report 09-05, citing Helwig & Tindal 2003). In Fuchs and colleagues' study, "teachers' decisions did not correspond to benefits students derived," and a brief data-based procedure predicted large-scale test performance better than teacher judgment (Fuchs et al. 2000 abstract). That procedure, the Dynamic Assessment of Test Accommodations (DATA), gives a student parallel forms with and without each accommodation and compares the student's boost to the typical boost for general-education peers. Teachers found it manageable (OSEP brief). The counterweight is that IEP teams are not usually wrong. Slightly more than 75% of IEP-team-recommended accommodation packages produced moderate-to-large effects in Elliott, Kratochwill and McKevitt's case-by-case experiment, though for a small share of students they did nothing (ASU record). The right stance for the product is to treat each accommodation as a reasonable hypothesis that the student's data can support, question, or leave undecided.

NCEO frames the classroom question the same way. It asks, "What are the results of classroom assignments and assessments when accommodations are used verses when accommodations are not used?" It then asks whether poor results came from lack of access to instruction, from the accommodation not being provided, or from the accommodation being ineffective (NCEO Accommodations Manual, Step 5). Against that standard, the candidate comparisons rank as follows.

ComparisonValid for "did it help this student?"Use in Clearpath
Same student, used vs. not used, on comparable main-version assessments, alternating over timeYes, the strongest availableHeadline comparison, but only where a legitimate "without" exists (see next section)
Same student, before vs. after an accommodation started (AB phase change)Suggestive only; confounded by growth, practice, and new unitsPhase line on the student's chart, never a verdict
Student's gap to the class median over time (main or accommodated scores only)No, but useful context ("closing the gap?")Secondary context band
Differential boost (student's gain vs. peers' gain)Conceptually ideal, but needs peers tested both ways, which Clearpath lacksDon't build
Accommodated score vs. class average on one assessmentNo; mixes the student's level with the accommodation's effectDescriptive only, never labeled as effectiveness
Pooled across all students with the same accommodationNo; students and implementations vary too muchRoster or fidelity view only
Tailored version vs. main version, same studentInvalid; different constructNever; tailored scores go against the IEP objective criterion

Because Clearpath's assessments vary in difficulty, any cross-assessment comparison should express each score relative to that assessment's median among students who took the main version. This is imperfect, since the class is a different population, but it removes the largest confound: "Unit 3 was easier than Unit 2." Bundles are a second confound. If Noah always gets picture cues, read-aloud, and extended time together, the data cannot separate them, and the tool must credit "his usual accommodations" rather than any single one. NCEO itself treats "what combinations of accommodations seem to be effective?" as a separate question (NCEO Accommodations Manual).

The law removes the "without" condition for mandated accommodations

This is where a DATA-style "try it both ways" feature collides with the law. The IEP must specify "individual appropriate accommodations" (34 CFR 300.320). Changes happen only through the IEP team or a written amendment agreed between parent and agency (34 CFR 300.324). OCR has investigated individual teachers' failures to implement plan provisions and obtained resolution agreements committing districts to implement plans "as written" (OCR 11161400; OCR 15221073). A teacher therefore cannot withhold an IEP accommodation to generate comparison data. If Clearpath prompted "try the next quiz without read-aloud," it would be prompting non-implementation and creating a record of it.

That means valid with-versus-without data comes from only four places. The first is supports a teacher is trialing that are not in the IEP, which teachers can generally offer in class. Here an alternating design is both legal and strong. The second is comparing one bundle against another (read-aloud versus answer-orally-first, both permitted). The third is optional supports the student sometimes chooses not to use. The fourth is scores from before the accommodation entered the plan, which are confounded by time and practice and must be labeled as such. For everything else, meaning most of a typical student's accommodations, the honest output is fidelity plus outcome trend: "Read-aloud provided on 6 of 6 assessments; scores relative to class have risen." No effect estimate should appear.

The design literature agrees on which comparisons are strong. The What Works Clearinghouse classes a single before/after (AB) design as not meeting standards. A causal claim requires at least three demonstrations of effect at different times, at least 3 data points per phase to count at all, and 5 per phase to meet standards without reservations. An alternating-treatments design needs 5 repetitions of the alternation (WWC SCD Technical Documentation). Accommodations are unusually well suited to alternating comparisons because they act on access rather than learning. Their effect should appear immediately and vanish when removed, unlike instruction, which does not reverse. So alternation is the strong design, and it belongs to trial supports. AB phase change is the weak design, and it is all that exists for mandated accommodations. The product should be explicit about which one a teacher is looking at.

Mandated accommodations still have a legitimate evidence path: to the team. IEP teams must revise the IEP "to address any lack of expected progress" at least annually (34 CFR 300.324). NCEO says formative evaluation of accommodations "is not the responsibility of just one individual," and warns against assuming the same accommodations remain appropriate year after year (NCEO Accommodations Manual). Clearpath's job is to assemble the evidence that meeting needs, and to leave the decision to the team.

Teachers will log one tap per student, so that tap must be "used?"

Any new field competes with paperwork teachers already resent. The federal SPeNSE benchmark found special educators spending a median 4.7 hours a week on paperwork versus 1.6 for general educators (CRS RS21226). In a 2023 survey of more than 900 Minnesota special education staff, 98.4% were moderately to extremely frustrated with paperwork and 20% spent more than 9 hours a week on it outside the classroom. Their chief complaint was that it takes time from students (Education Minnesota). Data collection methods are non-standardized, and lack of time, resources, and training are the main barriers. In high school, "progress data" mostly means grades (Swain et al. 2022, ERIC). The formats teachers build for themselves set the effort ceiling: weekly checkbox grids of student by accommodation, and sticker labels on student work (Celavora). State manuals use the same pattern, logging each accommodation weekly as refused, provided but not effective, or very effective (Colorado CDE).

The research shows which tap matters most. Assignment is not delivery, and delivery is not use. Teachers are "inconsistent in providing" accommodations (BRT Tech Report 09-05). Students often skip them. Among students with disabilities eligible for extended time on grades 3–8 state tests, relatively few used more than a typical amount of time (Witmer 2024). On a low-stakes math test, fewer than half of eligible students with ADHD used extended time, yet access to it, not use, was associated with higher completion (Bernard & Witmer 2025). IRIS warns that when fidelity is low, you cannot tell whether an intervention is working (IRIS EBP module). A comparison built on "assigned" rather than "used" mixes up those cases. A per-accommodation "used?" mark, defaulting to provided, is therefore the single most valuable field Clearpath can add. It is also the one field that turns natural variation, such as a student declining an optional support, into legitimate comparison data. Free-text notes on perceived benefit should stay optional and out of the analytics, since teachers' perceptions of benefit are unreliable.

Data is also more likely to be collected when it feeds something required. Teachers object most to paperwork that does not change services (Education Minnesota). Every comparable product leans on automatic graphs and progress-report generation to justify entry (IEP Insight). Clearpath should capture fidelity at the moment of score entry, which the teacher is already doing, and pay it back as the evidence page for the IEP review and the parent progress report required under 34 CFR 300.320(a)(3).

The competitive field confirms both the gap and the safe primitives. Panorama, FastBridge, SpedTrack, and EDPlan all share one model: a per-goal or per-intervention time series with an aim line, a trend line after about six points, traffic-light status, and vertical markers where an intervention starts (Panorama PM graphs; SpedTrack). Motivity separates intentional phase lines, which break the data path, from dotted event lines for things like absences or medication changes (Motivity). HiRasmus users have asked for exactly that separation (HiRasmus ideas). Accommodation logging exists in AbleSpace and Aequitas, but none of the products examined links accommodations to outcomes (AbleSpace; Aequitas). Clearpath already knows which version each student took, a condition label other tools lack.

Show every point, gate on n, and say it in words

Teachers find the core task here especially hard. In think-alouds with 23 teachers, the most difficult part of reading CBM graphs was interpreting relations between elements, especially comparing data across adjacent phases, and linking data to instruction (van den Bosch et al. 2017). Years of CBM experience did not predict comprehension, though short instruction improved it (same). Even trained visual analysts agree only about .76 of the time (Ninci et al. 2015 summary). Teachers "often do not respond to the data, at least not without decision-making supports," yet students improve when teachers do act on it (PMC6428331). A graph alone is not enough. Each comparison needs a sentence that states n, the medians, the overlap, and how strong the evidence is.

The chart itself should show every score as a dot, because a bar of averages hides the distribution when samples are small (Weissgerber et al. 2015). The primary view is one student over time, with scores relative to the class main-version median. Dots are filled when the accommodation was used and hollow when it was assigned but not used. A labeled vertical line marks the date an accommodation entered the plan, a per-phase median tick shows level, and a shaded range shows spread. Tailored-version scores go in a separate panel plotted against the IEP objective's criterion, never on the same axis. A used-versus-not-used comparison is a two-column strip of dots with median ticks. Error bars, confidence intervals, and p-values should be left out entirely, since with 3–5 points they mislead more than they inform.

Gates should follow the borrowed single-case and CBM thresholds. Below 3 scores per condition, phases "cannot be used to demonstrate existence or lack of an effect" (WWC), so the tool shows dots only, with no difference, arrow, or verdict. At 3–4 per condition it shows an "early signal" with ranges visible. At 5 or more it shows a median difference. Trend lines wait for about 8 points, matching CBM protocols that allow the four-point rule after about 6 points and trend analysis after about 8 (Kentucky decision rules). These thresholds come from research-design convention, not from any study of teacher-made classroom tests. And because unit assessments are far less frequent than weekly CBM, five alternations can take a semester. Teachers should be told this up front so that "not enough data yet" reads as normal rather than as a failure.

Wording carries most of the feature's value and most of its risk. Teacher notes and pattern outputs tied to a student are likely education records that parents can inspect, and agencies must answer "reasonable requests for explanations and interpretations" (34 CFR 300.613). Every string should be written as if a parent will read it. It should frame results as access gains, not deficits, consistent with the IEP's required consideration of "the strengths of the child" (34 CFR 300.324). It should never use causal or directive language about IEP accommodations.

SituationUse this wordingNever
Under 3 scores per condition"Not enough data yet. 2 scores with read-aloud so far; comparisons start at 3."Any difference, arrow, or color verdict
Early signal (trial or optional support)"With read-aloud: 3 scores, median 6 points above the class median. Without: 4 scores, median 9 points below. 1 of 3 overlaps. Early signal; 2 more alternations would make this clearer.""Read-aloud raised Noah's scores 14 points"
IEP-mandated accommodation"Picture cues provided on 6 of 6 assessments, used on 5. Scores relative to class have risen since October.""Picture cues are working" or "aren't working"
Bundled accommodations"These scores reflect Noah's usual accommodations together (read-aloud, extended time, picture cues); the data can't separate them."Crediting a single accommodation
Extended time often unused"Used extra time on 1 of 5 assessments.""Extended time unnecessary"
Handoff"Bring to Noah's IEP review: scores and delivery record for read-aloud.""Consider removing read-aloud"
Before/after only"Scores since the plan changed in October. Other things changed too (new units, more instruction), so this is a pattern, not proof."Before/after framed as an effect

Privacy and framing guardrails are product requirements

Per-student scores linked to accommodation status are FERPA education records that reveal disability status. Clearpath can receive them as a school official only under the district's direct control. Access must be limited to staff with a "legitimate educational interest" in each student (34 CFR 99.31). In practice, the general educator, the case manager, and an aide should see different students. IDEA adds its own requirements: confidentiality at every stage, training for everyone who uses the data, and destruction "at the request of the parents" once the data is no longer needed (34 CFR 300.623; 34 CFR 300.624). That requirement extends to derived artifacts and logs. California's SOPIPA bars building student profiles for non-school purposes (Cal. B&P §22584). New York's Ed Law 2-d requires data security plans in contracts (NYSED). Together they rule out cross-district "what works" models built from pooled data without contractual authorization or proper de-identification. If an LLM ever drafts summaries, it is a subprocessor that must be covered by the district agreement, with no training on the data and pseudonymized inputs. Disability status is also inferable from version assignment alone, so any class or school view needs small-cell suppression, and no student's view should reveal a classmate's accommodations.

Equity concerns point the same direction as the statistical ones. Modifications "can increase the gap" with grade-level expectations and reduce opportunity to learn (NCEO Accommodations Manual). The tool should therefore never suggest a tailored version or fewer items because scores are low. NCEO warns against choosing accommodations by category or checking boxes "to be safe," and CDT urges developers to "test tools for disability-related bias" (CDT testimony). Suggestions drawn from disability-category patterns ("students with autism benefit from...") are exactly what reviewers flag. Pooled rankings of accommodations by average gain would mislead for the same reason the research is inconsistent: students differ.

Prioritized recommendations for a minimal feature

The minimal feature has three layers. A fidelity record is captured during grading. A per-student evidence page applies gates and plain-language statements. A trial path, limited to non-IEP supports, allows alternating comparisons. Everything else should wait.

PriorityBuildDetails
P0: captureVersion type and assigned accommodations, auto-filledIEP accommodations are pre-attached and locked on for that student's version; nothing new to type
P0: capture"Used?" per accommodation per assessmentOne tap, defaulting to provided and used. For teacher-controlled presentation supports (read aloud, one question at a time, picture cues) it collapses to "provided"; a separate "used" mark matters for student-controlled ones (answer orally first, optional trials). Extended time is recorded as none / some / most. If assessments are digital, time-on-test is captured automatically
P0: capture"Not provided — reason"Saved as an implementation gap, shown to the case manager, never treated as a "without" data point
P0: captureAccommodation start dates and event markersPhase line when an accommodation enters the plan (from IEP dates). Optional dotted event marker (absence, new unit) that does not split the series
P0: displayStudent accommodation recordFidelity summary ("provided 6/6, used 4/6") plus a dot series relative to the class main-version median, with used/unused fill and phase lines. Tailored scores in a separate panel against the objective criterion
P0: displayGated plain-language statementsDots only below 3 per condition; "early signal" at 3–4; median difference at 5+; trend at 8+. Bundle caveat whenever accommodations always co-occur
P1Trial supports (non-IEP) with alternationTeacher marks a support as "trying"; Clearpath suggests alternating it across comparable main-version assessments and shows the DATA-style strip comparison with the same gates. Also used for bundle-vs-bundle comparisons
P1Used-vs-not-used for optional supportsSame comparison, built automatically from natural variation when the student declines
P1IEP review / parent exportDelivery record, data points, and the statements above, on one page; satisfies 300.320(a)(3) reporting and 300.613 explanation requests
P2Objective probe decision rulesFour-point rule status after 6 probes, trend vs. aim after 8, with phase lines for instructional changes (the DBI loop)
P2Construct-conflict flagTeacher can mark when a support (read-aloud on decoding, a cue that reveals the answer) changes what's measured; that version then moves to the tailored track

Several things should not be built. No "try without" prompts for IEP accommodations. No "consider removing" or "drop" suggestions. No accommodated-versus-class-average delta labeled as effectiveness. No single-number headlines like "+17 pts with picture cues." No rankings or averages of accommodations across students, classes, or districts. No differential-boost calculation. No p-values or confidence intervals. No combined grade average across main, accommodated, and tailored versions. No suggestion of modified versions or fewer items driven by low scores. No suggestions based on disability category. No LLM-written causal narratives. No pooled cross-district models.

Conclusion

The product insight here runs against the usual analytics instinct. The scarce, valuable data is not more scores. It is the one bit per accommodation per assessment that says whether the support actually happened. Researchers studying extended time reached the same conclusion when they started measuring time actually used. Once that bit exists, the same record serves fidelity documentation, the IEP team's annual question, parent reporting, and, in the narrow cases where it is legitimate, a real within-student comparison. Without it, every chart conflates "assigned" with "helped."

The honest ceiling is also lower than "learn what works." For most mandated accommodations, Clearpath can show only that a support was delivered and how the student's relative standing moved afterward. The decision belongs to the team. Where the product can go further is with trial supports and bundle comparisons, because there teachers are free to alternate. That area, which no current tool serves, is where Clearpath can earn its claim to help teachers learn which methods work for which students, one alternation at a time.