Darbe Schlosser, MA
Founder, Think & Move Well
Darbe Schlosser · Think & Move Well
Self-generated cueing series, paper 3 · Version 1.0, 2026-09-09 All corpus figures cite Measurement Standard v1.1 (306 clients / 1,580 coaching sessions).
What this paper is, and what it is not
This is a descriptive analysis of observations from Think & Move Well coaching practice data. It is not an efficacy trial. It establishes no learning, no retention, no transfer and no treatment effectiveness, and it makes no claim of any of them. The corpus contains no controlled comparison and cannot support one. Please read the numbers below with that boundary in place, along with the limits stated alongside each of them.
This paper describes a training architecture. It does not evaluate one.
Think & Move Well coaching teaches people with Parkinson's disease a digit sequence and attaches motor rules to it. The sequence persists across sessions and is also extended — it is not fixed, and this paper's own results record it being added to (C2, 137 clients). Across a de-identified corpus of 1,580 coaching sessions with 306 clients, five validated constructs measure components of that architecture.
The ordering of construct reach is the finding — and it is a fact about the constructs before it is a fact about the practice. A familiar scaffold carrying newly introduced material (C4) is detected for 267 clients; established digit recall (C1) for 115. That is 2.3 times the client reach — reach meaning detected at least once, which is neither a frequency nor a share of practice time. The two constructs also have very different detection surfaces, so the ratio must not be read as a comparison of how much of the coaching each represents.
Alongside those measurements, three individually anchored cases describe a phenomenon the corpus cannot currently measure: a client not producing material that was in use moments earlier, when the demand on the task changed. These are illustrations, not a rate. The proxy query for this phenomenon failed validation, so it is carried qualitatively and nowhere counted.
No claim of efficacy, learning, retention or transfer is made or implied. The corpus contains no controlled comparison and cannot support one. Motor-learning convention reserves learning for skill that is retained after a period without practice, and uses performance for change observed within a session (Hayes et al., 2015) — everything described here is on the performance side of that line, and none of it is offered as evidence of durable learning. The paper's contribution is a description of what the training consists of, and an honest account of one part of it that is visible but not yet measurable.
Self-generated cueing has an evidence base in Parkinson's disease. People with PD, with and without freezing of gait, can use both externally generated and self-generated cues to modify gait; across groups, external cues increased gait variability while self-generated cues did not (Harrison et al., 2020). Motor-cognitive training improves dual-task gait speed, with GRADE certainty rated high for dual-task gait speed, cadence, step length and dual-task-cost reduction (Johansson et al., 2023). Load-driven freezing has a described neuroimaging correlate in decoupling between cognitive-control networks and the basal ganglia (Shine et al., 2013).
What that literature does not settle is the question this paper sits inside: what these coaching sessions contain, and what happens to a strategy when the demand on it rises. The cueing literature names these as open on its own account — the Earhart-group work lists, as future research, durability after cue withdrawal, retrieval under real-world stress, and whether the capacity to generate a cue can be trained rather than supplied (Harrison et al., 2020).
That is evidence the question is open. It is not evidence that this paper is first to it, and no such claim is made here. Related work exists and is discussed below; what the citation establishes is that the field regards the matter as unresolved.
Paper ① of this series claims that retrieval under load decides whether a cue survives outside the clinic, and that retrieval is trainable. It states that claim and does not test it. This paper does not re-argue it. Its job is narrower and prior: to describe the architecture — what is held constant, what is varied, in what order — and to describe, without quantifying, what it looks like when a strategy that was available a moment ago is not produced.
That a person can perform a movement or produce a sequence is not the same as their producing it at a particular moment, under particular conditions.
This distinction is established, and it is not this paper's. In Parkinson's disease specifically, Lee and colleagues have shown that learned motor behaviour is context dependent: participants practised three finger sequences, each embedded in an incidental context, and performance worsened the next day when the sequence–context associations were changed (Lee et al., 2019).
Three features of that work bound what may be inherited from it:
What this paper describes is related but not equivalent, and the difference is the manipulation. Lee changes the context a sequence was learned in — colour, screen location — while the task itself is unchanged. The cases here involve changes in task demand: a postural correction with a rep count, ten newly learned digits alongside an alternating movement, and two light hand weights.
No claim is made that manipulating task demand is novel. It is a description of what these coaching sessions contained. The corpus has no validated measure of whether a strategy was reachable at the moment it was needed, and three described moments are not a measurement.
De-identified coaching transcripts held in the Think & Move Well knowledge engine. Analytic corpus: 1,580 sessions across 306 clients (Measurement Standard v1.1, adopted 2026-08-18). The client denominator is derived by classifying every roster entry and counting only those representing people — synthetic session keys, device labels and placeholder speaker labels are excluded — intersected with clients actually appearing in the indexed corpus.
Every count in this paper is a floor. Precision was measured for each construct; recall was never measured. Three exclusions of unknown or unassessed size apply: private meetings, which the recording platform's organisation-scope API structurally excludes and whose number is never estimated; 199 sessions held at the de-identification gate; and whatever a paused coverage audit would find. No completeness claim is available for any client or any month.
Five constructs are validated for automated corpus-wide counting, plus a gate applied before every category. Each was validated by drawing a seeded sample, reading every passage by hand, and computing precision; a query was promoted only at ≥0.80. Precision values are inherited from Measurement Standard v1.0 and were not re-measured under v1.1 — the amendment changed the corpus, not the instrument, and v1.1 must never be cited as a revalidation.
Two further categories are reviewer-scored and may never carry counts: self-initiated use outside sessions, and ambiguity. Both fail as retrieval targets for structural reasons, not fixable ones.
The candidate proxy query for failure-to-retrieve did not meet the 0.80 promotion threshold and is not used for counting anywhere in this paper.
On prior review, of 20 sampled passages 17 were on topic, but roughly half also matched administrative talk — validity approximately 50%. The query does not reliably separate a person failing to reach a strategy from ordinary forgetting, from scheduling conversation, or from a coach or a device being described as blank.
Independent inspection during case selection for this paper reproduced that figure. A broad search returned 364 candidate passages; tightening to explicit non-retrieval language left 84. Inspection of those found the coach blanking, a client's computer blanking, a note-taking application blanking, scheduling exchanges, and in-session slips on a movement instruction — which is performance failure, not failure to retrieve.
Therefore failure-to-retrieve is carried in this paper only through individually anchored qualitative cases, and carries no corpus-wide count, no prevalence, no denominator-based estimate and no frequency language of any kind. Formal corpus quantification would require a new measure validated under the seven-step protocol at ≥0.80, reporting its precision, and an amendment to the Measurement Standard. That is future work and is not this paper.
Cases were selected for representation of distinct failure shapes, not for frequency and not for drama. Each is one real de-identified client; clients are never merged. Candidates were rejected where the evidence was ambiguous rather than forced into classification — including one client's failure to reproduce newly assigned material, excluded because nothing established it had ever been available to lose. The number of cases presented reflects the number of distinct things that needed showing.
Digit-structured work is detected for 281 of 306 clients (91.8%), across 1,463 sessions and 5,676 passages. This is the gate applied before every construct below.
Each entry reads: passages / sessions / clients — share of the 306-client denominator — measured precision — number of hand-read passages the precision was measured on.
Client share and session counts are different denominators and are never compared to each other.
Construct names are conceptual labels for what each query detects — not direct measures of a client's experience, understanding, learning, or successful execution. C4's detector, for example, includes four ordinary conversational phrases behind the digit gate, and was validated on six hand-read passages.
Precision was measured on samples of six to nine hand-read passages. A precision of 100% on six passages is a small-sample result and should be read as such — it establishes that the query was not obviously wrong, not that it is exact. The sample size is reported here because a bare precision figure invites more confidence than these samples can carry.
C4 is detected for 2.3 times as many clients as C1 (267 against 115). Extending the sequence itself (C2, 137) and retrieving the rule attached to a digit's value (C3, 127) sit between them.
This is construct reach, and reach is a weak quantity. It must not be read as any of the following.
The two constructs being compared have asymmetric detection surfaces, and the asymmetry runs in the direction of the result. C1 matches eleven literal numeral strings, and its recorded false negatives are silent recall and runs that the transcription splits into isolated digits — so recitation that happened can systematically fail to register. C4 matches ten natural phrases, four of which are ordinary conversational language. A construct with a wider linguistic surface will reach more clients than one keyed to literal digit strings, whatever the practice actually contains.
What survives that caution is narrow and still worth stating: across this corpus, coaching language indicating that an existing structure is carrying newly introduced material is detected far more widely than transcribed recitation of the sequence itself. Whether that reflects how the coaching is distributed, or how the two constructs behave, is not resolved by these numbers. (expert-interpretation)
Nothing in the ordering is evidence that the training works. It says nothing about outcome, and nothing about whether any client improved.
They are not a prevalence of Parkinson's phenomena. They are detections of coaching activity in transcribed speech, and the governing lesson of the measurement work applies without exception: the vocabulary of a method is not evidence the method is being performed. A construct detects that something was said, at a measured precision. It does not establish that a client experienced anything.
Percentages under v1.1 are slightly lower than under v1.0. That is denominator and capture, not behaviour — every numerator grew, and the denominator grew faster, because the corpus expansion added disproportionately many clients with few captured sessions. A client with one captured session has fewer opportunities to be detected by any construct.
Each case stands alone. They are not summed, and their number carries no information about how often this happens. Full anchored versions accompany this paper.
With a postural correction added to the stepping pattern and the rep count set at four, the digit sequence was not produced. The client reported going blank. He asked not to be told, and the coach withheld the cue — he worked it out aloud and arrived at the correct run, which the coach confirmed. It went blank a second time on a later rep and was again recovered by the client. After the coach halved the rep count, execution was clean and continuous. (Case C127, session of 2025-09-24, t=06:07)
What this shows: on this occasion, material in use moments earlier was not accessible, and returned without the answer being supplied. What it does not show: that the added demand caused the failure, that reducing reps restored performance, or anything about learning, retention or mechanism.
Within a sequence he was producing correctly at length, place-keeping failed: he queried his position, was told he had already covered that portion, acknowledged it, and shortly afterwards reported a blank. He resumed and completed. The coach then named two concurrent demands that had just been added — ten newly learned digits, combined with alternating unilateral movement. The coach supplied position, not content; the sequence itself was produced by the client. (Case C057, session of 2025-10-11, t=28:37)
What this shows: a distinct retrieval problem — the sequence was available; the position within it was not. What it does not show: that those demands caused the lapse. The coach's attribution is a coaching statement made in the moment, not a measurement.
Practising curls and rows with the sequence spoken alongside, and holding two light hand weights, he lost track of the digits. He raised it himself, unprompted, identified the weights as the only change, and reported the effect as unexpected. The coach observed that the sequence had not been a problem on the other movements. He proposed practising that specific combination. (Case C022, session of 2025-07-25, t=00:00)
Anchor limitation, stated rather than hidden. This session's timestamps are collapsed in the source — 13 chunks carry 2 distinct values — so the anchor resolves to the chunk, not to a minute. The passage runs from line 96 to line 100 of the transcript file. This is a property of the recording, not of the evidence.
What this shows: a sequence holding during unweighted movements did not hold once light weights were added, by client self-report and consistent with the coach's observation. What it does not show: that the weights caused it, that the effect is specific to weight, or that practice resolved it — no follow-up is cited, and whether it returned is not established.
The measurements and the cases describe two different things, and the paper keeps them apart deliberately.
The constructs describe an architecture. A persisting structure has new material attached to it, and coaching language for that is detected across more clients than transcribed recitation of the structure itself. That is a claim about what is detectable in coaching transcripts — not about how the coaching is distributed, because the two constructs do not detect equally well. (expert-interpretation)
The cases describe moments of access. In each, material demonstrably in use was not produced when something about the task changed — a postural correction and a rep count, ten new digits with an alternating movement, two light hand weights. In three different task contexts, three different things about the demand changed, and three different retrieval problems appeared: the sequence not coming at all, the position within it being lost, and the hold on it being displaced.
That variety is the point, and it is also the limit. It is the reason the phenomenon resists a single query — the surface language of going blank covers at least three distinct failures, and shares that language with a coach blanking and a computer blanking. A measure that counted all of them would be measuring a word.
The general distinction is inherited from the context-dependency literature (Lee et al., 2019), not drawn here. What is this paper's own is narrower: a description of the same family of phenomenon arising when task demand changes rather than when the learned context changes — and that narrower version has no validated measure in this corpus.
C127 is the clearest instance available: the material was not produced, the cue was declined, the coach withheld the answer, and the client recovered it unaided. What that supports is limited and worth stating exactly: absent production at one moment did not, on this occasion, mean absent capability. It does not demonstrate a retrieval mechanism, and one case is not a demonstration of anything.
(This section is expert interpretation. Nothing in it is a measured finding.)
A meta-analysis of 15 studies (299 people with Parkinson's disease, 244 healthy controls) found impaired sequence-specific implicit motor performance in Parkinson's disease: standardized mean difference 0.83, 95% CI 0.30 to 1.36 (Hayes et al., 2015).
The interval matters as much as the point estimate. A lower bound of 0.30 is a small effect, and heterogeneity was high (I² = 87%) — so the pooled figure is a central estimate across markedly varied studies, not a precise quantity. Two scope limits come from the paper itself: 14 of the 15 studies used a serial reaction time task, and the authors state the results apply only to discrete, not continuous, tasks.
The challenge is direct and this paper does not soften it: the method's organising structure is a learned sequence, and this meta-analysis finds implicit, sequence-specific learning impaired in exactly the population the method serves. The qualifier is load-bearing and is kept everywhere: the finding is about implicit acquisition measured in serial reaction-time paradigms, not about sequence learning in general. A skeptical reader is entitled to ask why a method organised around a learned digit sequence should be plausible at all.
Three things must be distinguished in answering, and only the first two are on firm ground.
A · What Hayes et al., 2015 challenges. It challenges any expectation that a person with Parkinson's will pick up a motor sequence implicitly at a rate comparable to healthy adults. Interventions depending on sequence acquisition should expect a harder curve and should not assume implicit acquisition will carry the learning.
The authors bound their own conclusion, and the bound is load-bearing here. All but one of the included studies ran a single day of practice. Because motor-learning convention reserves learning for skill retained after a period without practice, they state that these studies "have only examined changes in motor performance," and that any assertion the results reflect persistent change "should be interpreted with caution." Their Conclusions call for research to determine whether the deficit is one of performance or of learning.
This does not dissolve the challenge — it locates it. The impairment is documented in within-session acquisition. Whether it extends to sequence knowledge built across many sessions with retention between them is not addressed by the included studies, and this paper does not assume either answer.
B · What this paper actually claims. Nothing about acquisition. This paper describes what is present in coaching sessions and what three moments of non-retrieval looked like. It makes no claim that the sequence is learned efficiently, learned implicitly, learned at all in a measurable sense, or learned better than by any alternative. The corpus contains no acquisition curve, no control group and no comparison. On the question the meta-analysis actually addresses, this paper is silent, and its silence is not an answer.
C · What remains hypothesis. The method's stated response — that the sequence is taught explicitly and chunked mnemonically rather than acquired implicitly, and may therefore sidestep the documented implicit deficit — is on record in this series as Hypothesis 1, with a stated falsification condition: it is falsified if acquisition rate tracks implicit sequence-learning measures (Counting as Coaching, earlier in this series). It is a hypothesis, not a demonstrated advantage, and nothing here resolves it.
The meta-analysis's own discussion raises explicit strategies — and that must be reported at its weakest defensible strength.
Its authors write that recognising the implicit deficit lets clinicians identify alternative resources that may augment motor learning, "such as providing explicit instruction, additional verbal cueing, or additional practice."
What that is: the authors themselves discussing explicit and verbal strategies as potential clinical responses to impaired implicit sequence performance.
What that is not, and the distinction matters: the meta-analysis does not validate the method described here, does not test a digit scaffold, and does not establish that explicit instruction bypasses the deficit it documents. No included study compared explicit instruction against implicit acquisition. A suggestion raised in a discussion section is not a test of that suggestion, and this paper does not treat it as one.
Two cautions on Hypothesis 1, both of which weaken it.
First, explicit teaching does not exempt a method from the finding. That a sequence is taught explicitly says how it is delivered, not how it is consolidated. Whatever consolidates it may still be the impaired system.
Second, the included studies measured explicit knowledge as an outcome, not as an instruction method — 11 of them reported no more than half of participants gained explicit knowledge of the repeating sequence. That is not evidence about teaching explicitly.
Hypothesis 1 remains untested, with its falsification condition intact. The challenge stands.
The case set does not resolve any of this, and may not be offered as if it did. Two of the three cases are neutral to the meta-analysis; C057 arguably complicates it, because the sequence there is explicitly recited — but that observation is descriptive, concerns one moment, and is evidence about a moment rather than about a learning system.
A narrative review reports that cue and attentional strategies have only short-term effects on gait, with stride length reverting once attention lapses (Wu et al., 2015; cited at abstract level) — relevant because this paper describes training and says nothing about durability. Dual-task training's benefits are established for gait outcomes this paper does not measure (Johansson et al., 2023); that evidence supports the general rationale for training under cognitive load and is not evidence for this method.
Think & Move Well coaching uses a digit sequence that persists across sessions — and is also extended — while varying the motor problem attached to it. Across 1,580 sessions with 306 clients, coaching language for a familiar structure carrying newly introduced material has the widest client reach of the five constructs, 2.3 times that of transcribed recitation of the sequence. Reach is not frequency, not a share of practice, and not comparable between constructs that detect unequally.
Alongside it sits something visible and unmeasured: material in use one moment and not produced the next, when the demand on the task changed. Three anchored cases describe three different versions of it. They are illustrations, and this paper counts nothing about them.
The contribution is a description of a training architecture and an honest account of a phenomenon inside it that the corpus can show but cannot yet measure. Whether any of it helps anyone is a different question, requiring a different design, and is not addressed here.
Read in full for this paper: Hayes et al. (2015) and Lee et al. (2019). Siegert et al. (2006) is cited from its abstract, and is noted as such where it appears.
Also referenced: Counting as Coaching, the earlier paper in this series, for the explicit-scaffolding hypothesis and its falsification condition; and Think & Move Well's internal storyboard digit-teaching hypothesis record.
Who Holds the Cue states the claim this paper deliberately does not test — that retrieval under load decides whether a cue survives outside the clinic, and that retrieval is trainable. This paper describes the training and the failure; it does not test that claim.
This paper describes coaching practice and observations from a de-identified coaching corpus. It is not medical advice, does not describe a treatment, and makes no claim of clinical benefit. Nothing here should inform a change to any person's medical care. Decisions about Parkinson's disease management belong with the treating clinician.