1. Introduction
It has been generally assumed that English /l/ is phonetically implemented as a lighter or darker variant depending on its position within the syllable, with a fronter, “clear” realization at syllable onsets and a backer, “dark” or velarized realization at syllable rimes. The gestural account of this variation attributes it to the differential timing of two component gestures of /l/—an apical gesture and a dorsal gesture—whose relative phasing is contingent upon prosodic position (Sproat & Fujimura, 1993). Within Articulatory Phonology, this timing relation is further understood as gestural competition and overlap, such that the phonetic output at any given point reflects the blending of co-active gestural targets rather than a discrete switch between categorically distinct segments (Browman & Goldstein, 1989, 1992).
A great deal of research has been conducted on how dark /l/-velarization becomes as a function of prosodic strengthening, boundary type, and adjacent context, and this is generally captured through continuous acoustic parameters such as F2, F2-F1, and the rate of their change across the /l/-segmentation. Comparatively less attention, however, has been given to a related but distinct question: whether the point that velarization is headed for—here termed the acoustic endpoint of the dorsal gesture—is fixed regardless of context, or whether it varies systematically depending on the vocalic environment preceding the lateral.
This paper reports a statistical reanalysis of a descriptive observation first made in Lee (2023), using the same acoustic data from American English speakers. When the steady-state and offset positions of word-final /l/ were plotted on the F2-F1 plane, it was observed that /l/s following front vowels (/i:, e, eɪ/) came around an area close to word-final schwa (cf. Flemming, 2009), whereas /l/s following non-front vowels (/ɔ:, oʊ, u:/) converged toward a separate, more peripheral area. Taken closely, this dichotomy did not simply follow raw acoustic proximity among the six vowels: on the specific measure used in Lee (2023)—F2-F1 distance at the steady state—/ɔ:/ was numerically closer to /e/ than to /u:/, yet patterned acoustically with the non-front group at the offset rather than with /e/; this vowel behaving “out of place” relative to that specific distance measure is not readily explained under a purely gradient, single-target account of coarticulation, and it is this asymmetry that motivated the present reanalysis.
On the grounds of the above, this paper puts the observation reported in Lee (2023) to formal statistical test. Specifically, the following issues are addressed.
First, this paper examines whether the steady-state and offset acoustic endpoints of word-final /l/ differ systematically according to the frontness of the preceding vowel.
Second, this paper directly tests whether this asymmetry differs between the steady state and the offset, rather than inferring a difference from separate significance tests at each point, and what the resulting pattern, if any, implies about the acoustic organization of /l/-velarization.
This paper departs from Lee (2023) in three respects. First, treating individual tokens as independent, as in the original descriptive comparison, is not appropriate given that measurements were taken repeatedly from the same speakers and the same vowels; the primary analysis reported here therefore uses linear mixed-effects models with crossed random intercepts for speaker and vowel. Second, the frontness-by-measurement-point interaction is tested directly, rather than compared informally across separate tests. Third, because the acoustic endpoints examined here do not uniquely identify an articulatory target, the results are reported and interpreted primarily as endpoint positions, with a gestural account presented as one possible, but not the only, interpretation.
The remainder of this paper is organized as follows. Section 2 reviews previous studies on the gestural organization of English /l/ and on target-driven as opposed to purely gradient accounts of coarticulatory variation. Section 3 describes the subjects, materials, and acoustic measurements used, together with the statistical procedures applied to the two issues raised above. Section 4 presents the results of the mixed-effects and token-level analyses. Section 5 discusses these results, addresses alternative explanations, and considers the relation between acoustic endpoint and articulatory target. Section 6 concludes with the limitations of the present paper and implications for further articulatory investigation.
2. Theoretical Background
English /l/ has traditionally been treated as having a small number of contextually determined allophones—a fronter, “clear” allophone in syllable onsets and a backer, “velarized” allophone in syllable rimes (Sproat & Fujimura, 1993). Sproat & Fujimura (1993), however, argued against treating these allophones as categorically distinct phonological entities, and proposed instead that /l/ is composed of two co-occurring gestures, namely an apical (consonantal) gesture and a dorsal (vocalic) gesture. It was further proposed that the vocalic gesture has a strong affinity for the syllable nucleus, whereas the consonantal gesture has a strong affinity for the syllable margin.
This gestural view was taken further within Articulatory Phonology (Browman & Goldstein, 1989, 1992), in which gestures are treated as the basic units of phonological contrast, each defined as a dynamically specified vocal-tract task with an intrinsic target and duration. Applied to /l/, this means that the degree of velarization observed at any given point of the segment emerges from the moment-to-moment competition and blending between the apical and dorsal gestures, rather than being read off a single phonological specification.
Building on this gestural view, Huffman (1997) examined intervocalic onset /l/ following schwa, a context in which the timing account of Sproat & Fujimura (1993) predicts little variation in backness. It was found, however, that onset /l/ following schwa does vary considerably in backness, and that this variation is related to gestural timing in a way not fully anticipated by the original affinity account. This finding suggests that the acoustic endpoint the dorsal gesture is drawn toward is not wholly independent of the surrounding vocalic context, even where prosodic position is held constant.
A separate but related line of research bears on whether phonetic variation of this kind is best understood as categorical or gradient. Lindblom (1963), in a classic study of vowel reduction, demonstrated that the acoustic realization of a vowel is systematically drawn toward the values of a preceding or following segment as duration is shortened, and interpreted this pattern as reflecting undershoot of a fixed underlying target rather than a change in the target itself.
Directly relevant to this question, Turton (2017) used ultrasound tongue imaging across nine varieties of English and found that /l/-darkening shows evidence of categorical allophony and gradient phonetic effects coexisting within the same grammar, rather than one pattern excluding the other. Related articulatory evidence for context-sensitive gestural targets is provided by Gick (2003), who found that the realized gestural target of a consonant is sensitive to its syllabic affiliation in ways not reducible to a single fixed target.
This raises the question of what an appropriate reference point for word-final /l/ would be. Word-final schwa has been reported to occupy a comparatively stable, mid-central region of the F1-F2 space (Flemming, 2009), and is accordingly a natural reference point against which the offset position of a velarized /l/ may be compared.
This convergence toward a schwa-like region for front-vowel contexts is not arbitrary. Gick & Wilson (2006) show, on cross-linguistic evidence, that when a vowel's fronting/advancing gesture conflicts with a following consonant's retraction gesture, the resulting compromise articulation is perceived as schwa-like—“a phonetic by-product of one strategy for reconciling an intrinsic conflict between articulatory targets” (Gick & Wilson, 2006:636). Applied specifically to English coda /l/, Stenbrenden (2024) argues that the resulting excrescent vowel resembles schwa or /ɔ/ precisely because these vowels share a post-oral gesture with the lateral's radical/dorsal component.
Finally, the degree of /l/-darkening is not fixed even within a single grammar, but varies gradiently with prosodic position; Sohn & Lim (2020) found that word-final /l/ is more strongly velarized at stronger prosodic boundaries for both native English speakers and Korean EFL speakers, with the two groups differing in degree but not direction.
Taken together, the gestural account of /l/ (Browman & Goldstein, 1989, 1992; Huffman, 1997; Sproat & Fujimura, 1993) and the undershoot account of coarticulatory variation (Lindblom, 1963) generate two distinguishable predictions for the acoustic endpoint of word-final /l/ across preceding-vowel contexts, and it is exactly this contrast that the descriptive observation reported in Lee (2023) speaks to. The present paper takes up this gap directly, treating the acoustic endpoint as the primary object of description and the gestural target as an interpretive layer over it, and is also in a position to speak to the more general question of whether coarticulatory variation of this kind is better characterized as gradient undershoot of a single endpoint or as variation between two acoustically distinguishable endpoints.
3. Method
Lee (2023) reports acoustic data from ten female native speakers of American English, recorded as part of a larger study spanning multiple prosodic and morphosyntactic positions of /l/. All ten were born and raised in the United States and were, at the time of recording, employed as English teachers at the Daegu Global Education Center under the Daegu Metropolitan Office of Education, having resided in Korea for two to four years; none showed evidence of Korean-influenced English, and none reported speech disfluencies. Of these ten speakers, seven provided complete, codable tokens across all six word-final vocalic contexts examined here (teal, tell, tale, tall, toll, and tool) and are accordingly included in the present analysis.
Word-final /l/ was elicited in six words, each instantiating a distinct pre-lateral vocalic context: teal /i:/, tell /e/, tale /eɪ/, tall /ɔ:/, toll /oʊ/, and tool /u:/. One token per speaker was obtained for each vocalic context, yielding 42 tokens in total (7 speakers×6 vocalic contexts). Each word was embedded in the carrier sentence “Say ___ for me” and recorded individually by each speaker in a quiet room; recordings were sampled at 44,000 Hz, 16-bit resolution.
Following the /l/-segmentation criteria used in Lee (2023), which in turn follow Sproat & Fujimura (1993), the /l/-segmentation was identified centering on rapid F2 movement: the onset of the /l/-transition was marked by a drop in intensity together with a loss of the clearly defined formant structure characteristic of the preceding vowel, and the offset was marked by the point at which apical occlusion is achieved (the end of the /l/-segmentation). F1, F2, and F3 were measured at three points—onset, steady state, and offset—with the steady-state measurement taken at the temporal midpoint of the /l/-segmentation's steady-state interval. All measurements were made in Praat v6.2.14 (Lee, 2023). The formant ceiling used in analysis was 5,000 Hz, as recoverable from the spectrogram display range in the original measurement records; the LPC order and analysis window length used are not separately documented and are noted as a limitation in Section 6. All formant frequency values were converted to the mel scale for the present analysis.
Given that repeated measurements were taken from the same seven speakers across the same six vocalic contexts, individual tokens cannot be treated as independent observations, nor can speaker and vowel/word be fully disentangled as separate sources of non-independence. The primary analysis therefore uses linear mixed-effects models, one for F2 and one for F1, with frontness (front vs. non-front), measurement point (steady state vs. offset), and their interaction as fixed effects, a random intercept for speaker, and a variance component for vowel (i.e., word/item). This structure follows standard practice for designs with repeated measures nested within both speakers and items, modeling both as random effects rather than analyzing them in separate, token-level and speaker-level, steps.
At the level of individual tokens, a one-way MANOVA was additionally conducted with front/non-front group as the independent variable and the four acoustic measures (F2std, F1std, F2off, and F1off) as dependent variables. A MANOVA, rather than four separate univariate ANOVAs, was used because the four acoustic measures are non-independent properties of the same /l/-segmentation; testing them jointly avoids the inflated Type I error rate that separate ANOVAs would produce, and evaluates whether the front/non-front groups differ in the acoustic pattern as a whole rather than on any single measure in isolation. Linear discriminant analysis (LDA) with leave-one-out cross-validation was used to assess token-level classification accuracy, both overall and broken down by vowel; and k-means clustering (k=2) was conducted on the standardized measures. As this analysis does not correct for the non-independence of tokens produced by the same speaker, it is reported as descriptive and exploratory, secondary to the mixed-effects analysis above.
Because a categorical front/non-front grouping is only one possible account of the acoustic pattern under investigation, a continuous coarticulatory alternative was also tested. For each of F2std and F1std, a mixed-effects model containing only continuous predictors—the /l/'s own onset formants (F2on, F1on), which closely track the preceding vowel's acoustic quality by coarticulatory continuity, and the preceding vowel's duration (Dur1)—was compared against a model that additionally included the categorical front/non-front term, both with random effects for speaker and vowel. The two models were compared via AIC, BIC, and a likelihood-ratio test for the added categorical term.
To evaluate whether a two-way grouping is empirically well-motivated relative to alternatives, silhouette coefficients (Rousseeuw, 1987) were compared across k=2, 3, and 4 cluster solutions. Variance inflation factors (VIF) were computed for the four acoustic measures to assess multicollinearity among the classification features used in the secondary analysis.
4. Results
This section reports the primary mixed-effects analysis (4.1), the secondary token-level analysis (4.2), and model-selection/robustness checks (4.3).
Figure 1 first situates the six vocalic contexts within the oral cavity by tracking the mean onset-to-steady-state-to-offset trajectory of /l/ for each context, oriented so that position in the F1-F2 plane corresponds intuitively to tongue-body position, following the display convention of Lee (2023). As shown there, the front-vowel contexts (/i:, e, eɪ/) trace trajectories that, on the F2 (frontness) dimension, remain fronter than the non-front contexts throughout; on the F1 (height) dimension, however, /e/ and /eɪ/ move toward a notably lower, more open offset position, whereas /i:/'s offset instead patterns with the higher, less open region occupied by /ɔ:/, /oʊ/, and /u:/—a dissociation discussed further in Section 5. Full descriptive statistics underlying this figure, including the onset values not used in the inferential analyses below, are given in Table 1.
Values in mel; SD in parentheses. F2on/F1on shown for reference (Figure 1); not used in inferential analyses.
Values in mel; SD in parentheses. F2on/F1on shown for reference (Figure 1); not used in inferential analyses.
Table 2 summarizes the linear mixed-effects models predicting F2 and F1 from frontness, measurement point, and their interaction, with random effects for speaker and vowel. For F2, the frontness effect was marginal (b=77.7, SE=40.1, p=.053); measurement point had a significant effect (b=117.0, SE=40.1, p=.004), reflecting an overall increase in F2 from steady state to offset across both groups; the frontness-by-measurement-point interaction was not significant (b=21.9, SE=56.7, p=.699). For F1, the frontness effect was significant (b=86.3, SE=34.6, p=.013); measurement point had no significant effect (b=–5.6, SE=34.2, p=.871); the interaction was again not significant (b=–13.4, SE=48.4, p=.783). Figure 2 plots the front and non-front group centroids at the steady state and offset directly in the F1-F2 plane (cf., Lee, 2023, figs. 5.9–5.10); Figure 3 presents the same comparison from the individual-speaker perspective.
Taken together, these models indicate a front/non-front asymmetry that holds overall—more clearly for F1 than for F2—but for which the present data provide no evidence of a reliable difference between the steady state and the offset, since neither interaction approached significance. This qualifies the descriptive impression, visible in Figures 2 and 3, that the front/non-front separation looks larger at the steady state than at the offset for F2 in particular: Table 2 shows that the F2 contrast in fact numerically increases at the offset, while the F1 contrast numerically decreases, and neither change is statistically reliable given the present sample size.
A further test asked whether the categorical grouping is in fact necessary, or whether a continuous coarticulatory account performs at least as well, following a reviewer's suggestion. Table 3 compares, for each of F2std and F1std, a mixed-effects model containing only continuous predictors (the /l/'s own onset formants, which closely track the preceding vowel's acoustic quality by coarticulatory continuity, and the preceding vowel's duration) against a model that additionally includes the categorical front/non-front term. For F2std, adding the categorical term did not significantly improve fit over the continuous-only model [χ² (1)=0.99, p=.319; AIC 489.6 vs. 490.6]; for F1std, the same was true [χ²(1)=0.09, p=.770; AIC 497.2 vs. 499.1]. Both comparisons favored the continuous-only model on AIC and BIC, indicating that, at least at the steady state, a continuous coarticulatory account is a more parsimonious description of the present data than the categorical front/non-front grouping emphasized in the Introduction.
Treating the 42 tokens as independent observations—a simplification not adopted in the primary analysis above—the one-way MANOVA on front/non-front group was significant [Wilks' λ=.658, F(4,37)=4.81, p=.003]. LDA with leave-one-out cross-validation classified tokens into front/non-front at 71.4% accuracy, against a chance baseline of 50%. As shown in Table 4, this accuracy was not uniform across vowels: /u:/ and /e/ were classified at 85.7%, /i:/ and /oʊ/ at 71.4%, while /eɪ/ and /ɔ:/—the two contexts adjacent to the front/non-front boundary—were classified at only 57.1%. Unsupervised k-means clustering (k=2) recovered the front/non-front partition at a comparable rate (30/42 tokens, 71.4% agreement). Figure 4 plots all 42 tokens in the F1-F2 plane at the steady state and at the offset, on identical axis scales in both panels, together with per-vowel and per-group centroids.
| Vowel | Group | Classification accuracy (%) |
|---|---|---|
| /u:/ | non-front | 85.7 |
| /e/ | front | 85.7 |
| /i:/ | front | 71.4 |
| /oʊ/ | non-front | 71.4 |
| /eɪ/ | front | 57.1 (boundary) |
| /ɔ:/ | non-front | 57.1 (boundary) |
Silhouette coefficients were .292 for k=2, .318 for k=3, and .317 for k=4 (Table 5), indicating that a three-cluster solution fits the token-level data marginally better than the theoretically motivated two-cluster solution, consistent with /ɔ:/ forming a quasi-independent third grouping. VIF values (1.10–1.64) indicate no problematic multicollinearity.
| Check | Metric | Value |
|---|---|---|
| Silhouette coefficient (k-means) | k=2 / k=3 / k=4 | .292 / .318 / .317 |
| Variance inflation factor (VIF) | F2std / F1std / F2off / F1off | 1.62 / 1.64 / 1.10 / 1.11 |
5. Discussion
Taken together, the results in Section 4 offer qualified support for the front/non-front asymmetry raised in Section 1, while also revealing a pattern more nuanced than a clean two-way split. At the steady state—the point of maximal constriction—the front/non-front distinction was directionally exceptionless across all seven speakers, though the multivariate test at the speaker level was only marginal (p=.067) given the small sample. At the offset, the same distinction was weaker and less consistent across speakers. This asymmetry is not naturally explained under a purely gradient, single-target undershoot account (Lindblom, 1963); it is more consistent with a competing-gestural-target account (Browman & Goldstein, 1989, 1992; Sproat & Fujimura, 1993), under which the offset is already subject to coarticulatory transition into whatever follows the /l/-segmentation—though, given the marginal statistics, this interpretation should be treated as suggestive rather than confirmed.
At the same time, the token-level analyses in 4.2 indicate that this asymmetry is not uniformly categorical across the six vocalic contexts. The two boundary contexts, /eɪ/ and /ɔ:/, were classified at close to chance level, while /e/ and /u:/ were classified with comparatively high accuracy. This is in line with the qualitative observation in Lee (2023) that /ɔ:/ patterns with the non-front group despite its acoustic proximity to /e/; the present results indicate that this is mirrored, in the opposite direction, by /eɪ/. The marginally better fit of a three-cluster solution (4.3) reinforces this picture, and raises the possibility that /ɔ:/ and /eɪ/ occupy an intermediate region rather than belonging discretely to either side of the asymmetry.
These findings together suggest that /l/-velarization following different preceding vowels is better characterized as a graded bias toward two regions of the F1-F2 space than as a fully categorical, two-locus system. A possible alternative explanation for the /eɪ/ pattern concerns measurement: as a diphthong, /eɪ/ does not have a single stable articulatory target, and the point selected as its “steady state” is to some extent conventional.
More broadly, the coexistence of a directionally robust component (at the steady state) with a more variable, boundary-sensitive component (at the offset, and for /ɔ:/ and /eɪ/) is not without precedent. Using direct articulatory data, Turton (2017) similarly found that categorical allophony and gradient phonetic effects coexist within a single grammar, interpreted through a life-cycle model in which a gradient phonetic process stabilizes into a more categorical phonological one while the earlier gradient layer persists alongside it. The present acoustic results are consistent with a comparable layering, though the present data cannot adjudicate between a synchronic gestural-competition account (2.1) and a diachronic life-cycle account; this is a question direct articulatory data would be better placed to resolve (Section 6).
6. Conclusion
This paper reanalyzed a descriptive observation from Lee (2023)—using the same acoustic data—that word-final /l/ following front versus non-front preceding vowels appears to converge on two distinguishable regions of the F1-F2 space. Treating speaker and vowel as random effects, a front/non-front asymmetry was significant for F1 and marginal for F2, collapsing across the steady state and offset; critically, the frontness-by- measurement-point interaction was not significant for either formant, so the present data do not support the claim, offered in an earlier version of this analysis, that this asymmetry is more robust at the steady state than at the offset. A model-comparison test further showed that a continuous coarticulatory account, based on the /l's own onset formants and preceding-vowel duration, fits the steady-state data at least as well as the categorical grouping, without needing the categorical term at all; this is, on balance, the most parsimonious reading of the present results, and the front/non-front framing used throughout this paper should be understood as a convenient, phonologically motivated redescription of part of this continuous pattern rather than as an independently established categorical distinction. Token-level classification analyses, treated as exploratory, corroborated the overall asymmetry and further indicated that /ɔ:/ and /eɪ/ sit near the boundary between the two regions, and that /i:/ dissociates from /e/ and /eɪ/ in F1 despite patterning with them in F2.
Several limitations should be noted. First, the number of speakers (n=7, of the ten American English speakers reported in Lee (2023); the remaining three did not provide complete tokens across all six word-final vocalic contexts and were accordingly not included) is small; replication with a larger sample is needed. Second, the acoustic endpoints discussed here are not direct observations of an articulatory target; the claims made are about F1/F2 position, and a gestural interpretation is offered as one possibility among others. Third, each vocalic context was instantiated by a single word, so vocalic context and lexical item are fully confounded; the present conclusions should accordingly be restricted to the sampled speakers and words rather than generalized to word-final /l/ across the board. Fourth, of the formant-analysis settings used in the original measurements in Lee (2023), the formant ceiling (5,000 Hz) was recoverable from the original spectrogram display records, but the LPC order and analysis window length are not separately documented and could not be independently verified for this reanalysis. Fifth, the reference schwa position (Flemming, 2009) is drawn from an external study rather than from the same speakers, since no independent schwa tokens were collected from them; this is a limitation of the comparison in Figures 1 and 2. The comparison is nonetheless reasonably well matched in one respect: Flemming's (2009) reference values (F1=665 Hz, F2=1,772 Hz) were themselves derived from nine female speakers of American English, matching the sex and national variety of the present sample, which somewhat mitigates the concern that the reference point reflects a mismatched population.
Future work might extend the present analysis in three directions: direct articulatory validation, for instance through ultrasound tongue imaging in the manner of Turton (2017), which could establish whether the asymmetry reported here reflects a genuinely bimodal gestural target or a unimodal target under variable coarticulatory pressure, and in particular could clarify the status of /i:/'s dissociation between F2 and F1; replication with a larger and more balanced speaker sample, with multiple items per vowel to disentangle vocalic context from lexical item; and comparison with the Korean-speaker data reported in Lee (2023), building on Sohn & Lim's (2020) finding that Korean EFL speakers implement word-final /l/-darkening less strongly than native English speakers while following the same direction of prosodic conditioning.






