Dynamically optimized pronunciation learning support system and method integrating acoustic statistics, visualization, and history control

The system addresses the limitations of conventional pronunciation learning by statistically evaluating and visually presenting errors, integrating history analysis, and dynamically optimizing materials, leading to improved pronunciation learning outcomes.

JP7784097B1Active Publication Date: 2025-12-11池上 さくら
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025109818
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-12-11
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

Conventional pronunciation learning support systems lack structural analysis of errors, visual presentation of errors, dynamic adaptation to individual learner's tendencies, and integration of learning history for personalized material presentation, leading to ineffective feedback and stagnation in pronunciation improvement.

Method used

A system that statistically evaluates acoustic characteristics, visually presents errors, and dynamically optimizes learning materials based on learner history by integrating error visualization, history analysis, and material presentation control.

Benefits of technology

Enables intuitive understanding of error direction and magnitude, promotes continuous and effective pronunciation improvement by adapting to individual learner needs, enhancing educational effectiveness and retention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007784097000001_ABST
    Figure 0007784097000001_ABST
Patent Text Reader

Abstract

We provide a pronunciation learning support system that improves the efficiency, acceptability, and continuity of pronunciation learning by acoustically quantitatively evaluating learners' pronunciation errors, visually presenting their directionality and magnitude, and individually optimizing teaching materials based on error history. [Solution] This invention processes speech on a phoneme-by-phoneme basis, extracts formants F1 and F2 from each phoneme, and performs error evaluation based on Z-scores and normalized distance Z_norm. Error information is visually displayed as points, vectors, and ellipses in F1F2 space, intuitively presenting the statistical error structure to the learner. Error information is recorded as a history and analyzed over time, and error trends are classified using clustering processing. The teaching material presentation means calculates presentation priority based on scoring based on the amount of error and frequency of recurrence, and is equipped with a mechanism for gradually controlling the difficulty of the teaching material according to the recurrence rate and improvement trend.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for supporting foreign language pronunciation learning, and more particularly to a pronunciation learning support system and method that statistically evaluates the acoustic characteristics of learners, visually visualizes errors, and individually optimizes learning materials based on the history. [Background technology]

[0002] Conventional pronunciation learning support systems typically use automatic speech recognition (ASR) to score the degree of agreement between the learner's voice and a model voice, and then present a correct / incorrect result and an overall score.

[0003] This method had the following technical issues: (1) There was no mechanism for evaluating and presenting the directionality or acoustic structure of pronunciation errors, and feedback was limited to a simple numerical score. (2) Information about the cause of the error or the direction of correction was not communicated to the learner, making it difficult for them to self-correct to improve their pronunciation. (3) There was no process for accumulating and analyzing the learner's pronunciation history or tendencies, and the presentation of teaching materials was uniform and static. (4) There was no structure for directly analyzing the acoustic characteristics of speech (formants F1, F2, etc.), and there was no support means for spatially and visually recognizing pronunciation discrepancies.

[0004] This meant that learners were unable to grasp the structure of their pronunciation or sense changes, and lacked the visual and intuitive information that would trigger self-correction. For these reasons, a technical framework was needed to statistically and spatially analyze errors and individually optimize learning content based on history.

[0005] As such, conventional pronunciation learning support systems lack structural analysis of errors and visual presentation, and have difficulty dynamically presenting teaching materials according to the pronunciation tendencies of individual learners. Therefore, new technological means are needed to statistically evaluate and visualize the direction and magnitude of errors, and to optimize teaching material presentation based on that history. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2022-078317 [Patent Document 2] Japanese Patent Publication No. 2020-158672 [Non-patent literature]

[0007] [Non-Patent Document 1] Boersma, P. & Weenink, D. “Praat: doing phonetics by computer”, https: / / www.fon.hum.uva.nl / praat / [Non-patent document 2] Watanabe et al. “Formant-based statistical modeling for L2 pronunciation feedback”, Proceedings of Interspeech 2020 Summary of the Invention [Problem to be solved by the invention]

[0008] This invention solves several unresolved technical problems in conventional pronunciation support systems that provide ASR scores. First, (1) it is difficult to understand the structure of pronunciation errors, and there has been no support system that focuses on acoustic structures (F1, F2, etc.).

[0009] Second, (2) learners were not provided with a means to visually recognize and understand errors, and feedback remained abstract and numerical.

[0010] Third, (3) there was no mechanism for individually optimizing learning materials based on learning history, and learning presentation was limited to static and uniform presentation.

[0011] Fourth, (4) there was no feedback control to encourage behavioral change in response to errors, that is, no educational design to induce cognitive change.

[0012] Fifth, (5) the accumulation and analysis of pronunciation learning history and the presentation of teaching materials were not integrated, and past learning was not utilized in selecting teaching materials.

[0013] Sixth, (6) there was a lack of consistency in pronunciation evaluation, visualization, and material selection based on a common scale (such as Z-score), making progress management and comparison difficult. To solve these problems, the present invention provides a configuration that integrates statistical error evaluation, visual presentation, history analysis, and optimal material presentation. [Means for solving the problem]

[0014] This invention employs a configuration in which error evaluation based on acoustic statistics, error visualization, history accumulation and trend analysis, and dynamic presentation control of teaching materials based on the analysis results are linked in a causal and interdependent manner.

[0015] In particular, error assessment is quantified as a Z-score by comparing statistical models of formants F1 and F2, and by plotting this in two-dimensional space, the "direction" and "distance" of the error can be visually presented, eliciting intuitive phonological awareness in learners.

[0016] Furthermore, error information is accumulated and classified as a history, and without this information dynamic optimization of teaching materials would not be possible. In other words, the means of error visualization, history analysis, and teaching material control are causally linked to each other, and are indispensable elements that cannot achieve the objectives of this invention alone.

[0017] These means form a four-layered learning support flow consisting of statistical error evaluation, visual feedback, history analysis, and teaching material control, which comprehensively enhances the accuracy, persuasiveness, and continuity of pronunciation improvement. Note that the present invention is not limited to acoustic features such as F1 and F2, and can also be applied to configurations that visually present error points and vectors in a speech feature space that includes articulatory features, and these are also included in the technical scope of the present invention.

[0018] The statistical error evaluation means extracts the formants F1 and F2 of each phoneme and compares them with a statistical model based on native speech to calculate Z-scores (first Z-score and second Z-score) and a normalized distance based on them (hereinafter referred to as Z-score normalized distance). This makes it possible to simultaneously evaluate the magnitude and direction of the error.

[0019] The visualization display method not only plots error points and vectors in the F1-F2 space, but also displays heat maps that reflect past trends and the centers of gravity of error clusters, making it easier for learners to notice errors through visual incongruity.

[0020] The history analysis method accumulates and classifies the Z-score sequence ΔZ(t) and the acoustic difference vector ΔF for each learning session in time series, and characterizes the error trend through clustering processing. This information is dynamically reflected in the presentation of learning materials.

[0021] The learning material presentation optimization means uses the score function shown in Figure 7 to calculate presentation priorities based on the amount of error and the frequency of erroneous speech. The score function is defined by adding the value obtained by multiplying the magnitude of the formant error for each phoneme by a first coefficient (α) and the value obtained by multiplying the frequency of erroneous speech for that phoneme by a second coefficient (β). Based on the calculated score, phonemes whose error index exceeds a predetermined threshold (hereinafter referred to as the error judgment threshold) are filtered out and reflected in the selection of learning materials to be presented next. [Effects of the Invention]

[0022] This invention quantitatively evaluates the error of each phoneme using a Z-score based on acoustic statistics and visually presents the error points, vectors, and heat maps in F1-F2 space, allowing learners to intuitively understand the "direction" and "distance" of pronunciation errors. This promotes structural understanding and spontaneous phonological awareness that could not be achieved with conventional numerical scores alone.

[0023] Furthermore, in this invention, by recording and clustering error scores and acoustic difference vectors ΔF over time, we analyze the error trends for each learner and optimize the presentation of teaching materials based on that analysis. This dynamic control allows us to prioritize phonemes with many errors and recurring errors, enabling individually adaptive pronunciation learning.

[0024] Furthermore, by integrating the three elements of statistical evaluation, visual visualization, and history optimization, it is possible to achieve high levels of accuracy, persuasiveness, and continuity in pronunciation learning support. This will enable the provision of highly effective and sustainable pronunciation correction support in a variety of educational settings, including school education, language training, and personal learning applications, and will also make a significant industrial contribution to the field of foreign language education. [Brief explanation of the drawings]

[0025] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of a pronunciation learning support system according to the present invention. [Figure 2] A diagram showing an example of the configuration of an error display UI with F1 and F2 as coordinate axes. [Figure 3] Overall processing flow diagram of error history analysis and teaching material presentation control [Figure 4] A diagram showing an example of the UI for displaying learning progress and outputting reports [Figure 5] Diagram showing the mutual feedback structure of error visualization, history analysis, and teaching material presentation [Figure 6] Control flow diagram for teaching material presentation based on moving average of Z scores and re-error rate [Figure 7] Definition formula of the score function for calculating the teaching material presentation score [Figure 8] Diagram showing the Z-score standardization formula [Figure 9] Formula for calculating the moving average Z score [Figure 10] Definition of the re-error rate and its components [Figure 11] Flowchart for determining the difficulty of learning materials based on the error rate and change rate DETAILED DESCRIPTION OF THE INVENTION

[0026] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. As shown in Fig. 1, a pronunciation learning support system according to the present invention comprises a voice input means 101, a phoneme division means 102, an acoustic feature extraction means 103, a statistical error evaluation means 104, a visualization display means 105, a history recording and analysis means 106, and a teaching material presentation optimization means 107. This configuration corresponds to the invention defined in claim 1.

[0027] The voice input means 101 is used to acquire the voice of the learner using a terminal device equipped with a microphone.

[0028] The phoneme dividing means 102 is for dividing the speech acquired by the speech input means 101 into phoneme units using a forced alignment method such as Montreal Forced Aligner.

[0029] The acoustic feature extraction means 103 extracts the first formant (F1) and the second formant (F2) from each phoneme section divided by the phoneme division means 102, and can be realized using software such as Praat or librosa.

[0030] The statistical error evaluation means 104 compares the values ​​of formant frequencies F1 and F2 obtained by the acoustic feature extraction means 103 with a preset statistical model (for example, an average value and standard deviation based on native speakers) and calculates a standardized score for each formant (hereinafter referred to as a first Z-score and a second Z-score) and a normalized error distance based on these. The standardized score is an index that quantitatively evaluates, based on the standard deviation, how much the target acoustic feature value deviates from the average value of the statistical model, and its calculation method is shown in Figure 8.

[0031] The visualization display means 105 is a means for displaying the error information calculated by the statistical error evaluation means 104 in a two-dimensional space with the formants F1 and F2 as the coordinate axes, and in this space, it draws index points indicating the error points of each phoneme, arrow lines indicating the error direction, and a color distribution diagram (so-called heat map) showing the error trend.

[0032] The history recording and analysis means 106 records the error score and the frequency of occurrence for each phoneme in chronological order, and classifies and analyzes pronunciation trends using analytical techniques such as clustering.

[0033] The learning material presentation optimization means 107 calculates the presentation priority for each phoneme or word using the score function shown in Fig. 7 based on the classification results and error index obtained by the history recording and analysis means 106, and dynamically presents learning materials to the learner. Note that the error evaluation and visualization configuration according to the present invention is not limited to the case where formant frequencies F1 and F2 are used as acoustic features, but can also be applied to visual presentation using spatial representation based on features related to articulatory movements such as tongue height, front-back position, and lip shape.

[0034] In the present invention, "error trend" refers to a phenomenon in which the change history of the Z-score calculated for a certain phoneme or word shows a certain pattern of temporal fluctuation (e.g., continuous increase, stagnant change, periodic oscillation, etc.). In order to quantitatively classify these error trends, an evaluation is performed using the moving average value of the Z-score over a predetermined period (e.g., the average of the most recent five points in time).

[0035] The moving average is defined by calculating the average of the normalized distances of the Z scores (Z score error index) for the last five sessions. This moving average is the arithmetic mean of five Z score error indexes, including the most recent value at time t and the data from the four sessions prior to that, and its calculation method is shown in Figure 9.

[0036] In the present invention, the "recurrence error rate" refers to the rate at which an utterance in which the normalized error index of the Z-score for a target phoneme or word exceeds a predetermined threshold (e.g., 1.5) reoccurs during a certain period of time. The recurrence error rate is calculated as the rate at which errors occur from the second time onward relative to the total number of times the phoneme is presented. The definition and calculation method are detailed in FIG. 10.

[0037] If the rate of change of the moving average value is 0 or less and no tendency for improvement is observed, or if the re-error rate for the target phoneme exceeds a predetermined threshold (e.g., 30 percent), the learning material presentation optimization means 107 selects learning materials containing the phoneme or a similar phoneme as a target for re-presentation. The calculation methods for the moving average value and the rate of change are shown in Figure 9, and the definition of the re-error rate is shown in Figure 10.

[0038] Furthermore, the difficulty level of the presented learning materials is adjusted in stages according to the learner's error history. Specifically, if the re-error rate for the target phoneme or word exceeds a predetermined upper threshold (e.g., 30 percent) and the rate of change in the moving average is positive or zero, it is determined that the learner has not made sufficient improvement, and the difficulty level of the learning materials is lowered by one level (e.g., from sentence to word, from word to monophoneme). On the other hand, if the re-error rate falls below a predetermined lower threshold (e.g., 10 percent) and the rate of change in the moving average is negative, it is determined that the learner has improved, and the difficulty level of the learning materials is raised by one level (e.g., from monophoneme to word, from word to sentence). If any other conditions are met, the current difficulty level is maintained. A detailed judgment flow for each condition is shown in Figure 11.

[0039] This enables dynamic and individually optimized control of teaching material presentation according to the learner's tendency to make errors, making it possible to provide a learning experience that avoids excessive load and stagnation and promotes a continuous sense of accomplishment. [Example]

[0040] For example, when a learner utters the word "bit," the speech is recorded by the speech input means 101, and the speech is divided into three phonemes (an initial plosive, a central short vowel, and a final plosive) by the phoneme division means 102. Next, the acoustic feature extraction means 103 extracts the first and second formants from the phoneme section corresponding to the central vowel, and the statistical error evaluation means 104 compares them with a preset native statistical model. From this comparison, a standardized score and a normalized error index corresponding to each formant are calculated. Details of these calculation methods are shown in FIG. 8.

[0041] These error values ​​are displayed by the visualization display means 105 in a two-dimensional space with the first and second formants as coordinate axes, and are drawn as index points indicating each error, with directional lines indicating the difference from the reference point. Furthermore, if the same or similar errors have been observed in the past, they are reflected in a color distribution diagram as a visual distribution, and cluster classification is performed by the history recording and analysis means 106. The learning material presentation optimization means 107 calculates presentation priorities using the score function shown in Figure 7 based on this error trend data, and dynamically selects and presents learning materials of the most appropriate difficulty level to the learner, taking into consideration the re-error rate and the changing trends of the moving average value.

[0042] This invention can be widely applied to foreign language pronunciation learning support systems, including English, and can achieve high educational effectiveness and retention rates, particularly in online learning platforms, mobile learning apps, pronunciation assessment software for educational institutions, and individually optimized e-learning materials. It can also be applied to the entire education industry, including developmental support education requiring pronunciation correction, adult English learning, and even global human resource development support. [Industrial Applicability]

[0043] This invention is extremely useful in the field of foreign language education, including English, as it is configured to achieve statistical error evaluation, visual visualization, and history-based individualized learning material optimization in pronunciation learning support. It can also be applied to e-learning materials, smartphone learning applications, online pronunciation assessment systems, and more, and is expected to be widely deployed in educational institutions, language schools, corporate training, and the individual learning support market. Furthermore, as a technological platform that contributes to the development of new services through the fusion of AI, speech technology, and educational engineering, this technology has great industrial value in the education industry as a whole. [Explanation of symbols]

[0044] 101: Voice input means 102: Phoneme segmentation means 103: Acoustic feature extraction means 104: Statistical error evaluation method 105: Visualization display means 106: History recording and analysis means 107: Teaching material presentation optimization method 201: Two-dimensional space UI display area with F1 and F2 as coordinate axes 202: Recommended teaching materials list display area 301: Z-score historical data 302: Clustering processing (unsupervised learning) 303: Calculation of Z-Cluster integration index (Vi) 304: Calculation of teaching material presentation score Si 305: Selection and presentation of teaching materials 401: Learning progress bar 402: Achieved Phoneme List 403: Z-score change graph 404: Error history display 405: PDF output button (1): Native speaker statistical ellipse (plus or minus 2 SD) (2): Learner's pronunciation position (Z score) (3): Error direction vector (4): Error history heat map 601: Ri evaluation unit (calculation of re-error rate) 602: Moving average calculation section 603: Decision flow branching section 604: Re-presentation processing unit 605: Output control unit 606: Teaching material selection control Z diff(t): The difference in the Z-score series (alternative notation for ΔZ(t)) mu5(t): The moving average Z score of the last 5 data points (alternative to μ5(t)) deltaMu: The change in the moving average (alternative to Δμ) Ri: Re-error rate (proportion of failed pronunciations for each phoneme) Vi: A phoneme characteristic index integrating Z-scores and cluster information Si: Teaching material presentation score (used to determine presentation order)

Claims

1. A pronunciation learning support system comprising: a speech input means for acquiring spoken speech; a phoneme division means for dividing the speech into phoneme units; an acoustic feature extraction means for extracting formants F1 and F2 of each phoneme; a statistical error evaluation means for comparing the extracted F1 and F2 with the statistical distribution of native speakers to calculate a Z-score and a normalized distance Z_norm; a visualization display means for visually displaying the error in a two-dimensional space with F1 and F2 as the coordinate axes; a history recording means for recording the error evaluation results as a history; an error trend analysis means for analyzing error trends based on the history; and a learning material presentation optimization means for selecting and presenting learning materials based on the analysis results, and characterized in that the system performs quantitative evaluation of pronunciation errors and controls learning material presentation without using automatic speech recognition (ASR).

2. 2. The pronunciation learning support system according to claim 1, wherein the teaching material presentation optimization means comprises scoring means for calculating a presentation priority of teaching materials based on an acoustic error vector and a frequency of erroneous speech, and the scoring means calculates the presentation priority based on the acoustic error vector and the frequency of erroneous speech using a first weighting factor for the acoustic error vector and a second weighting factor for the frequency of erroneous speech.

3. 3. The pronunciation learning support system according to claim 1, wherein the visualization display means comprises a user interface that simultaneously displays, on the same plane in a two-dimensional space with F1 and F2 as coordinate axes, an ellipse indicating the statistical distribution of reference phonemes, an error point indicating the learner's pronunciation position, and a vector indicating the error direction.

4. A pronunciation learning support system as described in claim 1 or 2, characterized in that the history recording means and error tendency analysis means record the acoustic difference vector ΔF, occurrence frequency f, and Z-score series ΔZ(t) for each learning session in chronological order, and have the function of classifying and characterizing the learner's error tendency through clustering processing.

5. A pronunciation learning support system as described in claim 1 or 2, characterized in that the teaching material presentation optimization means is equipped with a filter processing mechanism that preferentially extracts phonemes or words whose Z_norm calculated by the statistical error evaluation means exceeds a predetermined threshold, and presents teaching materials containing those phonemes or words.

6. A pronunciation learning support system as claimed in claim 1 or 2, characterized in that the history recording means and teaching material presentation optimization means are provided with a mechanism for automatically re-presenting teaching materials for the phoneme or similar phonemes when it is determined, based on the history information, that the moving average of the Z score has not improved within a predetermined period (e.g., the most recent five sessions) or that the re-error rate (the rate at which Z_norm exceeds a threshold) has exceeded a predetermined value (e.g., 30%).

7. A pronunciation learning support system as described in claim 1 or 2, characterized in that the teaching material presentation optimization means is equipped with a teaching material presentation control mechanism that gradually adjusts the difficulty level of the teaching materials to be presented (e.g., monophone → CVC → sentence) based on historical information, in accordance with the recurrence error rate (recurrence rate of exceeding the Z_norm threshold) and the error improvement trend (moving average change rate of the Z score).

Citation Information

Patent Citations

  • Dysarthria vowel evaluation template and method

    CN108670199A

  • Language learning support method

    JP2007065291A

  • Real time voice analysis and method for providing speech therapy

    US20070168187A1

  • Photocurable resin composition

    JP2020158672A

  • Compositions for immunizing against Staphylococcus aureus

    JP2022078317A