English pronunciation correction method, device, equipment and storage medium
By acquiring learners' speech and physiological data on English pronunciation, and utilizing cross-modal fusion feature data and error occurrence prediction models, the English pronunciation training program is adjusted in real time, solving the problem of poor dynamic adjustability in existing English pronunciation correction methods and achieving efficient English pronunciation correction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA ENERGY GRP NINGXIA COAL IND CO LTD
- Filing Date
- 2026-03-16
- Publication Date
- 2026-07-31
AI Technical Summary
Existing English pronunciation correction methods use a static feedback model, which has poor dynamic adjustment capabilities and affects learners' English pronunciation learning outcomes.
By acquiring learners' English pronunciation speech data, combined with physiological and cognitive feature data, and utilizing cross-modal fusion feature data and error prediction models, English pronunciation training programs can be adjusted in real time, including interventions for phoneme-level, prosodic-level, and mother tongue transfer errors.
It enables proactive intervention in English pronunciation errors, improves correction efficiency, reduces the risk of ingrained incorrect memories, and enhances learning outcomes.
Smart Images

Figure CN122493883A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, specifically to an English pronunciation correction method, an English pronunciation correction device, an electronic device, and a machine-readable storage medium. Background Technology
[0002] With the continuous growth in demand for English learning, learners' pursuit of standard pronunciation is becoming increasingly prominent. In international academic exchanges, transnational business cooperation, and daily cross-cultural communication, standard, fluent, and natural spoken English has become a core competitive advantage. Traditional English speaking learning mainly relies on face-to-face instruction from teachers or immersive language environments, but both methods have significant limitations: human tutoring is constrained by the number of teachers, their teaching level, and geographical distribution, making it difficult to meet the needs of a large number of learners on a large scale; while building a real language environment is costly, and learning outcomes are easily affected by individual differences.
[0003] Current English pronunciation correction methods mostly employ a static feedback model, which evaluates learners' pronunciation retrospectively based on preset standards, failing to dynamically adapt and adjust according to learners' real-time speech signals. This lagging and fixed correction method is difficult to effectively correct incorrect pronunciation habits, easily leading to a gradual deterioration of learners' spoken English pronunciation and other aspects during the learning process. Summary of the Invention
[0004] The purpose of this application is to provide an English pronunciation correction method, apparatus, device, and storage medium to solve the problem that the dynamic adjustability of existing English pronunciation correction methods using static feedback mode is poor, which affects learners' learning of English pronunciation.
[0005] To achieve the above objectives, the first aspect of this application provides a method for correcting English pronunciation, the method comprising: Acquire the speech data of learners' English pronunciation; Based on the speech data, cross-modal fusion feature data is determined; wherein, the cross-modal fusion feature data is the fusion data of acoustic dimension and physiological dimension; The cross-modal fusion feature data is input into a pre-trained error occurrence prediction model, and the error occurrence prediction model outputs error occurrence prediction results; wherein, the error occurrence prediction results include the types of pronunciation errors that the learner may make in the future; If the error occurrence prediction result meets the preset dynamic adjustment conditions, a target early intervention plan is determined based on the error occurrence prediction result, and the learner's current English pronunciation training plan is adjusted according to the target early intervention plan.
[0006] In one embodiment of this application, the method further includes: Acquire physiological and cognitive characteristic data of learners' English pronunciation, as well as specific parameter data used to quantify the learner's emotional state; Based on the specific parameter data, the learner's emotional state is determined; Based on the physiological feature data, the cognitive feature data, the emotional state, the cross-modal fusion feature data, and the error occurrence prediction results, the learner's current English pronunciation training program is adjusted using a pre-established adaptive decision engine.
[0007] In one embodiment of this application, determining cross-modal fusion feature data based on the speech data includes: Feature extraction is performed on the speech data to obtain phoneme-level feature data and prosodic feature data; wherein, the phoneme-level features include dynamic formant trajectory parameters, voice onset time parameters, and fricative energy distribution parameters, and the prosodic features include speech contour parameters, stress distribution parameters, and rhythm deviation parameters; Based on the speech data, physiological motor characteristic data are determined; among which, physiological motor characteristics include tongue tip height, nasal pronunciation integrity, and lip opening and closing curve; The phoneme-level feature data, the prosodic feature data, and the physiological motion feature data are fused to obtain cross-modal fused feature data.
[0008] In one embodiment of this application, pronunciation errors are categorized into phoneme-level errors, prosodic-level errors, and mother tongue transfer errors. Based on the error occurrence prediction results, a target early intervention plan is determined, including: If the pronunciation error type in the error occurrence prediction result includes phoneme-level errors, then preset contrastive phoneme training is added; If the pronunciation error type in the error prediction result includes prosodic errors, then standard pronunciation demonstration and rhythm guidance will be added; If the pronunciation error type in the error occurrence prediction result includes mother tongue transfer errors, then the analysis of differences between mother tongue and English pronunciation will be added.
[0009] In one embodiment of this application, specific parameters include anxiety index and attention distraction index; Based on the specific parameter data, the learner's emotional state is determined, including: Based on anxiety index data and attention distraction index data, a weighted calculation method was used to quantify the emotion index data. The learner's emotional state is determined based on the aforementioned emotion index data.
[0010] In one embodiment of this application, the adaptive decision engine includes: The state space, which includes multi-dimensional parameters, is used to characterize the learner's real-time state. These multi-dimensional parameters include pronunciation accuracy, emotional state, potential pronunciation error patterns, physiological characteristics, and cognitive characteristics. Action space, including training intensity adjustment, learning feedback method selection, and knowledge point injection; The reward function is used to guide the adaptive decision engine to generate an optimized pronunciation training strategy based on indicators such as improved pronunciation accuracy, reduced error repetition rate, and optimized learning time, so as to adjust the learner's current English pronunciation training plan.
[0011] In one embodiment of this application, the method further includes: Collect core data on learners completing a full English pronunciation training task; the core data includes pronunciation accuracy, error repetition rate, emotion change curve, and training duration. Based on the core data, the node data of the learner's pre-built personalized pronunciation tree is updated, and the parameters of the error occurrence prediction model and the reward function of the adaptive decision engine are optimized; wherein, the personalized pronunciation tree is used to create new English pronunciation training tasks for learners.
[0012] A second aspect of this application provides an English pronunciation correction device, the device comprising: The speech data acquisition module is used to acquire speech data of learners' English pronunciation; A cross-modal fusion module is used to determine cross-modal fusion feature data based on the speech data; wherein the cross-modal fusion feature data is fusion data of acoustic dimension and physiological dimension; The potential error prediction module is used to input the cross-modal fusion feature data into a pre-trained error occurrence prediction model, and output error occurrence prediction results through the error occurrence prediction model; wherein, the error occurrence prediction results include the types of pronunciation errors that the learner may make in the future; The training program adjustment module is used to determine a target early intervention plan based on the error occurrence prediction results when the error occurrence prediction results meet the preset dynamic adjustment conditions, and to adjust the learner's current English pronunciation training plan according to the target early intervention plan.
[0013] A third aspect of this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the English pronunciation correction method described in the first aspect above.
[0014] A fourth aspect of this application provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the English pronunciation correction method described in the first aspect.
[0015] The English pronunciation correction method, device, equipment, and storage medium provided in this application determine the fusion data of acoustic and physiological dimensions based on the acquired speech data, realizing cross-modal fusion of acoustic and physiological aspects. This effectively improves the accuracy of subsequent error prediction models in predicting potential pronunciation errors in advance. When the error prediction results output by the error prediction model meet the preset dynamic adjustment conditions, a corresponding target early intervention plan is determined to adjust the learner's current English pronunciation training plan in real time. This achieves proactive intervention in the learner's pronunciation errors. Compared with the passive error correction method of traditional pronunciation correction, this application enables rapid and dynamic adjustment of the English pronunciation training plan based on the learner's pronunciation training situation. It can quickly switch between different pronunciation training strategies to correct the learner's pronunciation, effectively improving correction efficiency and reducing the risk of error memory solidification.
[0016] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings: Figure 1 The schematic diagram illustrates a flowchart of an English pronunciation correction method according to an embodiment of this application; Figure 2 This schematic diagram illustrates the structural block diagram of an English pronunciation correction device according to an embodiment of this application; Figure 3 The diagram illustrates the internal structure of a computer device according to an embodiment of this application.
[0018] Explanation of reference numerals in the attached figures A01 - Processor; A02 - Network Interface; A03 - Internal Memory; A04 - Display Screen; A05 - Input Device; A06 - Non-volatile Storage Media; B01 - Operating System; B02 - Computer Program. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0020] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0021] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0022] In view of the poor dynamic adjustability of English pronunciation correction methods using static feedback mode in related technologies, which affects learners' learning of English pronunciation, this application provides an English pronunciation correction method, device, equipment, and storage medium. The English pronunciation correction method, device, equipment, and storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and implementation methods.
[0023] Figure 1 The schematic diagram illustrates a flowchart of an English pronunciation correction method according to an embodiment of this application. For example... Figure 1 As shown in one embodiment of this application, an English pronunciation correction method is provided, which may include the following steps.
[0024] Step 200: Obtain the learner's English pronunciation audio data.
[0025] It is important to understand that traditional English pronunciation correction methods often use single-channel microphones to capture speech signals. This method has three major drawbacks: First, it has weak sound source location tracking capabilities and is easily affected by environmental noise, leading to speech signal distortion. Second, it cannot effectively eliminate nonlinear echoes, especially in indoor learning scenarios, where echoes can severely affect the accuracy of subsequent feature extraction. Third, the signal acquisition dimension is singular, making it difficult to capture subtle acoustic differences during the pronunciation process.
[0026] Therefore, in this embodiment, to obtain high-quality voice data, a four-channel ring microphone array is deployed in the learner's environment before executing this method. Based on this, step 200 includes the following steps.
[0027] Step 210: The learner's speech signal is captured in real time using a four-channel ring microphone array to obtain multiple original speech signals.
[0028] Step 220: The nonlinear echoes in each original speech signal are eliminated in layers using the Subband Normalized Least Mean Square (S-NLMS) algorithm to obtain the speech signal after first-level processing.
[0029] Step 230: The speech signals after each primary processing stage are sequentially framed and windowed to obtain the speech signals after secondary processing.
[0030] Step 240: Based on the secondary processed speech signal, the Minimum Variance Distortionless Response (MVDR) algorithm is used to dynamically track and locate the learner's sound source location in order to enhance the speech component in the direction of the learner's sound source, thereby obtaining the speech data.
[0031] Step 220, which utilizes the S-NLMS algorithm to perform layered elimination of nonlinear echoes (such as echoes caused by room reflections and equipment delays) commonly encountered in spoken language training, includes: decomposing the speech signal into multiple sub-bands and adaptively adjusting the filtering coefficients through normalization processing. Notably, step 220, by eliminating nonlinear echoes, avoids echo interference with the extraction of nasal resonance signals.
[0032] In addition, step 230 converts the signal into frame-level speech data by performing frame segmentation and windowing processing, which ensures that each frame of the signal can reflect the instantaneous details of the pronunciation (such as the instantaneous nasal resonance state).
[0033] It is important to understand that this embodiment uses a four-channel ring microphone array instead of a traditional single-channel microphone. Through a spatially distributed layout, it achieves 360° sound source coverage. Combined with the MVDR algorithm to dynamically track the learner's pronunciation direction, it ensures accurate targeting of the sound source regardless of the learner's posture or distance, while filtering out irrelevant environmental noise. Compared to a single-channel microphone, the signal-to-noise ratio of the four-channel ring layout in this embodiment can be improved by more than 40%, and the orientation recognition error can be controlled within ±5°. This optimized hardware deployment provides high-fidelity data support for subsequent feature analysis.
[0034] Furthermore, it's important to understand that compared to the traditional NLMS algorithm, the S-NLMS algorithm has the advantages of faster convergence speed and better nonlinear echo suppression. This embodiment, by introducing the S-NLMS algorithm, can reduce the distortion of the echo-cancelled speech signal to less than 5%, ensuring complete preservation of pronunciation details.
[0035] As can be seen, step 200, by combining a four-channel ring microphone array, the S-NLMS algorithm, and the MVDR algorithm, can obtain a high-fidelity speech signal. This effectively solves the technical pain points of traditional speech signal capture, such as high noise, large distortion, and directional bias, laying a solid foundation for subsequent feature extraction and error prediction, and providing a guarantee for achieving accurate correction.
[0036] Step 300: Determine cross-modal fusion feature data based on the speech data.
[0037] The cross-modal fusion feature data is a fusion of acoustic and physiological dimensions.
[0038] It is important to understand that speech feature extraction is the core of speech correction. Traditional speech feature extraction methods are mostly limited to phoneme-level acoustic features (i.e., limited to the acoustic dimension) and have a single parameter dimension, which cannot fully reflect the physiological mechanism and rhythmic rules of pronunciation. This leads to superficial localization of pronunciation errors and difficulty in uncovering the deeper causes.
[0039] Based on this, this embodiment constructs a comprehensive pronunciation feature map through feature dimension expansion and cross-modal fusion technology to achieve accurate source tracing of pronunciation defects.
[0040] In this embodiment, step 300 includes the following steps.
[0041] Step 310: Extract features from the speech data to obtain phoneme-level feature data and prosodic feature data.
[0042] Among them, phoneme-level features include dynamic formant trajectory parameters, voice onset time parameters, and fricative energy distribution parameters.
[0043] In this embodiment, step 310 uses the Mel-scale Frequency Cepstral Coefficients (MFCC) algorithm to extract features from the speech data to obtain the phoneme-level feature data.
[0044] It's important to understand that traditional MFCC algorithms rely solely on 39 basic parameters, making it difficult to capture subtle differences such as vowel gliding and consonant articulation timing. Therefore, this embodiment expands upon the traditional 39-dimensional parameters by adding three core parameters tailored to the characteristics of English pronunciation, achieving deep optimization of phoneme-level feature extraction: Dynamic formant trajectory parameters: By tracking the rate of change of F1-F4 formants on the time axis, vowel glide anomalies (such as unnatural vowel transitions during the pronunciation of / ai / ) can be accurately identified, thus making up for the shortcomings of traditional parameters that can only reflect static features. Voice Onset Time (VOT) parameter: By determining the time interval (which can be measured) from the release of a stop consonant to the vibration of the vocal cords, it effectively distinguishes the pronunciation differences of easily confused consonants such as / p / and / b / , / t / and / d / , in order to solve the problem of unclear consonant pronunciation caused by mother tongue transfer; Friction sound energy distribution parameters: By analyzing the differences in high-frequency energy concentration areas of fricatives such as / s / and / ʃ / based on Bark spectrum, the deviation of the fricative sound production position can be accurately located, so as to avoid the correction error caused by the unclear energy differentiation of traditional methods.
[0045] In a specific example, the parameter values of the three newly added core parameters are determined as follows.
[0046] Dynamic formant trajectory parameters: The preprocessed high-fidelity vowel speech signal is extracted, and framing and windowing are performed to adapt to the temporal characteristics of English vowel pronunciation. Based on the processed signal, frequency data of formants F1-F4 are continuously extracted and organized along the time axis to form the original time-series data of the formants. Based on the original time-series data of the formants, the frequency change process of the formants on the time axis is tracked, and the frequency change rate and transition trend of the formants are extracted to obtain the original parameter values of the dynamic formant trajectory. A standard formant trajectory parameter library for different English vowels, constructed based on standard pronunciation data of native English speakers, is retrieved. The original parameter values of the dynamic formant trajectory are matched and calibrated with the standard formant trajectory parameter library for the corresponding vowels to correct deviations caused by the environment and pronunciation acquisition, resulting in the final dynamic formant trajectory parameter values. Simultaneously, corresponding standard thresholds are matched to determine vowel slip anomalies.
[0047] Voice onset time parameters: Locate the pronunciation interval of English plosives from the preprocessed speech signal, and separate the silent segment from the plosive pronunciation segment; identify the release time of the plosive, mark the start point of the plosive segment, and simultaneously detect the onset time of vocal cord vibration to clarify the vibration initiation node; measure the time interval from the release time to the onset time of vocal cord vibration to obtain the original VOT parameter values; combine the different usage contexts of plosives at the beginning, middle, and end of words to perform context-adapted calibration of the original VOT parameter values; retrieve the standard VOT parameter library of English voiceless and voiced plosives covering different contextual standard parameter ranges; compare the calibrated parameter values with the standard library to determine the final VOT parameter values, and simultaneously determine the pronunciation deviation of plosives based on the corresponding standard thresholds.
[0048] Fricative energy distribution parameters: Bark spectrum analysis was performed on the preprocessed fricative speech signal to divide the high-frequency bands suitable for English fricatives; the energy values in each high-frequency band were extracted, the energy distribution in different frequency bands was statistically analyzed, and the frequency band where the energy peak was located was located; based on the energy distribution and the frequency band where the energy peak was located, the energy proportion and peak frequency band coordinates were extracted to obtain the original parameter values of fricative energy distribution; a standard energy distribution parameter library for different English fricatives was retrieved based on the Bark spectrum analysis results of standard pronunciation by native English speakers; the original parameter values of fricative energy distribution were matched and calibrated with the corresponding standard parameter library of fricatives to eliminate the energy deviation caused by signal acquisition, and the final fricative energy distribution parameter values were obtained. At the same time, the deviation of the fricative articulation position was determined by combining the corresponding standard threshold.
[0049] Among them, prosodic features include speech contour parameters, stress distribution parameters, and rhythm deviation parameters.
[0050] It is important to understand that prosody (including stress, rhythm, and intonation) is a core manifestation of fluency in spoken English. Based on this, this embodiment employs a pre-trained LSTM-RNN-based temporal analysis model to selectively extract the following three types of prosodic features in a temporal manner. This aims to capture pitch variations, stress distribution, and rhythmic patterns relevant to spoken fluency, while filtering out irrelevant and redundant features: Speech contour parameters: By calculating the dynamic time warping (DTW) distance between the pitch curve and the preset standard template, the intonation deviation is quantified to solve the "Chinese intonation" problem; Stress distribution parameters: Based on the ratio of the relative intensity to the duration of syllable energy peaks, missing stresses or stress misalignments are detected to improve the naturalness of spoken expression; Rhythm deviation parameter: Statistical variation coefficient of vowel duration to identify isochronic interference caused by native language transfer (such as the "one word at a time" phenomenon in English pronunciation by native Chinese speakers) and optimize pronunciation rhythm.
[0051] Step 320: Based on the voice data, determine the physiological motion characteristic data.
[0052] Among them, physiological motor characteristics include tongue tip height, nasal pronunciation integrity, and lip opening and closing curve.
[0053] In this embodiment, the time change rate of formants F1 and F4 is used to calculate the height of the tongue tip and quantify the tongue movement state; the nasalization index is used to judge the pronunciation integrity of nasal sounds (such as / m, n, ŋ / ) to avoid problems of insufficient or excessive nasal resonance; and the lip opening curve is reconstructed to reflect the amplitude of lip movement.
[0054] It is important to understand that nasalization degree is essentially the proportion of nasal resonance energy to total vocal energy. In this embodiment, the calculation logic of the nasalization degree index is as follows: First, key acoustic features are extracted: oral and nasal resonance energies in the speech data are separated through spectral analysis (such as Bark spectral analysis and MFCC). It is worth noting that oral resonance energy mainly corresponds to the fundamental frequency and harmonics of vowels, while the core characteristics of nasal sounds are enhanced low-to-mid-frequency energy and attenuated high-frequency energy, with nasal resonance peaks at ~250Hz and ~500Hz.
[0055] Then, the extracted key acoustic features are exponentially quantified: the nasalization index is set as nasal resonance energy / (oral resonance energy + nasal resonance energy), which means the value of the nasalization index is normalized to the range of [0,1].
[0056] It should be noted that in this embodiment, by binding the calculated nasalization index with the physiological state of nasal resonance, a cross-modal mapping from acoustic signal to physiological state is achieved. A higher nasalization index indicates stronger nasal resonance, and vice versa.
[0057] In this embodiment, the specific operation of determining the integrity of nasal pronunciation using the nasalization index is as follows: The calculated nasalization index is compared with the standard threshold range in the corresponding nasal pronunciation context (such as single pronunciation or pronunciation in connected speech). If the nasalization index is within the standard threshold range, it indicates that the nasal resonance intensity is moderate and the learner's pronunciation is complete. If the nasalization index is less than the left boundary value of the standard threshold range, it indicates that the nasal resonance is insufficient and the learner's pronunciation is incomplete. If the nasalization index is greater than the right boundary value of the standard threshold range, it indicates that the nasal resonance is excessive and the learner's pronunciation is incomplete.
[0058] In a specific example, when pronouncing / m / , the lips are closed but the nasal cavity is not fully opened, indicating insufficient nasal resonance in the learner; when pronouncing / n / , the nasal cavity is overly open, resulting in excessive nasalization that masks the details of oral pronunciation, indicating excessive nasal resonance in the learner.
[0059] It is worth noting that the nasal resonance intensity of different nasal sounds varies physiologically, therefore their corresponding standard threshold ranges need to be set differently. Based on this, in this embodiment, a standard nasalization database is pre-constructed based on standard nasal sound pronunciation data from native English speakers. Using this database, the DTW algorithm is employed to determine the standard threshold ranges for different nasal sounds in different pronunciation contexts, ensuring the accuracy of nasal sound integrity judgment in various contexts.
[0060] In a specific example, based on a large amount of pronunciation data of / m, n, ŋ / from native English speakers, the distribution of nasalization of different nasal sounds in different pronunciation contexts (including solo pronunciation, connected speech, and stress position) was statistically analyzed to determine the standard threshold range corresponding to "complete pronunciation". For example, in the pronunciation context of solo pronunciation, the standard threshold range corresponding to complete pronunciation of / m / (bilabial nasal) is [0.65, 0.85]; the standard threshold range corresponding to complete pronunciation of / n / (alveolar nasal) is [0.55, 0.75]; and the standard threshold range corresponding to complete pronunciation of / ŋ / (velar nasal) is [0.50, 0.70].
[0061] In this embodiment, the lip opening curve is reconstructed as follows.
[0062] While acquiring speech signals, the system also acquires visual data of the learner's lips and precisely aligns the timing of speech and lip movements. The visual data of the lips can be acquired using a facial motion capture system or a high-definition lip-shape camera. The visual data of the lips is preprocessed (including grayscale conversion, noise reduction, and edge detection), the region of interest of the lips is extracted, and key point detection algorithms (such as Dlib library and MediaPipe facial mesh model) are used to locate multiple core lip feature points such as the upper lip edge, lower lip edge, cupid's bow, and corner of the mouth. Based on the location of multiple core lip feature points, the vertical span value of the lip contour is determined. Using the Euclidean distance between the midpoint feature point of the upper lip and the midpoint feature point of the lower lip among the located core lip feature points as an indicator, combined with the determined vertical span value of the lip contour, the lip opening and closing value of each frame is calculated. The lip opening and closing values of each frame are concatenated along the timeline to form the original temporal curve of lip opening and closing. The original time-series curve of lip opening and closing is smoothed (moving average filtering can be used) to extract the dynamic features of the processed curve, such as peak value, valley value, rise rate, fall rate, and duration, and the curve is transformed into a quantifiable original parameter set to form the original parameter set of lip opening and closing. The system retrieves a standard curve library of lip opening and closing for different phonemes (primarily vowels, semi-vowels, and labial consonants) based on the standard pronunciation of native English speakers. This library contains the standard parameter range for each phoneme in different contexts. The original parameter set of lip opening and closing is matched and calibrated with the standard curves to correct errors caused by individual lip shape differences and sampling angle deviations. The final lip opening and closing curve and parameter values are obtained. Subsequently, the system combines the corresponding standard thresholds to judge deviations such as insufficient, excessive, or misaligned lip opening and closing amplitude during pronunciation, thereby enabling the determination of lip shape-related pronunciation abnormalities.
[0063] Step 330: The phoneme-level feature data, the prosodic feature data, and the physiological motion feature data are fused to obtain cross-modal fused feature data.
[0064] In this embodiment, for determined physiological motion feature data, it is transformed into a multi-dimensional numerical data combination such as datasets and vectors, and then merged with the phoneme-level feature data and the prosodic feature data to obtain cross-modal fusion feature data. It is worth noting that different physiological motion features have their own corresponding preset transformation rules for data transformation.
[0065] It is important to understand that traditional speech feature extraction methods do not consider the movement state of the speech organs, resulting in an incomplete analysis of the causes of pronunciation errors, which in turn affects the localization of pronunciation errors. Based on this, this embodiment, leveraging the intrinsic correlation between acoustic features and the physiological movements of speech, associates acoustic signals with physiological states such as tongue tip height, nasal pronunciation integrity, and lip opening / closing, forming an "acoustic-physiological" cross-modal mapping and constructing a cross-modal feature fusion mechanism.
[0066] In a specific example, the method constructs a feature extraction model. The input to this model includes the speech data (a high-fidelity speech signal) and standard pronunciation feature templates. The output is a "multi-dimensional, stereoscopic, and cross-modal" pronunciation feature map (i.e., a cross-modal feature set), which covers all pronunciation details and can comprehensively support subsequent pronunciation error tracing. In other words, this model has the function of determining phoneme-level feature data, prosodic feature data, and physiological motion feature data, as well as fusing these types of feature data.
[0067] The standard pronunciation feature template consists of pre-determined input data used as a comparison benchmark. It includes: standard phoneme acoustic features (such as the nasalization range of standard / m, n, ŋ / , and the VOT standard values of / p / and / b / ), standard prosodic curves (such as pitch change trajectory, stress distribution pattern, and rhythmic sequence template), and standard physiological correlation parameters (such as the tongue tip height range and lip opening curve during standard vowel pronunciation). These parameters are used to quantify the deviation from the learner's pronunciation features.
[0068] For the extended phoneme-level features, the model adds three key parameters to the traditional 39-dimensional MFCC parameters: dynamic formant trajectory parameters, voice onset time parameters, and fricative energy distribution parameters. This not only preserves the basic acoustic features but also makes up for the shortcomings of the traditional model in capturing subtle differences in pronunciation.
[0069] For temporal prosodic features, the model extracts them based on the temporal analysis model of LSTM-RNN, focusing on the core of spoken fluency, namely speech contour parameters, stress distribution parameters and rhythm deviation parameters, to achieve dynamic and quantitative presentation of prosodic features.
[0070] For cross-modal physiological motion characteristics, the physiological motion state is inferred from acoustic features. Specifically, this includes: tongue tip height parameters calculated based on the rate of change of formants F1 and F4 (used to quantify tongue movement), nasalization index for judging the integrity of nasal pronunciation (used to avoid abnormal nasal resonance), and reconstructed lip opening curve (used to reflect the amplitude of lip movement), so as to achieve cross-modal mapping from "acoustic signal" to "physiological motion".
[0071] These outputs together constitute a comprehensive pronunciation feature map, providing multi-dimensional data support for subsequent error prediction models, ensuring that pronunciation error localization can delve from the phoneme level to the speech organ movement level, and achieving accurate discovery of pronunciation errors.
[0072] As can be seen, step 300 upgrades speech feature extraction from a "single acoustic dimension" to a cross-modal fusion of "acoustic-physiological" features. It also expands from a "single dimension" to a "multi-dimensional" approach, realizing a full-link characterization of the pronunciation process. This provides multi-dimensional data support for subsequent pronunciation error prediction and pronunciation training strategies (i.e., early intervention plans for errors).
[0073] Furthermore, it's important to understand that by conducting multi-level quantitative analysis of learners' acoustic characteristics (such as vowel formant distribution, vocal onset time, and speech rate rhythm fluctuations), physiological motor parameters (such as tongue tip height, lip opening and closing, and soft palate closure), and potential pronunciation error patterns (such as phoneme confusion caused by native language transfer), a learner-specific pronunciation feature atlas can be constructed. This enables end-to-end reverse modeling from acoustic signals to speech organ movements, allowing for precise localization of phoneme-level pronunciation defects and revealing the correlation networks of error causes (such as the mechanical correlation between / r / sounds errors in native Chinese speakers and insufficient tongue tip retraction). This provides cross-dimensional data support for developing personalized pronunciation training strategies for learners. Potential pronunciation error patterns can be determined based on the learner's historical English pronunciation training.
[0074] Step 400: Input the cross-modal fusion feature data into the pre-trained error occurrence prediction model, and output the error occurrence prediction result through the error occurrence prediction model.
[0075] The error prediction results include the types of pronunciation errors that learners may make in the future.
[0076] Among them, pronunciation errors are categorized into phoneme-level errors, prosodic-level errors, and mother tongue transfer errors.
[0077] In one specific example, the error occurrence prediction model takes cross-modal fusion feature data as input and outputs a list of high-probability pronunciation errors and pronunciation error types.
[0078] In this specific example, each training data in the training dataset used to train the error occurrence prediction model includes cross-modal fusion feature data, a list of high-probability pronunciation errors, and pronunciation error types.
[0079] In this specific example, step 400 is performed as follows: The cross-modal fusion feature data is input into the error occurrence prediction model. The error occurrence prediction model dynamically analyzes the risk of phoneme confusion (such as the spectral overlap tendency of / θ / and / s / ), prosodic deviation trends (such as the rules of stress position shift), and mother tongue transfer paths (such as the interference of Chinese tones on English stress). It predicts the specific pronunciation errors that learners may make in the next preset number of pronunciation training sessions (the value can be selected from 3 to 5) (such as insufficient retraction of the tongue tip of / r / , omission of linking rules) and their probability of occurrence.
[0080] It should be noted that each type of pronunciation error includes multiple specific pronunciation error points. Furthermore, the confidence level of the probability of occurrence predicted by the error occurrence prediction model is greater than a preset confidence threshold (selectable as 85%). That is, for a specific pronunciation error point, the error occurrence prediction model will only consider it a "possible specific pronunciation error point" and determine its probability of occurrence if the confidence level of its prediction of the learner's likelihood of encountering that specific pronunciation error point in a preset number of future pronunciation training sessions is greater than the preset confidence threshold.
[0081] Then, based on the specific pronunciation errors that learners may make in future preset pronunciation training sessions and their probabilities, the types of pronunciation errors that learners may make in the future are determined.
[0082] Specifically, the pronunciation error type to which a specific pronunciation error point with a probability greater than or equal to a preset probability threshold belongs is determined as the pronunciation error type that the learner may make in the future.
[0083] In another specific example, the error occurrence prediction model also employs a deep learning framework based on a bidirectional spatiotemporal attention mechanism. Its inputs include cross-modal fusion feature data, learners' potential pronunciation error patterns, and learners' native language background encoding (unique and pre-compiled). Its outputs include a list of high-probability pronunciation errors and pronunciation error types.
[0084] In this specific example, each training data in the training dataset used to train the error occurrence prediction model includes cross-modal fusion feature data, potential pronunciation error patterns, native language background encoding, a list of high-probability pronunciation errors, and pronunciation error types.
[0085] In this specific example, a data storage module is deployed to store the learner's historical pronunciation error types predicted by the error occurrence prediction model, with each historical pronunciation error type corresponding to a historical moment.
[0086] In this specific example, the method for identifying the learner's potential pronunciation error patterns is as follows: If the data storage module stores the learner's historical pronunciation error types in this English pronunciation training task, then the learner's latest historical pronunciation error type is identified as a potential pronunciation error pattern, or the proportion of each of the learner's different pronunciation error types is determined, and the pronunciation error type with the highest proportion is identified as a potential pronunciation error pattern. If the data storage module does not store the learner's historical pronunciation error types in this English pronunciation training task, then the preset pronunciation error type corresponding to the learner's native language background or the potential pronunciation error pattern corresponding to the learner in the previous English pronunciation training task stored in the data storage module will be determined as the potential pronunciation error pattern.
[0087] Among them, the potential pronunciation error patterns of learners in the previous English pronunciation training task can be determined based on the pronunciation error types predicted in the previous English pronunciation training task (specifically, it can be determined based on the proportion of different pronunciation error types).
[0088] In this specific example, step 400 is performed as follows: First, after determining the learner's potential pronunciation error patterns and native language background encoding, the cross-modal fusion feature data, the learner's potential pronunciation error patterns, and the learner's native language background encoding are input into the error occurrence prediction model. The error occurrence prediction model dynamically analyzes the risk of phoneme confusion (such as the spectral overlap tendency of / θ / and / s / ), prosodic deviation trends (such as the rules of stress position shift), and native language transfer paths (such as the interference of Chinese tones on English stress). It predicts the specific pronunciation errors that the learner may make in the next preset number of pronunciation training sessions (the value can be selected from 3 to 5) (such as insufficient retraction of the tongue tip of / r / , omission of linking) and their probability of occurrence.
[0089] Similar to the previous specific example, each type of pronunciation error includes multiple specific pronunciation error points, and the confidence level of the occurrence probability predicted by the error occurrence prediction model is greater than the preset confidence threshold.
[0090] Then, based on the specific pronunciation errors that learners may make in future preset pronunciation training sessions and their probabilities, a list of high-probability pronunciation errors and the types of pronunciation errors that learners may make in the future are determined.
[0091] In this specific example, a list of high-probability pronunciation errors is generated based on specific pronunciation errors that have a probability of occurrence greater than or equal to a preset probability threshold. The pronunciation error type to which each specific pronunciation error point in the list of high-probability pronunciation errors belongs is determined as the pronunciation error type that the learner may make in the future.
[0092] As can be seen, compared to the traditional reactive "pronunciation → detection → correction" model, learners are more likely to develop muscle memory for incorrect pronunciation. Step 400, by identifying potential pronunciation errors in advance, allows for proactive intervention, transforming passive error correction into proactive risk prediction. This provides a "targeted goal" for adjusting the learner's English pronunciation training program. This embodiment's shift from "post-correction" to "proactive intervention" reduces the risk of error entrapment at the source.
[0093] Step 500: If the error occurrence prediction result meets the preset dynamic adjustment conditions, determine the target early intervention plan based on the error occurrence prediction result, and adjust the learner's current English pronunciation training plan according to the target early intervention plan.
[0094] In this embodiment, the preset dynamic adjustment condition includes the fact that the learner's potential future pronunciation error types are not empty. That is, intervention in the learner's pronunciation is only carried out when the error occurrence prediction model outputs at least one of the pre-defined pronunciation error types.
[0095] Taking the specific example provided in step 400, if the learner's confidence in a certain pronunciation error point in a future preset pronunciation training session is greater than the preset confidence threshold, step 500 is triggered, that is, a preventive intervention strategy is triggered (such as preset contrastive phoneme training, inserting auditory interference to break the automatic error response), thereby blocking its formation link before the error actually occurs.
[0096] In this embodiment, step 500 determines a target early intervention plan based on the error occurrence prediction result, including the following steps.
[0097] Step 510: If the pronunciation error type in the error occurrence prediction result includes phoneme-level errors, then add preset contrastive phoneme training (such as practicing the pronunciation of / θ / and / s / in pairs); if the pronunciation error type in the error occurrence prediction result includes prosodic errors, then add standard pronunciation demonstrations and rhythm guidance; if the pronunciation error type in the error occurrence prediction result includes mother tongue transfer errors, then add analysis of differences between mother tongue and English pronunciation to help learners establish correct pronunciation cognition.
[0098] It can be seen that, compared with the "passive correction" mode of traditional pronunciation correction, the English pronunciation correction method provided in this application adopts the "prediction → active intervention" mode to dynamically adjust the learner's current English pronunciation training plan, which can significantly improve correction efficiency and reduce the risk of erroneous memory solidification.
[0099] Optionally, in this embodiment, the method further includes the following steps.
[0100] Acquire physiological and cognitive characteristic data of learners' English pronunciation, as well as specific parameter data used to quantify the learner's emotional state; Based on the specific parameter data, the learner's emotional state is determined; Based on the physiological feature data, the cognitive feature data, the emotional state, the cross-modal fusion feature data, and the error occurrence prediction results, the learner's current English pronunciation training program is adjusted using a pre-established adaptive decision engine.
[0101] It should be noted that the operation of obtaining physiological characteristic data, cognitive characteristic data and specific parameter data for quantifying the learner's emotional state of English pronunciation can be performed simultaneously with step 400 to reduce wasted time.
[0102] It is important to understand that traditional pronunciation correction focuses solely on the acoustic characteristics of pronunciation, neglecting the differences in learners' physiological conditions and cognitive patterns. This results in a lack of personalized adaptation in pronunciation training strategies (such as using a uniform English pronunciation training program for learners with insufficient tongue flexibility). Therefore, this embodiment acquires physiological characteristic data, cognitive characteristic data, and specific parameter data to provide learners with "personalized + emotionally adapted" pronunciation correction solutions.
[0103] In this embodiment, physiological characteristics include tongue flexibility and respiratory support, while cognitive characteristics include pronunciation mastery speed and forgetting rate. It is important to understand that pronunciation is the result of physiological movement, and tongue flexibility and respiratory support are core physiological factors affecting pronunciation quality. This embodiment acquires physiological and cognitive characteristic data to accurately pinpoint the underlying causes of errors, i.e., to determine whether they are due to physiological limitations or cognitive deficiencies.
[0104] In a specific example, a dynamic detection mechanism for vocal tract parameters is established to detect the learner's tongue flexibility and respiratory support, thereby obtaining physiological characteristic data constructed from the comprehensive data of tongue flexibility and respiratory support. Based on a pre-constructed Hidden Markov Model (HMM), the learner's learning process is tracked, and the accuracy improvement curve and forgetting rate are obtained. Cognitive characteristic data are then constructed based on the accuracy improvement curve and forgetting rate.
[0105] Specifically, this example quantifies tongue flexibility based on rapid alternating pronunciation tasks (such as rapid switching between / la-li-lu / ) to identify whether pronunciation problems (such as difficulty pronouncing / r / ) are caused by insufficient tongue movement flexibility; and obtains respiratory support parameters based on the analysis of vowel sound pressure level attenuation curves to determine whether there are problems such as weak pronunciation or dropout of the final consonant due to insufficient breath.
[0106] This specific example uses Hidden Markov Models (HMMs) to track learners' learning process throughout its entire lifecycle. It statistically analyzes the accuracy improvement curve of new pronunciations (a fitted exponential function) to reflect the learner's speed of mastering specific pronunciations. Based on the HMM, it calculates the learner's forgetting rate through spaced repetitive testing to identify easily forgotten pronunciation knowledge points (such as special linking rules). In this example, the HMM input is a time series of correct pronunciations, and the output is the accuracy improvement curve and the forgetting rate.
[0107] It is important to understand that this embodiment, through dual-dimensional detection of physiological and cognitive characteristics, can accurately pinpoint the underlying causes of pronunciation errors (i.e., determine whether it is due to physiological limitations or insufficient cognitive memory), providing a personalized basis for subsequent generation of adaptive pronunciation training strategies based on an adaptive decision engine, thus avoiding a "one-size-fits-all" training model.
[0108] Furthermore, it's important to understand that in English oral training, learners' anxiety and distraction can severely impact training effectiveness (e.g., anxiety leads to pronunciation tension, and lack of concentration increases the error rate), while traditional pronunciation correction neglects the influence of emotional factors. Therefore, this implementation incorporates emotional factors into the pronunciation correction system, establishing an emotion quantification model to achieve dynamic adaptation between pronunciation training strategies and learners' emotional states.
[0109] In this embodiment, specific parameters used to quantify the learner's emotional state include an anxiety index and an attention distraction index. Based on the specific parameter data, the learner's emotional state is determined, including the following steps.
[0110] Based on anxiety index data and attention distraction index data, a weighted calculation method was used to quantify the emotion index data. The learner's emotional state is determined based on the aforementioned emotion index data.
[0111] In a specific example, eye-tracking technology is used to acquire learners' anxiety index data (such as pupil changes and blink frequency) and attention distraction index data (such as the duration of time the gaze is off the screen). Based on the anxiety index data and attention distraction index data, and combined with the preset weights corresponding to the anxiety index and attention distraction index respectively, the emotion index data is quantified using a weighted calculation method.
[0112] In this specific example, the sum of the preset weights for the anxiety index and the attention distraction index is 1, and the numerical range of the emotion index data is [0,1]. The emotion state includes three levels: mild, moderate, and severe. The emotion index data range for mild is [0,0.6); the emotion index data range for moderate is [0.6,0.8); and the emotion index data range for severe is [0.8,1].
[0113] Specifically, a mild state indicates that the learner's anxiety / distraction level is low, so native language comparison demonstrations can be introduced to strengthen the recognition of pronunciation differences and maintain the training rhythm; a moderate state indicates that the learner's anxiety / distraction level is moderate, so music can be switched and the learner's breathing can be guided to reduce anxiety through soothing music and adjust the training intensity (such as shortening the duration of a single training session); a severe state indicates that the learner's anxiety / distraction level is high, so the learner can be guided to perform a preset duration (3-5 minutes) of breathing relaxation and attention focus training, and pronunciation training can be resumed after the learner's emotions have calmed down.
[0114] As can be seen, this embodiment introduces emotional factors into the pronunciation correction system and adjusts the English pronunciation training program in real time in a targeted manner. It can quickly achieve dynamic adaptation between "pronunciation training strategy and emotional state", effectively improve the learner's training experience and concentration, and avoid the decline in training efficiency caused by the learner's emotional problems.
[0115] In this embodiment, the adaptive decision engine includes a state space, an action space, and a reward function, wherein: The state space, including multi-dimensional parameters, is used to comprehensively characterize the learner's real-time state.
[0116] The multi-dimensional parameters include pronunciation accuracy, emotional state, potential pronunciation error patterns, physiological characteristics, and cognitive characteristics.
[0117] The action space includes adjusting training intensity (such as increasing / decreasing the number of practice sessions), selecting learning feedback methods (such as voice demonstration comparison / text analysis / animation demonstration of the articulatory organs), and injecting knowledge points (such as inserting explanations of pronunciation techniques) to meet diverse pronunciation correction needs.
[0118] The reward function is used to guide the adaptive decision engine to generate an optimized pronunciation training strategy based on indicators such as improved pronunciation accuracy, reduced error repetition rate, and optimized learning time, so as to adjust the learner's current English pronunciation training plan.
[0119] Among them, the learning feedback method selection provides learners with a variety of feedback methods, making it easier for learners to understand the advantages and areas for improvement of the current pronunciation training.
[0120] In a specific example, the adaptive decision engine is built based on the Proximal Policy Optimization (PPO) algorithm to achieve intelligent and personalized dynamic generation of pronunciation training strategies. The state space contains 28 dimensions of parameters, and the action space contains 12 classes of operations. It is worth noting that the PPO algorithm is a reinforcement learning algorithm used to improve policy performance while avoiding drastic fluctuations during training, making it suitable for handling problems with high-dimensional state spaces and complex decision-making.
[0121] It is important to understand that traditional pronunciation correction training strategies are mostly based on preset fixed procedures, which cannot be dynamically adjusted according to the learner's real-time state (such as pronunciation accuracy, emotional state, and pronunciation error patterns), resulting in low training efficiency or accumulated learner fatigue. Based on this, this embodiment establishes an adaptive decision engine, based on a four-dimensional data system of "acoustics-physiology-cognition-emotion," to achieve optimization from "fixed strategies" to "dynamic adaptation." It can dynamically adjust the English pronunciation training plan according to the learner's real-time state and individual differences, realizing personalized pronunciation correction with "one plan for one person," ensuring that subsequent pronunciation training operations are accurate and tailored to the learner's individual state.
[0122] Specifically, the training difficulty can be adjusted according to physiological characteristics (e.g., if the tongue is not flexible enough, the proportion of rapid pronunciation tasks can be reduced), the training frequency can be optimized according to cognitive characteristics (e.g., if the forgetting rate is high, the interval repetition test can be increased), the training rhythm can be adjusted according to emotional state (e.g., if there is moderate anxiety, the duration of each session can be shortened and soothing music can be played), and the training focus can be optimized according to the error occurrence prediction results (e.g., if the connected speech error is predicted, connected speech-specific practice can be increased), thereby achieving comprehensive dynamic adaptation.
[0123] Furthermore, in this embodiment, the method further includes the following steps.
[0124] Collect core data on learners completing a full English pronunciation training task.
[0125] The core data includes pronunciation accuracy, error repetition rate, emotion change curve, and training duration.
[0126] Based on the core data, the learner's potential pronunciation error patterns, cognitive feature curves, and node data of the pre-constructed personalized pronunciation tree are updated, and the parameters of the error occurrence prediction model and the reward function of the adaptive decision engine are optimized.
[0127] The personalized pronunciation tree is used to create new English pronunciation training tasks for learners; each node of the personalized pronunciation tree includes a heatmap of historical potential pronunciation error frequencies, the optimal pronunciation correction path, and warnings of potential related pronunciation errors.
[0128] In a specific example, a personalized pronunciation tree for different learners is constructed based on a random forest model. By determining the distribution of high-frequency errors, the optimal pronunciation correction path, and warnings of potential related pronunciation errors, the pronunciation training content is ensured to focus on risk points while also being tailored to individual differences.
[0129] Within the nodes of the personalized pronunciation tree: Historical potential pronunciation error frequency heatmap: used to visually present the predicted high-frequency pronunciation error types and distribution of learners; Optimal Pronunciation Correction Path: This is used to recommend the best English pronunciation training program based on pronunciation correction data of similar learners and the learner's historical pronunciation training results. Potential related pronunciation error warning: Used to indicate other possible errors associated with a specific pronunciation error that has been predicted so far (such as the / r / sound error that may cause linking problems).
[0130] This embodiment optimizes the next English pronunciation training task by performing data feedback update processing to collect core data and update data based on the core data, forming a closed loop of "training-feedback-update". This makes subsequent pronunciation training more suitable for learners, thereby continuously improving the accuracy and efficiency of correction.
[0131] As can be seen, this embodiment, based on an adaptive decision engine, ensures that the adjusted English pronunciation training program is personalized and emotionally adaptable, and the data feedback operation based on core data enables continuous optimization of pronunciation correction. This embodiment, with its overall operational sequence and logical relationship of "extracting cross-modal fusion feature data based on speech data → supplementing physiological, cognitive, and emotional feature data in parallel → predicting pronunciation errors based on cross-modal fusion feature data → integrating all feature data for early intervention and dynamic training adjustments → executing the English pronunciation training program + data feedback closed loop," ultimately achieves the core objectives of accurate risk prediction, personalized adjustment strategies, and closed-loop optimization.
[0132] The English pronunciation correction method provided in this embodiment has the following advantages: Improved accuracy of correction: By using a four-channel ring microphone array and cross-modal feature fusion technology, the accuracy of speech signal capture and feature extraction is greatly improved. The accuracy of pronunciation error localization is refined from the phoneme level of traditional methods to the movement level of the speech organs, which increases the accuracy of pronunciation error localization to over 92%. Through the error occurrence prediction model, potential pronunciation errors can be identified in advance, which increases the success rate of pronunciation error correction by over 35%.
[0133] Personalized adaptation capabilities are comprehensively enhanced: The combination of physiological and cognitive dual-dimensional feature detection and adaptive decision engine enables pronunciation training strategies to accurately adapt to the physiological conditions (such as tongue flexibility and respiratory support), cognitive patterns (such as learning speed and forgetting rate) and learning habits of different learners, completely changing the limitations of the traditional "one-size-fits-all" approach and achieving personalized correction for each individual.
[0134] Dual optimization of training efficiency and learning experience: The error early intervention mechanism shortens the correction cycle (average correction cycle shortened by 40%), and the dynamic adjustment strategy effectively reduces learners' anxiety and fatigue (anxiety index decreased by an average of 28%); diversified feedback methods and training intensity adjustments enhance the fun and focus of learning, and learners' average daily training time and persistence rate are significantly improved.
[0135] Strong generalization ability: It is adaptable to learning systems in multiple scenarios. The signal processing algorithms and model architecture adopted can adapt to noise interference in different learning environments (such as home, library, office), support the personalized needs of learners at different levels (from beginner to advanced), and be compatible with different devices (mobile phones, tablets, computers), with a wide range of application scenarios and user coverage.
[0136] Figure 1 This is a flowchart illustrating an English pronunciation correction method in one embodiment. It should be understood that, although... Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise explicitly stated herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0137] Figure 2 A schematic block diagram illustrating the structure of an English pronunciation correction device according to an embodiment of this application is shown. Figure 2 As shown in one embodiment of this application, an English pronunciation correction device is provided, which may include the following functional modules.
[0138] The speech data acquisition module is used to acquire speech data of learners' English pronunciation; A cross-modal fusion module is used to determine cross-modal fusion feature data based on the speech data; wherein the cross-modal fusion feature data is fusion data of acoustic dimension and physiological dimension; The potential error prediction module is used to input the cross-modal fusion feature data into a pre-trained error occurrence prediction model, and output error occurrence prediction results through the error occurrence prediction model; wherein, the error occurrence prediction results include the types of pronunciation errors that the learner may make in the future; The training program adjustment module is used to determine a target early intervention plan based on the error occurrence prediction results when the error occurrence prediction results meet the preset dynamic adjustment conditions, and to adjust the learner's current English pronunciation training plan according to the target early intervention plan.
[0139] In one embodiment, the apparatus further includes an optimization adjustment module, the optimization adjustment module being used for: Acquire physiological and cognitive characteristic data of learners' English pronunciation, as well as specific parameter data used to quantify the learner's emotional state; Based on the specific parameter data, the learner's emotional state is determined; Based on the physiological feature data, the cognitive feature data, the emotional state, the cross-modal fusion feature data, and the error occurrence prediction results, the learner's current English pronunciation training program is adjusted using a pre-established adaptive decision engine.
[0140] In one embodiment, determining cross-modal fusion feature data based on the speech data includes: Feature extraction is performed on the speech data to obtain phoneme-level feature data and prosodic feature data; wherein, the phoneme-level features include dynamic formant trajectory parameters, voice onset time parameters, and fricative energy distribution parameters, and the prosodic features include speech contour parameters, stress distribution parameters, and rhythm deviation parameters; Based on the speech data, physiological motor characteristic data are determined; among which, physiological motor characteristics include tongue tip height, nasal pronunciation integrity, and lip opening and closing curve; The phoneme-level feature data, the prosodic feature data, and the physiological motion feature data are fused to obtain cross-modal fused feature data.
[0141] In one embodiment, pronunciation error types are categorized into phoneme-level errors, prosodic errors, and mother tongue transfer errors; Based on the error occurrence prediction results, a target early intervention plan is determined, including: If the pronunciation error type in the error occurrence prediction result includes phoneme-level errors, then preset contrastive phoneme training is added; If the pronunciation error type in the error prediction result includes prosodic errors, then standard pronunciation demonstration and rhythm guidance will be added; If the pronunciation error type in the error occurrence prediction result includes mother tongue transfer errors, then the analysis of differences between mother tongue and English pronunciation will be added.
[0142] In one embodiment, specific parameters include an anxiety index and an attention distraction index; Based on the specific parameter data, the learner's emotional state is determined, including: Based on anxiety index data and attention distraction index data, a weighted calculation method was used to quantify the emotion index data. The learner's emotional state is determined based on the aforementioned emotion index data.
[0143] In one embodiment, the adaptive decision engine includes: The state space, which includes multi-dimensional parameters, is used to characterize the learner's real-time state. These multi-dimensional parameters include pronunciation accuracy, emotional state, potential pronunciation error patterns, physiological characteristics, and cognitive characteristics. Action space, including training intensity adjustment, learning feedback method selection, and knowledge point injection; The reward function is used to guide the adaptive decision engine to generate an optimized pronunciation training strategy based on indicators such as improved pronunciation accuracy, reduced error repetition rate, and optimized learning time, so as to adjust the learner's current English pronunciation training plan.
[0144] In one embodiment, the apparatus further includes an update module, the update module being used to: Collect core data on learners completing a full English pronunciation training task; the core data includes pronunciation accuracy, error repetition rate, emotion change curve, and training duration. Based on the core data, the node data of the learner's pre-built personalized pronunciation tree is updated, and the parameters of the error occurrence prediction model and the reward function of the adaptive decision engine are optimized; wherein, the personalized pronunciation tree is used to create new English pronunciation training tasks for learners.
[0145] Since the English pronunciation correction device provided in this application embodiment is a virtual device corresponding to the English pronunciation correction method of the above embodiment, it can also solve the problem that the dynamic adjustment of the English pronunciation correction method using the static feedback mode in the prior art is poor, which affects learners' learning of English pronunciation.
[0146] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the English pronunciation correction method described in the above embodiments.
[0147] The electronic device provided in this application embodiment includes a processor capable of running the English pronunciation correction method of the aforementioned embodiment. Therefore, it can also solve the problem that the dynamic adjustment of the English pronunciation correction method using the static feedback mode in the prior art is poor, which affects learners' learning of English pronunciation.
[0148] This application provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the English pronunciation correction method described in the above embodiments.
[0149] The machine-readable storage medium provided in this application embodiment stores instructions for causing the machine to execute the English pronunciation correction method of the above embodiment. Therefore, it can also solve the problem that the dynamic adjustment of the English pronunciation correction method using the static feedback mode in the prior art is poor, which affects learners' learning of English pronunciation.
[0150] Figure 3 The diagram schematically illustrates the internal structure of a computer device according to an embodiment of this application. Figure 3 As shown in one embodiment of this application, a computer device is provided, which can be a terminal. The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05, and a memory (not shown) connected via a system bus. The processor A01 provides computing and control capabilities. The memory includes internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 stored in the non-volatile storage medium A06. The network interface A02 is used to communicate with an external terminal via a network connection. When the computer program B02 is executed by the processor A01, it implements an English pronunciation correction method. The display screen A04 of the computer device can be an LCD screen or an e-ink screen. The input device A05 of the computer device can be a touch layer covering the display screen A04, or a button, trackball, or touchpad set on the computer device casing, or an external keyboard, touchpad, or mouse, etc.
[0151] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0152] In one embodiment, the English pronunciation correction device provided in this application can be implemented as a computer program, which can be implemented in the form of, for example, Figure 3 The device operates on the computer shown. The computer's memory can store various program modules that make up the English pronunciation correction device, and the computer program composed of these program modules causes the processor A01 to execute the steps in the English pronunciation correction methods of the various embodiments of this application described in this specification.
[0153] Figure 3 The computer equipment shown can be used as follows Figure 2 The speech data acquisition module in the English pronunciation correction device shown performs step 200, the cross-modal fusion module performs step 300, the potential error prediction module performs step 400, and the training scheme adjustment module performs step 500.
[0154] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0155] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0156] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0157] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0158] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0159] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0160] Computer-readable media include both permanent and non-permanent, removable and non-removable media, which can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0161] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0162] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An English pronunciation correction method, characterized by, The method includes: Acquire the speech data of learners' English pronunciation; Based on the speech data, cross-modal fusion feature data is determined; wherein, the cross-modal fusion feature data is the fusion data of acoustic dimension and physiological dimension; The cross-modal fusion feature data is input into a pre-trained error occurrence prediction model, and the error occurrence prediction model outputs error occurrence prediction results; wherein, the error occurrence prediction results include the types of pronunciation errors that the learner may make in the future; If the error occurrence prediction result meets the preset dynamic adjustment conditions, a target early intervention plan is determined based on the error occurrence prediction result, and the learner's current English pronunciation training plan is adjusted according to the target early intervention plan.
2. The English pronunciation correction method of claim 1, wherein, The method further includes: Acquire physiological and cognitive characteristic data of learners' English pronunciation, as well as specific parameter data used to quantify the learner's emotional state; Based on the specific parameter data, the learner's emotional state is determined; Based on the physiological feature data, the cognitive feature data, the emotional state, the cross-modal fusion feature data, and the error occurrence prediction results, the learner's current English pronunciation training program is adjusted using a pre-established adaptive decision engine.
3. The English pronunciation correction method of claim 1, wherein, Based on the speech data, cross-modal fusion feature data is determined, including: Feature extraction is performed on the speech data to obtain phoneme-level feature data and prosodic feature data; wherein, the phoneme-level features include dynamic formant trajectory parameters, voice onset time parameters, and fricative energy distribution parameters, and the prosodic features include speech contour parameters, stress distribution parameters, and rhythm deviation parameters; Based on the speech data, physiological motor characteristic data are determined; among which, physiological motor characteristics include tongue tip height, nasal pronunciation integrity, and lip opening and closing curve; The phoneme-level feature data, the prosodic feature data, and the physiological motion feature data are fused to obtain cross-modal fused feature data.
4. The method of claim 1, wherein, Pronunciation errors are categorized into phoneme-level errors, prosodic-level errors, and errors resulting from mother tongue transfer. Based on the error occurrence prediction results, a target early intervention plan is determined, including: If the pronunciation error type in the error occurrence prediction result includes phoneme-level errors, then preset contrastive phoneme training is added; If the pronunciation error type in the error prediction result includes prosodic errors, then standard pronunciation demonstration and rhythm guidance will be added; If the pronunciation error type in the error occurrence prediction result includes mother tongue transfer errors, then the analysis of differences between mother tongue and English pronunciation will be added.
5. The method of claim 2, wherein the method further comprises: Specific parameters include anxiety index and attention deficit index; Based on the specific parameter data, the learner's emotional state is determined, including: Based on anxiety index data and attention distraction index data, a weighted calculation method was used to quantify the emotion index data. The learner's emotional state is determined based on the aforementioned emotion index data.
6. The English pronunciation correction method according to claim 2, characterized in that, The adaptive decision engine includes: The state space, which includes multi-dimensional parameters, is used to characterize the learner's real-time state. These multi-dimensional parameters include pronunciation accuracy, emotional state, potential pronunciation error patterns, physiological characteristics, and cognitive characteristics. Action space, including training intensity adjustment, learning feedback method selection, and knowledge point injection; The reward function is used to guide the adaptive decision engine to generate an optimized pronunciation training strategy based on indicators such as improved pronunciation accuracy, reduced error repetition rate, and optimized learning time, so as to adjust the learner's current English pronunciation training plan.
7. The English pronunciation correction method according to claim 2, characterized in that, The method further includes: Collect core data on learners completing a full English pronunciation training task; the core data includes pronunciation accuracy, error repetition rate, emotion change curve, and training duration. Based on the core data, the node data of the learner's pre-built personalized pronunciation tree is updated, and the parameters of the error occurrence prediction model and the reward function of the adaptive decision engine are optimized; wherein, the personalized pronunciation tree is used to create new English pronunciation training tasks for learners.
8. An English pronunciation correction device, characterized in that, The device includes: The speech data acquisition module is used to acquire speech data of learners' English pronunciation; A cross-modal fusion module is used to determine cross-modal fusion feature data based on the speech data; wherein the cross-modal fusion feature data is fusion data of acoustic dimension and physiological dimension; The potential error prediction module is used to input the cross-modal fusion feature data into a pre-trained error occurrence prediction model, and output error occurrence prediction results through the error occurrence prediction model; wherein, the error occurrence prediction results include the types of pronunciation errors that the learner may make in the future; The training program adjustment module is used to determine a target early intervention plan based on the error occurrence prediction results when the error occurrence prediction results meet the preset dynamic adjustment conditions, and to adjust the learner's current English pronunciation training plan according to the target early intervention plan.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the English pronunciation correction method according to any one of claims 1 to 7.
10. A machine-readable storage medium storing instructions thereon, characterized in that, When executed by a processor, the instruction causes the processor to be configured to perform the English pronunciation correction method as described in any one of claims 1 to 7.