English auxiliary pronunciation training method and system based on multi-task learning
By adopting a multi-task learning framework and multi-modal data fusion technology in the English pronunciation training system, the shortcomings of error positioning and training path adjustment in the existing system are solved, precise evaluation and personalized training are achieved, and training efficiency and effect are improved.
Patent Information
- Application Number
- CN202510164169.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing English pronunciation training system lacks deep integration of multimodal data, making it difficult to achieve accurate mispositioning and personalized feedback, and the training path is fixed and cannot be adjusted dynamically, resulting in inefficient training.
Using a framework based on multi-task learning, voice signals, lip motion videos and text content are collected through the multi-modal data input module, cross-modal features are generated using a shared feature extraction layer, dynamically allocate the weight of the sub-task model in combination with the attention mechanism between tasks, and personalized training strategies are generated based on the sub-task output.
Accurate evaluation and personalized training of English pronunciation are realized, training effects and efficiency are improved, and an efficient, convenient and personalized English pronunciation training environment is provided.
Smart Images

Figure CN119993197A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of language learning, and in particular to an English auxiliary pronunciation training method and system based on multi-task learning. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, English pronunciation auxiliary training systems based on speech recognition have gradually become a research hotspot in the field of language learning. Traditional methods mostly build models based on a single modality (such as audio signals) or a single task (such as pronunciation scoring), such as using acoustic feature extraction combined with a hidden Markov model (HMM) to evaluate pronunciation accuracy. However, such methods have significant limitations: on the one hand, pronunciation training involves the synergy of acoustics, vision (such as lip shape, tongue position) and semantics, and the existing technology lacks deep fusion of multimodal data; on the other hand, there are various types of pronunciation errors (such as phoneme confusion and intonation deviation), and a single task model is difficult to achieve accurate error location and personalized feedback. In addition, traditional systems mostly use fixed training paths and cannot dynamically adjust training strategies according to the user's learning progress, resulting in low training efficiency. Therefore, there is an urgent need for an English auxiliary pronunciation training method and system based on multi-task learning. Summary of the invention
[0003] The object of the present invention is to provide an English auxiliary pronunciation training method and system based on multi-task learning to solve the problems raised in the above background technology.
[0004] In order to solve the above technical problems, the present invention provides the following technical solutions: an English auxiliary pronunciation training method based on multi-task learning, comprising the following steps:
[0005] Constructing a multi-task learning framework, the framework comprising at least two pronunciation-related sub-task models, including pronunciation accuracy scoring, pronunciation error type location, and intonation fluency analysis;
[0006] Collect the user's voice signal, lip movement video and text content through a multimodal data input module;
[0007] Use a shared feature extraction layer to jointly encode multimodal data to generate cross-modal features that integrate acoustics, vision, and semantics;
[0008] Dynamically allocate the weights of sub-task models through the inter-task attention mechanism to optimize feature sharing and task collaboration;
[0009] Generate personalized training strategies based on the comprehensive results of subtask outputs and adaptively adjust the user's pronunciation training path.
[0010] First, by constructing a multi-task learning framework, the user's pronunciation can be comprehensively evaluated and analyzed, and the collection of multimodal data can obtain the user's pronunciation information from multiple angles. The shared feature extraction layer can effectively fuse these multimodal data and generate cross-modal features. The use of the inter-task attention mechanism can reasonably allocate the weights of the sub-task models and improve the training effect.
[0011] Preferably, the shared feature extraction layer adopts an improved Transformer architecture, including:
[0012] 1D convolutional neural network (CNN) and bidirectional LSTM module for speech signal branch;
[0013] 3D CNN and spatiotemporal attention module for video data branch;
[0014] BERT semantic encoder for text branch;
[0015] The cross-modal fusion module adaptively weights the output of each branch through a gating mechanism.
[0016] Preferably, the pronunciation error type localization subtask adopts an adversarial training strategy, simulates typical pronunciation error patterns through a generative adversarial network (GAN), and compares and analyzes them with the user's actual pronunciation, and aligns the temporal relationship between the tongue position pressure signal and the acoustic features through a cross-modal attention mechanism. The generative adversarial network includes a discriminator module based on the phoneme confusion matrix.
[0017] By simulating typical pronunciation error patterns through generative adversarial networks (GANs), a more comprehensive reference can be provided for comparative analysis with users' actual pronunciation. The use of cross-modal attention mechanisms can effectively align the temporal relationship between tongue pressure signals and acoustic features, further improving the accuracy of locating pronunciation error types. The discriminator module based on the phoneme confusion matrix contained in the generative adversarial network can more accurately determine the type of pronunciation errors, provide users with more targeted pronunciation training suggestions, and help users improve pronunciation problems more effectively.
[0018] Preferably, the adaptively adjusting the training path specifically includes:
[0019] Real-time monitoring of the heat map distribution of user pronunciation errors;
[0020] Construct a reward function based on the reinforcement learning algorithm, and prioritize the training intensity of high-frequency incorrect phonemes;
[0021] Generate voice training materials with adaptive difficulty based on the user's historical performance.
[0022] Preferably, the method also includes collecting the user's tongue position pressure signal and laryngeal vibration data through a wearable device, and jointly modeling the physiological signal features and speech features.
[0023] Preferably, the multi-task learning framework adopts a curriculum learning strategy, and gradually increases the complexity of subtasks during the training process: in the initial stage, the pronunciation accuracy score is optimized first, and in the later stage, the intonation fluency and emotional expression are jointly optimized.
[0024] Preferably, the generative adversarial network is connected to a pronunciation error knowledge graph database, associates and infers the user's pronunciation error pattern with the phoneme opposition and co-pronunciation rules in linguistic theory, and generates explainable error correction suggestions.
[0025] By collecting the user's tongue pressure signal and laryngeal vibration data through wearable devices and jointly modeling the physiological signal characteristics with the speech characteristics, we can have a more comprehensive understanding of the user's pronunciation and obtain more information from a physiological level, thereby providing support for more accurate pronunciation training. The joint modeling of physiological signal characteristics and speech characteristics can more deeply analyze problems in the pronunciation process and provide users with more targeted training suggestions.
[0026] An English auxiliary pronunciation training system based on multi-task learning, comprising:
[0027] Multimodal input devices (microphone arrays, cameras, wearable sensors);
[0028] Edge computing module, deploying lightweight multi-task learning models to achieve real-time reasoning;
[0029] Augmented reality (AR) feedback interface, dynamically displaying standard pronunciation movements through a 3D oral model of a virtual teacher;
[0030] Distributed training platform that supports federated learning of multi-user data to update model parameters.
[0031] Preferably, the AR feedback interface integrates a tactile feedback device. When the user pronounces something incorrectly, the micro-vibration of the wearable gloves prompts the correct movement trajectory of the corresponding vocal organ. The micro-vibration prompt includes a pulse sequence encoded based on the movement trajectory of the vocal organ, and the vibration frequency is proportional to the pronunciation duration of the target phoneme.
[0032] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0033] First, the present invention, by first constructing a multi-task learning framework, can comprehensively evaluate and analyze the user's pronunciation, and the collection of multimodal data can obtain the user's pronunciation information from multiple angles, and the shared feature extraction layer can effectively fuse these multimodal data to generate cross-modal features. The use of the inter-task attention mechanism can reasonably allocate the weights of the sub-task models and improve the training effect. Finally, a personalized training strategy is generated according to the comprehensive results of the sub-tasks, which can adaptively adjust the user's pronunciation training path to meet the needs of different users and enable them to more effectively improve their English pronunciation level. Through strategies such as multimodal data input, shared feature extraction, inter-task attention mechanism, and adaptive training path adjustment, accurate evaluation and personalized training of English pronunciation are achieved.
[0034] Second, in the present invention, the multimodal input device can comprehensively collect the pronunciation-related information of the user; the lightweight multi-task learning model of the edge computing module realizes real-time reasoning, which improves the response speed of the system; the augmented reality (AR) feedback interface dynamically displays the standard pronunciation movements through the 3D oral model of the virtual teacher, providing users with intuitive learning guidance; the distributed training platform supports the federated learning of multi-user data to update the model parameters, which can continuously optimize the model and improve the training effect, providing users with an efficient, convenient and personalized English pronunciation training environment, and providing intuitive and efficient training experience and data update mechanism through the augmented reality feedback interface and distributed training platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flowchart of the English auxiliary pronunciation training method based on multi-task learning of the present invention;
[0036] Figure 2 A flowchart of a multi-task learning framework and multi-modal data acquisition is constructed for the present invention;
[0037] Figure 3 A flowchart of generating a personalized training strategy for the comprehensive results output by the subtasks of the present invention;
[0038] Figure 4 It is a block diagram of the English auxiliary pronunciation training system based on multi-task learning of the present invention. DETAILED DESCRIPTION
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] The present invention provides the following technical solutions:
[0041] See also Figure 1 , Figure 2 , Figure 3 , an English auxiliary pronunciation training method based on multi-task learning, comprising the following steps:
[0042] Build a multi-task learning framework that includes at least two pronunciation-related subtask models, including pronunciation accuracy scoring, pronunciation error type location, and intonation fluency analysis;
[0043] Collect the user's voice signal, lip movement video and text content through a multimodal data input module;
[0044] Use a shared feature extraction layer to jointly encode multimodal data to generate cross-modal features that integrate acoustics, vision, and semantics;
[0045] Dynamically allocate the weights of sub-task models through the inter-task attention mechanism to optimize feature sharing and task collaboration;
[0046] Generate personalized training strategies based on the comprehensive results of subtask outputs and adaptively adjust the user's pronunciation training path.
[0047] Through the above technical solution, firstly, by constructing a multi-task learning framework, the user's pronunciation can be comprehensively evaluated and analyzed, and the collection of multimodal data can obtain the user's pronunciation information from multiple angles. The shared feature extraction layer can effectively fuse these multimodal data and generate cross-modal features. The use of the inter-task attention mechanism can reasonably allocate the weights of the sub-task models and improve the training effect. Finally, a personalized training strategy is generated according to the comprehensive results of the sub-tasks, which can adaptively adjust the user's pronunciation training path to meet the needs of different users and enable them to more effectively improve their English pronunciation level.
[0048] The shared feature extraction layer uses an improved Transformer architecture, including:
[0049] 1D convolutional neural network (CNN) and bidirectional LSTM module for speech signal branch;
[0050] 3D CNN and spatiotemporal attention module for video data branch;
[0051] BERT semantic encoder for text branch;
[0052] The cross-modal fusion module adaptively weights the output of each branch through a gating mechanism.
[0053] Through the above technical solutions, the 1D convolutional neural network (CNN) and bidirectional LSTM module of the speech signal branch can effectively process the characteristics of the speech signal; the 3D CNN and spatiotemporal attention module of the video data branch can well capture the spatiotemporal information in the video; the BERT semantic encoder of the text branch can deeply understand and encode the text content, and the cross-modal fusion module adaptively weights the output of each branch through the gating mechanism, realizing the effective fusion of different modal information, providing richer and more accurate feature representation for subsequent pronunciation training.
[0054] The pronunciation error type localization subtask adopts an adversarial training strategy. It simulates typical pronunciation error patterns through a generative adversarial network (GAN) and compares and analyzes them with the user's actual pronunciation. It aligns the temporal relationship between the tongue position pressure signal and the acoustic features through a cross-modal attention mechanism. The generative adversarial network includes a discriminator module based on the phoneme confusion matrix.
[0055] Through the above technical solution, typical pronunciation error patterns are simulated through the generative adversarial network (GAN), which can provide a more comprehensive reference for comparative analysis with the user's actual pronunciation. The use of the cross-modal attention mechanism can effectively align the temporal relationship between the tongue pressure signal and the acoustic features, further improving the accuracy of locating the pronunciation error type. The discriminator module based on the phoneme confusion matrix contained in the generative adversarial network can more accurately determine the type of pronunciation error, provide users with more targeted pronunciation training suggestions, and help users improve pronunciation problems more effectively.
[0056] Adaptive adjustment of the training path specifically includes:
[0057] Real-time monitoring of the heat map distribution of user pronunciation errors;
[0058] Construct a reward function based on the reinforcement learning algorithm, and prioritize the training intensity of high-frequency incorrect phonemes;
[0059] Generate voice training materials with adaptive difficulty based on the user's historical performance.
[0060] Through the above technical solution, by real-time monitoring of the heat map distribution of users' pronunciation errors, we can intuitively understand the distribution of users' pronunciation problems, build a reward function based on the reinforcement learning algorithm, and prioritize the training intensity of high-frequency erroneous phonemes, so as to improve users' pronunciation accuracy in a more targeted manner. According to the user's historical performance, voice training materials with adaptive difficulty can be generated to meet the personalized needs of different users, making training more efficient and better helping users improve their English pronunciation.
[0061] It also includes collecting the user's tongue position pressure signal and laryngeal vibration data through wearable devices, and jointly modeling the physiological signal characteristics and speech characteristics.
[0062] Through the above technical solution, the user's tongue position pressure signal and throat vibration data are collected through wearable devices, and the physiological signal characteristics and speech characteristics are jointly modeled, so as to have a more comprehensive understanding of the user's pronunciation and obtain more information from the physiological level, thereby providing support for more accurate pronunciation training. The joint modeling of physiological signal characteristics and speech characteristics can more deeply analyze the problems in the pronunciation process, provide users with more targeted training suggestions, help users better improve their pronunciation, and improve the accuracy and fluency of English pronunciation.
[0063] The multi-task learning framework adopts a curriculum learning strategy, gradually increasing the complexity of subtasks during training: in the initial stage, pronunciation accuracy scores are optimized first, and in the later stage, intonation fluency and emotional expression are jointly optimized.
[0064] Through the above technical solution, the complexity of subtasks is gradually increased during the training process, which is in line with the principle of gradual learning. In the initial stage, the pronunciation accuracy score is optimized first, which helps learners lay a solid foundation. In the later stage, the fluency of intonation and emotional expression are jointly optimized, which can enable learners to improve their pronunciation at a higher level, making their English pronunciation more natural, fluent and emotional, and better meeting the needs of learners at different stages.
[0065] The generative adversarial network connects the pronunciation error knowledge graph database, associates and infers the user's pronunciation error patterns with the phoneme opposition and co-pronunciation rules in linguistic theory, and generates explainable error correction suggestions.
[0066] Through the above technical solution, the pronunciation error knowledge graph database is connected through the generative adversarial network, and the user's pronunciation error pattern is associated with the phoneme opposition and co-pronunciation rules in linguistic theory to generate explainable error correction suggestions. The generation ability and data processing capabilities of GAN are utilized. The generator of GAN can try to generate pronunciation data similar to the real pronunciation pattern, while the discriminator can determine whether the input pronunciation data is accurate. By associating the pronunciation error pattern with the linguistic theory rules, it can provide users with more targeted and explainable error correction suggestions.
[0067] See also Figure 4 , an English auxiliary pronunciation training system based on multi-task learning, comprising:
[0068] Multimodal input devices (microphone arrays, cameras, wearable sensors);
[0069] Edge computing module, deploying lightweight multi-task learning models to achieve real-time reasoning;
[0070] Augmented reality (AR) feedback interface, dynamically displaying standard pronunciation movements through a 3D oral model of a virtual teacher;
[0071] Distributed training platform that supports federated learning of multi-user data to update model parameters.
[0072] Through the above technical solutions, the multimodal input device can comprehensively collect the user's pronunciation-related information; the lightweight multi-task learning model of the edge computing module realizes real-time reasoning, which improves the response speed of the system; the augmented reality (AR) feedback interface dynamically displays the standard pronunciation movements through the 3D oral model of the virtual teacher, providing users with intuitive learning guidance; the distributed training platform supports federated learning of multi-user data to update model parameters, which can continuously optimize the model and improve the training effect, providing users with an efficient, convenient and personalized English pronunciation training environment.
[0073] The AR feedback interface integrates a tactile feedback device. When the user pronounces something incorrectly, the micro-vibration of the wearable gloves will prompt the correct movement trajectory of the corresponding vocal organs. The micro-vibration prompt contains a pulse sequence encoded based on the movement trajectory of the vocal organs, and the vibration frequency is proportional to the pronunciation duration of the target phoneme.
[0074] Through the above technical solution, when the user pronounces incorrectly, the micro-vibration of the wearable gloves can prompt the correct movement trajectory of the corresponding vocal organs. This method can provide users with more intuitive and specific feedback, helping users to better understand and correct pronunciation errors. The micro-vibration prompt contains a pulse sequence encoded based on the movement trajectory of the vocal organs, making the prompt more accurate and effective. The vibration frequency is proportional to the pronunciation duration of the target phoneme, which can better guide users to master the correct pronunciation rhythm and duration.
[0075] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and alterations may be made to the embodiments without departing from the principles and spirit thereof, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An English auxiliary pronunciation training method based on multi-task learning, characterized in that: The following steps are involved: Constructing a multi-task learning framework, the framework comprising at least two pronunciation-related sub-task models, including pronunciation accuracy scoring, pronunciation error type location, and intonation fluency analysis; Collect the user's voice signal, lip movement video and text content through a multimodal data input module; Use a shared feature extraction layer to jointly encode multimodal data to generate cross-modal features that integrate acoustics, vision, and semantics; Dynamically allocate the weights of sub-task models through the inter-task attention mechanism to optimize feature sharing and task collaboration; Generate personalized training strategies based on the comprehensive results of subtask outputs and adaptively adjust the user's pronunciation training path.
2. The English auxiliary pronunciation training method based on multi-task learning according to claim 1, characterized in that: The shared feature extraction layer adopts an improved Transformer architecture, including: 1D convolutional neural network (CNN) and bidirectional LSTM module for speech signal branch; 3D CNN and spatiotemporal attention module for video data branch; BERT semantic encoder for text branch; The cross-modal fusion module adaptively weights the output of each branch through a gating mechanism.
3. The English auxiliary pronunciation training method based on multi-task learning according to claim 1, characterized in that: The pronunciation error type localization subtask adopts an adversarial training strategy, simulates typical pronunciation error patterns through a generative adversarial network (GAN), and compares and analyzes them with the user's actual pronunciation. The temporal relationship between the tongue position pressure signal and the acoustic features is aligned through a cross-modal attention mechanism. The generative adversarial network includes a discriminator module based on a phoneme confusion matrix.
4. The English auxiliary pronunciation training method based on multi-task learning according to claim 1, characterized in that: The adaptive adjustment of the training path specifically includes: Real-time monitoring of the heat map distribution of user pronunciation errors; Construct a reward function based on the reinforcement learning algorithm, and prioritize the training intensity of high-frequency incorrect phonemes; Generate voice training materials with adaptive difficulty based on the user's historical performance.
5. The English auxiliary pronunciation training method based on multi-task learning according to claim 1 is characterized in that: It also includes collecting the user's tongue position pressure signal and laryngeal vibration data through wearable devices, and jointly modeling the physiological signal characteristics and speech characteristics.
6. The English auxiliary pronunciation training method based on multi-task learning according to claim 1, characterized in that: The multi-task learning framework adopts a curriculum learning strategy, and gradually increases the complexity of subtasks during training: in the early stage, pronunciation accuracy scores are optimized first, and in the later stage, intonation fluency and emotional expression are jointly optimized.
7. The English auxiliary pronunciation training method based on multi-task learning according to claim 1, characterized in that: The generative adversarial network is connected to a pronunciation error knowledge graph database, associates and infers the user's pronunciation error pattern with the phoneme opposition and co-pronunciation rules in linguistic theory, and generates explainable error correction suggestions.
8. An English auxiliary pronunciation training system based on multi-task learning, characterized in that: include: Multimodal input devices (microphone arrays, cameras, wearable sensors); Edge computing module, deploying lightweight multi-task learning models to achieve real-time reasoning; Augmented reality (AR) feedback interface, dynamically displaying standard pronunciation movements through a 3D oral model of a virtual teacher; Distributed training platform that supports federated learning of multi-user data to update model parameters.
9. The English pronunciation assistance training system based on multi-task learning according to claim 8, characterized in that: The AR feedback interface integrates a tactile feedback device. When the user pronounces something incorrectly, the micro-vibration of the wearable gloves prompts the correct movement trajectory of the corresponding vocal organ. The micro-vibration prompt includes a pulse sequence encoded based on the movement trajectory of the vocal organ, and the vibration frequency is proportional to the pronunciation duration of the target phoneme.
Citation Information
Cited By
Spoken English pronunciation correction auxiliary system based on speech recognition
CN120412648A
English spoken pronunciation correction auxiliary system based on speech recognition
CN120412648B
Foreign language teaching effect evaluation method and system based on artificial intelligence
CN120636446A
Artificial intelligence-based foreign language teaching effect evaluation method and system
CN120636446B
Mandarin pronunciation real-time correction method and system based on multi-modal streaming learning
CN121393456A