Personalized English listening and speaking interactive training system based on large language model
By adopting a large language model and TTS engine in the English interactive training system, integrating multiple modules, providing personalized feedback and diversified dialogue practice scenarios, the problem of poor functional richness of the existing system is solved, and learning interest and teaching quality are improved.
Patent Information
- Application Number
- CN202510451322.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing English interactive training system has poor functional richness in evaluation and interactive teaching, and lacks personalized and diverse learning experiences.
It adopts a personalized English listening and speaking interactive training system based on large language models, integrating training selection modules, voice acquisition modules, voice judgment modules, voice recognition modules, voice error correction modules, training analysis modules, interactive teaching modules, promotion scoring modules and voice interaction modules. Through the cloud platform and API, LLM and TTS engines are integrated, providing personalized feedback and diversified dialogue practice scenarios.
It improves learners' interest and enthusiasm in learning, enhances practical ability and independent learning habits, provides personalized learning reports and feedback, and improves teaching quality and learning efficiency.
Smart Images

Figure CN119993208A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of English learning, and more specifically to a personalized English listening and speaking interactive training system based on a large language model. Background Art
[0002] Given the current society's growing demand for high-quality foreign language education resources, the increasing maturity of relevant technical support, and the booming online education industry, it is very necessary to help users improve their English listening and speaking skills through technical means; The existing English interactive training system "As shown in Application No.: CN202210458460.3, an interactive system for English listening and speaking learning", has the advantages of interacting with learners, evaluating the pronunciation quality of learners, pointing out learners' mistakes, and reminding learners in time when they are not in good condition; However, in the application process, the assessment of learners' learning is relatively limited and there are fewer forms of interactive teaching, resulting in poor functional richness. Summary of the invention
[0003] The purpose of the present invention is to solve the above technical problems and provide a personalized English listening and speaking interactive training system based on a large language model.
[0004] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions: The present invention proposes a personalized English listening and speaking interactive training system based on a large language model, including a training selection module, a database, a voice collection module, a voice judgment module, a cloud platform, a voice recognition module, a voice error correction module, a training analysis module and an interactive teaching module; The training selection module is used for learners to select the type of speech to be read from the database for pronunciation training as needed; The voice collection module is used to collect the learner's voice data and transmit the collected voice data to the cloud platform; The voice judgment module is connected to the cloud platform and is used to judge the validity of the voice data. If it is valid, the voice data will be sent to the voice recognition module; if it is invalid, the voice data will be collected again; The speech recognition module is used to transmit the erroneous speech to the speech correction module so that the learner can perform correction exercises; The speech error correction module is used to store and record learners' erroneous speech and draw charts according to preset rules to provide reference for learners to correct their exercises; The training analysis module is connected to the speech judgment module, and is used to obtain the judgment result of the speech judgment module to analyze the training status of the learner and judge whether the corresponding learner needs to rest; wherein the judgment result is a valid mark and an invalid mark; The training analysis module is used to send reminder information to the terminal of the corresponding learner to remind the corresponding learner to take a break for a period of time before continuing training to improve learning efficiency; The interactive teaching module is used for teachers and learners to log in to the education platform and conduct online interactive teaching, so that teachers and students can form effective interaction, improve teaching efficiency and students' learning enthusiasm; The cloud platform integrates LLM and TTS engines through APIs. The large language model can be Alibaba Cloud Tongyi Qianwen and Baidu Wenxin Yiyan, and the language understanding module is Alibaba Cloud's Tongyi series. It also includes a promotion scoring module, which is connected to the cloud platform; The specific application of the promotion scoring module is as follows: Setting a language understanding module, which is connected to the promotion scoring module. The language understanding module can be the Tongyi series of Alibaba Cloud. The language understanding module is used to set the standards for the appropriateness of vocabulary use and grammatical correctness; An algorithm for weighting the scores of pronunciation accuracy, fluency, appropriateness of vocabulary use, and grammatical correctness is set, and the algorithm may be a weighted average; and the database is set with multiple levels and promotion is possible when the required score values are reached for multiple consecutive times; When checking pronunciation accuracy: the pronunciation sample is recorded through the voice collection module, the voice judgment module is used to ensure that the audio quality meets the requirements, the voice recognition module converts the audio into text and compares it with the standard answer to evaluate the accuracy, and the voice error correction module provides specific error analysis and improvement suggestions; When checking fluency: the voice collection module collects the user's spoken expressions, the voice judgment module cleans and optimizes the audio, and the voice recognition module recognizes the content, analyzes the speech speed and pauses to measure fluency; When checking the appropriateness of vocabulary and grammatical correctness: the text information obtained after speech recognition is sent to the language understanding module, and the LLM gives scores on vocabulary use and grammar; Comprehensive scoring and personalized feedback: The training analysis module continuously monitors and records the user's score data, pronunciation accuracy, fluency, vocabulary appropriateness and grammatical correctness indicators, and then generates an evaluation report for the indicators through LLM; As a preferred technical solution of the present invention, it also includes a voice interaction module, which is connected to the TTS engine, generates natural and fluent English speech through the TTS engine and the large language model, and supports simulated dialogue exercises in various scenarios. The cloud platform integrates an open source framework through an API; The specific applications of the voice interaction module are as follows: The TTS engine generates natural and fluent English speech; The TTS engine can use Google TTS, Amazon Polly, or Microsoft Azure TTS; And optimize voice quality through TTS engine; Collected through scenario scripting or LLM and stored in the database; Define multiple virtual characters and store their speaking styles and background stories in the database; build with an open source framework to understand and process user input; use the speech recognition module in conjunction with the speech judgment module to accurately capture user voice commands and provide instant feedback through TTS technology.
[0005] The beneficial effects of the present invention are as follows: By integrating a large language model and TTS engine on the cloud platform, adding a promotion scoring module and a voice interaction module, learners can have a clear understanding of their current ability level. It also provides teachers with an effective tool to evaluate students' progress to improve teaching quality, set different levels, and provide an incentive mechanism. Combined with the powerful processing capabilities of LLM, the system can provide each user with personalized learning reports, personalized feedback and suggestions, provide a variety of dialogue practice scenarios and a vivid, natural and smooth voice experience, which can improve learning interest, increase participation enthusiasm, enhance practical ability and promote learners' habits of independent learning, so as to improve functional richness. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 It is a system block diagram of the personalized English listening and speaking interactive training system based on a large language model of the present invention. DETAILED DESCRIPTION
[0007] like Figure 1 As shown, this embodiment proposes: a personalized English listening and speaking interactive training system based on a large language model, including a training selection module, a database, a voice acquisition module, a voice judgment module, a cloud platform, a voice recognition module, a voice error correction module, a training analysis module and an interactive teaching module, and its application is based on the use of "CN202210458460.3" in the background technology; The training selection module is used for learners to select the type of speech to be read from the database for pronunciation training as needed; The voice collection module is used to collect the learner's voice data and transmit the collected voice data to the cloud platform; The speech judgment module is connected to the cloud platform and is used to judge the validity of the speech data. If it is valid, the speech data is sent to the speech recognition module; if it is invalid, the speech data is collected again; The speech recognition module is used to transmit erroneous speech to the speech correction module so that the learner can perform correction exercises; The speech error correction module is used to store and record learners' erroneous speech and draw charts according to preset rules to provide reference for learners to correct their exercises; The training analysis module is connected to the speech judgment module, and is used to obtain the judgment result of the speech judgment module to analyze the training status of the learner and judge whether the corresponding learner needs to rest; wherein the judgment result is a valid mark and an invalid mark; The training analysis module is used to send reminder information to the terminal of the corresponding learner to remind the corresponding learner to take a break for a period of time before continuing training to improve learning efficiency; The interactive teaching module is used for teachers and learners to log in to the education platform and conduct online interactive teaching, so that effective interaction can be formed between teachers and students, improving teaching efficiency and students' learning enthusiasm.
[0008] The cloud platform integrates a large language model (LLM) and a TTS engine through APIs. The large language model can be Alibaba Cloud Tongyi Qianwen, Baidu Wenxin Yiyan, etc. API (Application Programming Interface) is a software intermediary that allows different software to interact with each other. In the cloud platform environment, large language models and TTS (text-to-speech) are integrated through API to call specific service functions; It also includes a promotion scoring module, which is connected to the cloud platform and is used to judge and score the pronunciation accuracy information, fluency information, vocabulary usage appropriateness information and grammatical correctness information collected by the cloud platform; The specific application of the promotion scoring module is as follows: Setting a language understanding module, which is connected to the promotion scoring module. The language understanding module can be the Tongyi series of Alibaba Cloud. The language understanding module is used to set the standards for the appropriateness of vocabulary use and grammatical correctness; An algorithm for weighting the scores of pronunciation accuracy, fluency, appropriateness of vocabulary use, and grammatical correctness is set, which may be a weighted average, etc. Multiple levels are set in advance in the database, and students can advance to the next level when the required score values are reached multiple times in a row; The weighted average calculation formula is: Given a set of values (x1, x2, ..., xn) and their corresponding weights (w1, w2, ..., wn) (assuming all weights are positive), the weighted average (A_w) can be calculated using the following formula: [Aw = \frac{\sum{i=1}{n} (xi \times wi)}{\sum_{i=1}{n} w_i}\] The numerator is the sum of all values multiplied by their corresponding weights; The denominator is the sum of all weights.
[0009] Example: When the scores for pronunciation accuracy, fluency, vocabulary appropriateness, and grammatical correctness are 80, 90, 70, and 80 points respectively, and the weights of these courses are 0.4, 0.2, 0.2, and 0.2 respectively, the above formula is used to calculate as follows: Values: (x1=80, x2=90, x3=70, x4=80); Weights (credits): (w1=0.4, w2=0.2, w3=0.2, w3=0.2); Substituting into the formula we get: [A_w = \frac{(80 \times 0.4)+(90 \times 0.2)+(70 \times 0.2)+(80 \times 0.2)}{0.4+0.2+0.2+0.2} = \frac{32+18+14+16}{1} = \frac{80}{1} \approx80\] It can be obtained that the learner's score is 80 points.
[0010] When checking pronunciation accuracy: users input their own pronunciation samples through the voice collection module, the voice judgment module performs preprocessing to ensure that the audio quality meets the requirements, the voice recognition module converts the audio into text and compares it with the standard answer to evaluate the accuracy, and the voice error correction module can also provide specific error analysis and improvement suggestions; When checking fluency: the voice collection module collects the user's spoken expressions, the voice judgment module cleans and optimizes the audio, and the voice recognition module not only recognizes the content, but also analyzes factors such as speech speed and pauses to measure fluency; When checking the appropriateness of vocabulary and grammatical correctness: the text information obtained after speech recognition is sent to the language understanding module, which can understand and analyze sentence structure, vocabulary selection, etc., and then give scores on vocabulary use and grammar; Comprehensive scoring and personalized feedback: The training analysis module continuously monitors and records the user's learning performance. Based on the above score data and the above indicators (pronunciation, fluency, vocabulary, grammar), LLM generates a comprehensive evaluation report for the indicator. The report not only contains the score, but also detailed improvement directions and personalized suggestions. After the application of the above-mentioned promotion scoring module: it can help learners clearly understand their current ability level, and also provide teachers with an effective tool to evaluate students' progress and improve teaching quality; setting different levels provides an incentive mechanism; combined with the powerful processing capabilities of LLM, the system can provide each user with personalized learning reports, personalized feedback and suggestions.
[0011] It also includes a voice interaction module, which is connected to the TTS engine, generates natural and fluent English speech through the TTS engine and a large language model, and supports simulated conversation exercises in various scenarios. The cloud platform integrates open source frameworks through APIs. The specific applications of the voice interaction module are as follows: 1. The TTS engine generates natural and fluent English voice; The TTS engine can use Google TTS, Amazon Polly, or Microsoft Azure TTS; The voice quality is optimized through the TTS engine to adjust parameters to optimize the naturalness and fluency of the voice, such as pitch, speaking speed, emotion, etc.
[0012] 2. Design dialogue scenarios to provide a variety of dialogue practice scenarios, such as shopping, traveling, interviews, etc.; these scenarios can be written directly through scenario scripts or collected through LLM, including various life-like or professional dialogue scripts involving different roles and dialogue contents; these scripts are stored in the database for user selection and loading.
[0013] 3. Real-time dialogue: users can have real-time conversations with virtual characters to improve their oral skills; By defining multiple virtual characters, each with its own unique way of speaking and background story, which are stored in the database; using open source frameworks (such as Rasa) to build, responsible for understanding and processing user input to generate appropriate responses based on the current conversation context; the speech recognition module cooperates with the speech judgment module to accurately capture the user's voice commands and provide instant feedback through TTS technology, ensuring that the entire communication process is smooth and unimpeded; Rasa is an open source machine learning framework specifically designed for building context-aware conversational systems, allowing the creation of applications that can understand natural language and carry out conversations. After the above-mentioned voice interaction module is applied: through diversified dialogue practice scenarios and vivid, natural and smooth sound experience, it can improve learning interest and increase participation enthusiasm. At the same time, the real-time dialogue function allows learners to communicate with different virtual characters in a virtual environment and enhance practical ability; it provides a platform for English listening and speaking practice anytime and anywhere to promote the habit of independent learning.
[0014] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention. Therefore, no matter from which point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the attached claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
[0015] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A personalized English listening and speaking interactive training system based on a large language model, including a training selection module, a database, a voice acquisition module, a voice judgment module, a cloud platform, a voice recognition module, a voice error correction module, a training analysis module, and an interactive teaching module It is characterized in that The cloud platform integrates LLM and TTS engines through APIs respectively; It also includes a promotion scoring module, which is connected to the cloud platform; The specific application of the promotion scoring module is as follows: Setting up a language understanding module, which is connected to the promotion scoring module, and the language understanding module is used to set the standards for appropriateness of vocabulary use and grammatical correctness; An algorithm for weighting the scores of pronunciation accuracy, fluency, appropriateness of vocabulary use, and grammatical correctness is set, and the algorithm may be a weighted average; and the database is set with multiple levels and promotion is possible when the required score values are reached for multiple consecutive times; When checking pronunciation accuracy: the pronunciation sample is recorded through the voice collection module, the voice judgment module is used to ensure that the audio quality meets the requirements, the voice recognition module converts the audio into text and compares it with the standard answer to evaluate the accuracy, and the voice error correction module provides specific error analysis and improvement suggestions; When checking fluency: the voice collection module collects the user's spoken expressions, the voice judgment module cleans and optimizes the audio, and the voice recognition module recognizes the content, analyzes the speech speed and pauses to measure fluency; When checking the appropriateness of vocabulary and grammatical correctness: the text information obtained after speech recognition is sent to the language understanding module, and the LLM gives scores on vocabulary use and grammar; Comprehensive scoring and personalized feedback: The training analysis module continuously monitors and records the user's score data, pronunciation accuracy, fluency, vocabulary appropriateness and grammatical correctness indicators, and then generates an evaluation report for the indicators through LLM; Also includes a voice interaction module.
2. The personalized English listening and speaking interactive training system based on a large language model according to claim 1 is characterized in that: The voice interaction module is connected to the TTS engine, and generates natural and fluent English speech through the TTS engine and the large language model, and supports simulated dialogue exercises in various scenarios. The open source framework is integrated through the API on the cloud platform.
3. The personalized English listening and speaking interactive training system based on a large language model according to claim 2 is characterized in that: The specific application of the voice interaction module is as follows: The TTS engine generates natural and fluent English speech; The TTS engine can use Google TTS, Amazon Polly, or Microsoft Azure TTS; And optimize voice quality through TTS engine; Collected through scenario scripting or LLM and stored in the database; Define multiple virtual characters and store their speaking styles and background stories in the database; build with an open source framework to understand and process user input; use the speech recognition module in conjunction with the speech judgment module to accurately capture user voice commands and provide instant feedback through TTS technology.
4. The personalized English listening and speaking interactive training system based on a large language model according to claim 3 is characterized in that: The open source framework is Rasa.
5. The personalized English listening and speaking interactive training system based on a large language model according to claim 1 is characterized in that: The large language model is Alibaba Cloud Tongyi Qianwen or Baidu Wenxin Yiyan.
6. The personalized English listening and speaking interactive training system based on a large language model according to claim 1 is characterized in that: The language understanding module is the Tongyi series of Alibaba Cloud.
7. The personalized English listening and speaking interactive training system based on a large language model according to claim 1 is characterized in that: The training selection module is used for learners to select the type of speech to be read from the database for pronunciation training as needed; The voice collection module is used to collect the learner's voice data and transmit the collected voice data to the cloud platform; The voice judgment module is connected to the cloud platform and is used to judge the validity of the voice data. If it is valid, the voice data will be sent to the voice recognition module; if it is invalid, the voice data will be collected again; The speech recognition module is used to transmit the erroneous speech to the speech correction module so that the learner can perform correction exercises; The speech error correction module is used to store and record learners' erroneous speech and draw charts according to preset rules to provide reference for learners to correct their exercises; The training analysis module is connected to the speech judgment module, and is used to obtain the judgment result of the speech judgment module to analyze the training status of the learner and determine whether the corresponding learner needs a rest; The judgment result is a valid mark and an invalid mark; The training analysis module is used to send reminder information to the terminal of the corresponding learner to remind the corresponding learner to take a break for a period of time before continuing training to improve learning efficiency; The interactive teaching module is used for teachers and learners to log in to the education platform and conduct online interactive teaching, so that effective interaction can be formed between teachers and students, improving teaching efficiency and students' learning enthusiasm.
Citation Information
Patent Citations
Interactive language learning system and method
CN101739870A
Method for automatic evaluation based on generalized fluent spoken language fluency
CN101740024A
Interaction system for English listening and speaking learning
CN114743541A
Training device and equipment for shaping language fluency of children and storage medium
CN115662242A
Oral language evaluation method, oral language evaluation model training method and related device
CN119517080A