Intelligent generation method of English listening audio for junior middle school students based on large model
Through an intelligent generation method based on large language model, combined with adaptive learning and TTS tone configuration, the problems of low efficiency and inability to personalize listening material generation in the existing technology are solved, and high-quality and diverse listening audio generation is achieved, which improves the learning effect.
Patent Information
- Application Number
- CN202411565030.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-05
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-05
AI Technical Summary
In the existing junior high school English listening teaching, the generation and use of listening materials rely on manual recording and pre-edited audio content, resulting in low generation efficiency, high cost, and inability to flexibly respond to personalized needs in teaching.
An intelligent generation method based on a large language model is adopted. By collecting and annotating samples of junior high school English listening test papers, the large language model is trained and tuned, and an adaptive learning mechanism is used to adjust the audio content, speech speed and tone in real time according to student feedback, and speech synthesis is performed in combination with the TTS tone configuration and conditional probability model.
The dynamic generation of listening content is realized, and the audio content is adjusted according to students' learning progress and language level is improved, the quality and diversity of listening materials are enhanced, and the learning effect and students' sense of participation are enhanced.
Smart Images

Figure CN119380697B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of English for junior middle schools, and in particular to a method for intelligently generating audio for English listening comprehension for junior middle schools based on a large model. Background Art
[0002] The generation and use of listening materials in the existing junior high school English listening teaching mainly rely on manually recorded and pre-edited audio content. Audio materials are usually recorded in advance by teachers or professional dubbing personnel and organized and used according to a fixed teaching syllabus. However, the traditional way of generating listening materials has obvious limitations. Since the process of recording audio involves topic selection, dubbing, recording and post-editing, it often takes a lot of manpower, time and cost, resulting in the inability of listening materials to flexibly respond to personalized needs in teaching, and it is difficult for students to obtain targeted listening practice resources according to their own learning progress and level.
[0003] Although some existing teaching assistance systems have attempted to use semi-automatic methods to generate listening content, such as generating some audio content through text template filling and speech synthesis, the flexibility and intelligence of such technologies are limited. The system is usually unable to make real-time adjustments and optimizations based on students' learning feedback. The generated listening materials are also difficult to cover diverse voice roles, accents, and situational requirements. In addition, the existing listening audio generation systems have poor adaptability in different teaching scenarios and lack the function of dynamically adjusting audio parameters such as speech speed and intonation, which makes it difficult to ensure the quality and diversity of listening materials.
[0004] In summary, the existing technology has obvious shortcomings in the generation, management and personalized application of listening audio. There is an urgent need for an intelligent solution that combines large language models and speech synthesis technology to improve generation efficiency and audio quality, and to achieve real-time optimization based on student feedback. Summary of the invention
[0005] One purpose of the present invention is to propose a method for intelligently generating English listening audio for junior high school students based on a large model. The present invention can quickly generate required listening materials according to different teaching situations.
[0006] According to an embodiment of the present invention, a method for intelligently generating junior high school English listening audio based on a large model comprises the following steps:
[0007] S1. Collect junior high school English listening test paper samples;
[0008] S2. Clean the text of the junior high school English listening test paper sample, remove irrelevant characters and formats, annotate the Chinese text and the English text respectively, and obtain the annotated junior high school English listening test paper sample;
[0009] S3. Based on the junior high school English listening test paper samples that have been cleaned and annotated, the large language model is trained to enable the large language model to distinguish between Chinese and English texts, and the large language model is tuned;
[0010] S4. Use the trained large language model to automatically classify the text in the junior high school English listening test paper sample and distinguish the Chinese text from the English text;
[0011] S4, select different TTS timbres for Chinese text and English text respectively, the TTS timbres include different pronunciation timbres suitable for Chinese text and English text, and configure different speech speed, pitch and timbre parameters for Chinese text and English text according to actual needs;
[0012] S5, inputting the extracted Chinese text and English text into the corresponding TTS synthesis engine to generate Chinese audio and English audio respectively, and adjusting the speech speed and pitch parameters of the synthesized speech during the generation process to obtain the best auditory effect;
[0013] S6. According to the structure of the junior high school English listening test paper sample, the synthesized Chinese audio and English audio are organized according to the test paper structure to generate a complete listening audio test paper;
[0014] S7. Store the generated listening audio materials in a database, and update the audio materials in real time according to the teaching progress and learner feedback data.
[0015] Optionally, the S1 includes the following steps:
[0016] S11. Determine the sample range of English listening test papers for junior high schools, which includes official test papers provided by the education department, publicly published English listening textbook test papers, and mock test papers independently prepared by schools;
[0017] S12. Screen the collected junior high school English listening test papers and only retain the sample data that conforms to the junior high school English teaching syllabus. The sample data includes the complete test paper structure and content.
[0018] S13. Perform structured storage on the collected junior high school English listening test paper samples to construct a junior high school English listening test paper sample set T:
[0019] T={(D t ,Q s ,C d ,N m )∣D t ∈L,Q s ∈L,C d ∈L,N m ∈L};
[0020] Among them, Dt Indicates the title part of the test paper, which is used to guide students to understand the question type. s Indicates the content of the question, clarifies the question or the task required of students. d Indicates the dialogue part in the test paper, simulating the communication between characters in real situations, N m It indicates the monologue part, which is used to train students' continuous listening comprehension ability. L indicates the language set, including Chinese L c and English L e , that is, L = L c ∪L e , used to identify the language type of each part.
[0021] Optionally, S2 includes the following steps:
[0022] S21. Perform preliminary cleaning on the collected junior high school English listening test paper samples, remove non-text symbols, format codes and invisible characters in the junior high school English listening test paper samples, and retain the Chinese and English titles, questions, dialogues and monologues;
[0023] S22, according to the language set L = L c ∪L e Perform language identification detection on sample data, where L c Indicates Chinese text, L e Represents English text, ensuring that the language type of each part of the content is correctly identified during subsequent processing;
[0024] S23. Divide the preliminarily cleaned junior high school English listening test paper sample into headings D according to the content blocks t 、Question Q s 、Dialogue C d and monologue content N m , and use metadata tags to annotate the corresponding language type to construct a language annotation dataset T mark :
[0025] T mark ={(D t ,L i ),(Q s ,L j ),(C d ,L k ),(N m ,L m )∣L i ,L j ,L k ,L m ∈{L c ,L e}};
[0026] Among them, Li ,L j ,L k ,L m Indicates the language type label of the corresponding content block, used to distinguish Chinese text from English text;
[0027] S24, performing structure verification on the divided junior high school English listening test paper samples, wherein the structure verification rules include character length, sentence end identifier and language consistency, and excluding incomplete or incorrectly formatted data of the junior high school English listening test paper samples;
[0028] S25. Fine-grained annotation is performed on the verified junior high school English listening test paper samples, using the tag set {tag1, tag2,…, tag n}Tag the key elements in the junior high school English listening test paper sample, where tag1 represents the name of a person or role, tag2 represents the time or date, tag3 represents the place name or place name, and other tags ttag n Used to identify key words in a specific context;
[0029] S26, storing the marked junior high school English listening test paper samples according to language type and content block to generate the final junior high school English listening test paper sample T final :
[0030] T final ={(D t ,Q s ,C d ,N m )∣D t ∈L i ,Q s ∈L j ,C d ∈L k ,N m ∈L m}.
[0031] Optionally, S3 includes the following steps:
[0032] S31, the junior high school English listening test sample data set T final Input the initial training module of the large language model to pre-train the large language model so that it has basic language understanding capabilities and learns the grammatical differences between Chinese and English texts;
[0033] S32. Define the adaptive optimization loss function Loss(θ), which is adaptively adjusted according to student feedback:
[0034]
[0035] Among them, N is the number of samples, θ is the parameter set of the large language model, and y i represents the true label of the i-th sample, represents the predicted output of the model, α is the feedback adjustment coefficient, and F s Provide feedback data to students;
[0036] S33, fine-tuning the large language model using an adaptive learning algorithm, inputting the student's feedback data into the large language model in real time after each round of training, updating the feedback adjustment coefficient α in the loss function, so that the parameter optimization process of the large language model gradually adapts to the student's learning rhythm and language level;
[0037] S34. During the training of the large language model, the classification accuracy Acc of the large language model is monitored in real time:
[0038]
[0039] in, is an indicator function, which takes the value of 1 when the predicted result is consistent with the true label, otherwise it takes the value of 0;
[0040] S35. Automatically adjust the difficulty and speech speed of the audio content generated by the large language model based on student feedback data and real-time monitoring results, so that the generated listening audio matches the student's actual ability level;
[0041] S36, repeat the large language model training, feedback input, loss function update and fine-tuning process until the classification accuracy Acc of the model reaches the preset threshold T acc , and make the student's feedback index F s Meet the standards for adaptive learning.
[0042] Optionally, S4 includes the following steps:
[0043] S41, load the trained large language model and the junior high school English listening test sample dataset T final Input the large language model, and the large language model is processed by initializing the parameter set θ0 to enter the inference state:
[0044]
[0045] in, is the initial predicted text and language type pairing set, x i For the i-th input text, L i is the Chinese or English language label, θ0 represents the initial large language model parameter set;
[0046] S42, using the conditional probability model with language context embedding for each text x iCalculate the joint probability that the text belongs to Chinese or English:
[0047]
[0048] Among them, P(L i ∣x i ,C(x i )) means in the context C(x i )Next text x i Belongs to language tag L i The probability of and For language tag L i The corresponding weight matrix and bias term, h(x i ) is the text x i The embedding representation of φ(C(x i )) is the text context C(x i ), M is the number of possible language types;
[0049] S43. Using the joint probability result in S42, assign a language label to each text based on maximum likelihood estimation:
[0050]
[0051] Among them, L i For text x i The final language tag of
[0052] S44, classify the language-classified text according to the title, topic, dialogue and monologue to generate a classified data set T class ;
[0053] S45. Check the classified data set T class Whether it complies with the context consistency rule, if inconsistency is detected, it is based on the consistency detection function Ψ(T class ) to reclassify:
[0054]
[0055] Among them, Ψ(T class ) is the consistency detection function, Indicates that under the context C(x), the text x should satisfy the corresponding language tag L, It is an indicator function. If the condition is met, the value is 1, otherwise it is 0. If the result is 0, the large language model is called for reclassification.
[0056] Optionally, S4 includes the following steps:
[0057] S61, select different TTS voice sets for Chinese text and English text, according to the language tag L i ∈{L c ,L e}Choose the appropriate pronunciation tone set S tts :
[0058] S tts ={S c ,S e}Where S c Applicable to Chinese text, S e Applicable to English texts;
[0059] Among them, S c and S e A collection of TTS voices for Chinese and English pronunciations;
[0060] S62, configuring the speech speed, pitch and timbre parameters for the selected TTS timbre to establish an audio parameter set P;
[0061] S63. Language label L based on input text i For each text segment x i Choose the appropriate sound set S i And the audio parameter P i :
[0062]
[0063] Among them, S i Represents text x i The corresponding timbre set, P i Represents text x i The speaking speed, pitch and timbre parameters;
[0064] S64, input the selected TTS timbre and configured parameters into the TTS synthesis engine to generate the corresponding audio signal y i :
[0065] y i =TTS(x i ,S i ,P i );
[0066] Among them, y i For text x i For the corresponding audio output, the TTS synthesis engine generates a natural and fluent speech signal based on the input text, timbre and parameters;
[0067] S65, the generated audio signal y i Quality assessment is performed based on the audio feature consistency function Φ(y i) Determine whether the audio meets the predetermined quality standard:
[0068]
[0069] in, It is an indicator function. When the audio signal meets the standards of speech rate, pitch and timbre parameters, the value is 1, otherwise it is 0. If the detection result does not meet the standards, the TTS engine is called to re-synthesize the audio.
[0070] Optionally, the S6 includes the following steps:
[0071] S71, modularizing the generated Chinese audio and English audio according to the title, question stem, dialogue and monologue parts in the structure of the junior high school English listening test paper, so that the language label of each audio clip is consistent with the corresponding text part;
[0072] S72. According to the teaching objectives of the junior high school English listening test and the learning needs of students, configure the audio playback order to set appropriate interval time and transition effects for each module audio to simulate the real listening test scene;
[0073] S73. When splicing the audio of each module, adjust the audio connection method according to the content logic to ensure that the dialogue part remains smooth and coherent, and the monologue part is clear and complete;
[0074] S74. Generate multiple versions of audio junior high school English listening test papers according to different student needs, including standard speed and slow speed versions. At the same time, generate customized audio junior high school English listening test papers containing only specific dialogues or monologues according to teacher needs;
[0075] S75. Combine the metadata tags in the audio data to generate a playlist and control panel for the listening audio, supporting teachers and students to adjust the play order, repeat play and pause functions as needed;
[0076] S76. Store the complete organized listening audio junior high school English listening test paper in a database and generate a unique identifier for it.
[0077] The beneficial effects of the present invention are:
[0078] (1) The present invention introduces an adaptive learning mechanism into the large language model and performs real-time fine-tuning of the model according to students' feedback data to achieve dynamic generation of listening content. Compared with traditional listening training materials with fixed content, the model of the present invention can dynamically adjust the content, speaking speed and pitch of the listening audio according to the students' learning progress, language level and weak links, ensuring that students always train at a learning rhythm that suits them. The adaptive mechanism avoids the disadvantage that traditional audio materials cannot be flexibly adjusted, making the listening content generated by the model more targeted and personalized, thereby improving learning effects and students' sense of participation.
[0079] (2) The present invention adopts a modular organization method in the audio generation process to modularly classify Chinese audio and English audio according to the title, question stem, dialogue and monologue parts, and splice them into a complete listening test paper in a logical order. By designing an audio splicing algorithm, the dialogue part remains fluent and coherent, and the monologue part is clear and complete, simulating the real test scenario. At the same time, the system supports the generation of multiple versions of audio test papers and allows teachers to customize audio test papers with specific content according to needs, which not only improves the efficiency of audio management, but also enhances the flexibility of teachers in the classroom, enabling them to quickly generate the required listening materials according to different teaching situations.
[0080] (3) The present invention combines text context embedding with TTS timbre configuration by introducing a conditional probability model, thereby achieving highly natural speech synthesis. During the generation process, the model selects the most appropriate timbre and speaking speed according to the context of the text, so that the generated audio has a consistent expression effect in a variety of situations. The system supports intelligent switching between Chinese and English timbre, ensuring the natural transition of bilingual audio content and avoiding the problems of harsh tones and timbre mismatch in traditional speech synthesis. Through innovative context embedding and multimodal configuration algorithms, the listening audio generated by the model has achieved a higher level in speech naturalness, coherence and diversity, effectively improving students' listening comprehension ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0082] Figure 1 A flowchart of a method for intelligently generating audio for English listening comprehension in junior high school based on a large model proposed by the present invention;
[0083] Figure 2 This is a schematic diagram of the training and adaptive learning and tuning process of a large language model in the large-model-based intelligent generation method for junior high school English listening audio proposed in the present invention. DETAILED DESCRIPTION
[0084] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0085] refer to Figure 1-Figure 2 , a method for intelligently generating audio for junior high school English listening based on a large model, comprising the following steps:
[0086] S1. Collect junior high school English listening test paper samples;
[0087] S2. Clean the text of the junior high school English listening test paper sample, remove irrelevant characters and formats, annotate the Chinese text and the English text respectively, and obtain the annotated junior high school English listening test paper sample;
[0088] S3. Based on the junior high school English listening test paper samples that have been cleaned and annotated, the large language model is trained to enable the large language model to distinguish between Chinese and English texts, and the large language model is tuned;
[0089] S4. Use the trained large language model to automatically classify the text in the junior high school English listening test paper sample and distinguish the Chinese text from the English text;
[0090] S4, select different TTS timbres for Chinese text and English text respectively, the TTS timbres include different pronunciation timbres suitable for Chinese text and English text, and configure different speech speed, pitch and timbre parameters for Chinese text and English text according to actual needs;
[0091] S5, inputting the extracted Chinese text and English text into the corresponding TTS synthesis engine to generate Chinese audio and English audio respectively, and adjusting the speech speed and pitch parameters of the synthesized speech during the generation process to obtain the best auditory effect;
[0092] S6. According to the structure of the junior high school English listening test paper sample, the synthesized Chinese audio and English audio are organized according to the test paper structure to generate a complete listening audio test paper;
[0093] S7. Store the generated listening audio materials in a database, and update the audio materials in real time according to the teaching progress and learner feedback data.
[0094] In this implementation, S1 includes the following steps:
[0095] S11. Determine the sample range of English listening test papers for junior high schools, which shall include official test papers provided by the education department, publicly published English listening textbook test papers, and mock test papers independently prepared by the school;
[0096] S12. Screen the collected junior high school English listening test papers and only retain the sample data that conforms to the junior high school English teaching syllabus. The sample data includes the complete test paper structure and content.
[0097] S13. Perform structured storage on the collected junior high school English listening test paper samples to construct a junior high school English listening test paper sample set T:
[0098] T={(D t ,Q s ,C d ,N m )∣D t∈L,Q s ∈L,C d ∈L,N m ∈L};
[0099] Among them, D t Indicates the title part of the test paper, which is used to guide students to understand the question type. s Indicates the content of the question, clarifies the question or the task required of students. d Indicates the dialogue part in the test paper, simulating the communication between characters in real situations, N m It indicates the monologue part, which is used to train students' continuous listening comprehension ability. L indicates the language set, including Chinese L c and English L e , that is, L = L c ∪L e , used to identify the language type of each part.
[0100] In this implementation, S2 includes the following steps:
[0101] S21. Perform preliminary cleaning on the collected junior high school English listening test paper samples, remove non-text symbols, format codes and invisible characters in the junior high school English listening test paper samples, and retain the Chinese and English titles, questions, dialogues and monologues;
[0102] S22, according to the language set L = L c ∪L e Perform language identification detection on sample data, where L c Indicates Chinese text, L e Represents English text, ensuring that the language type of each part of the content is correctly identified during subsequent processing;
[0103] S23. Divide the preliminarily cleaned junior high school English listening test paper sample into headings D according to the content blocks t 、Question Q s 、Dialogue C d and monologue content N m , and use metadata tags to annotate the corresponding language type to construct a language annotation dataset T mark :
[0104] T mark ={(D t ,L i ),(Q s ,L j ),(C d ,L k ),(N m ,L m )∣L i ,L j ,Lk ,L m ∈{L c ,L e}};
[0105] Among them, L i ,L j ,L k ,L m Indicates the language type label of the corresponding content block, used to distinguish Chinese text from English text;
[0106] S24, performing structure verification on the divided junior high school English listening test paper samples, wherein the structure verification rules include character length, sentence end identifier and language consistency, and excluding incomplete or incorrectly formatted data of the junior high school English listening test paper samples;
[0107] S25. Fine-grained annotation is performed on the verified junior high school English listening test paper samples, using the tag set {tag1, tag2,…, tag n}Tag the key elements in the junior high school English listening test paper sample, where tag1 represents the name of a person or role, tag2 represents the time or date, tag3 represents the place name or place name, and other tags ttag n Used to identify key words in a specific context;
[0108] S26, storing the marked junior high school English listening test paper samples according to language type and content block to generate the final junior high school English listening test paper sample T final :
[0109] T final ={(D t ,Q s ,C d ,N m )∣D t ∈L i ,Q s ∈L j ,C d ∈L k ,N m ∈L m}.
[0110] In this implementation, S3 includes the following steps:
[0111] S31, the junior high school English listening test sample data set T final Input the initial training module of the large language model to pre-train the large language model so that it has basic language understanding capabilities and learns the grammatical differences between Chinese and English texts;
[0112] S32. Define the adaptive optimization loss function Loss(θ), which is adaptively adjusted according to student feedback:
[0113]
[0114] Among them, N is the number of samples, θ is the parameter set of the large language model, and y i represents the true label of the i-th sample, represents the predicted output of the model, α is the feedback adjustment coefficient, and F s Provide feedback data to students;
[0115] S33, fine-tuning the large language model using an adaptive learning algorithm, inputting the student's feedback data into the large language model in real time after each round of training, updating the feedback adjustment coefficient α in the loss function, so that the parameter optimization process of the large language model gradually adapts to the student's learning rhythm and language level;
[0116] S34. During the training of the large language model, the classification accuracy Acc of the large language model is monitored in real time:
[0117]
[0118] in, is an indicator function, which takes the value of 1 when the predicted result is consistent with the true label, otherwise it takes the value of 0;
[0119] S35. Automatically adjust the difficulty and speech speed of the audio content generated by the large language model based on student feedback data and real-time monitoring results, so that the generated listening audio matches the student's actual ability level;
[0120] S36, repeat the large language model training, feedback input, loss function update and fine-tuning process until the classification accuracy Acc of the model reaches the preset threshold T acc , and make the student's feedback index F s Meet the standards for adaptive learning.
[0121] In this implementation, S4 includes the following steps:
[0122] S41, load the trained large language model and the junior high school English listening test sample dataset T final Input the large language model, and the large language model is processed by initializing the parameter set θ0 to enter the inference state:
[0123]
[0124] in, is the initial predicted text and language type pairing set, x i For the i-th input text, Li is the Chinese or English language label, θ0 represents the initial large language model parameter set;
[0125] S42, using the conditional probability model with language context embedding for each text x i Calculate the joint probability that the text belongs to Chinese or English:
[0126]
[0127] Among them, P(L i ∣x i ,C(x i )) means in the context C(x i )Next text x i Belongs to language tag L i The probability of and For language tag L i The corresponding weight matrix and bias term, h(x i ) is the text x i The embedding representation of φ(C(x i )) is the text context C(x i ), M is the number of possible language types;
[0128] S43. Using the joint probability result in S42, assign a language label to each text based on maximum likelihood estimation:
[0129]
[0130] Among them, L i For text x i The final language tag of
[0131] S44, classify the language-classified text according to the title, topic, dialogue and monologue to generate a classified data set T class ;
[0132] S45. Check the classified data set T class Whether it complies with the context consistency rule, if inconsistency is detected, it is based on the consistency detection function Ψ(T class ) to reclassify:
[0133]
[0134] Among them, Ψ(T class ) is the consistency detection function, Indicates that under the context C(x), the text x should satisfy the corresponding language tag L, It is an indicator function. If the condition is met, the value is 1, otherwise it is 0. If the result is 0, the large language model is called for reclassification.
[0135] In this implementation, S4 includes the following steps:
[0136] S61, select different TTS voice sets for Chinese text and English text, according to the language tag L i ∈{L c ,L e}Choose the appropriate pronunciation tone set S tts :
[0137] S tts ={S c ,S e}Where S c Applicable to Chinese text, S e Applicable to English texts;
[0138] Among them, S c and S e A collection of TTS voices for Chinese and English pronunciations;
[0139] S62, configuring the speech speed, pitch and timbre parameters for the selected TTS timbre to establish an audio parameter set P;
[0140] S63. Language label L based on input text i For each text segment x i Choose the appropriate sound set S i And the audio parameter P i :
[0141]
[0142] Among them, S i Represents text x i The corresponding timbre set, P i Represents text x i The speaking speed, pitch and timbre parameters;
[0143] S64, input the selected TTS timbre and configured parameters into the TTS synthesis engine to generate the corresponding audio signal y i :
[0144] y i =TTS(x i ,S i ,P i );
[0145] Among them, y i For text x iFor the corresponding audio output, the TTS synthesis engine generates a natural and fluent speech signal based on the input text, timbre and parameters;
[0146] S65, the generated audio signal y i Quality assessment is performed based on the audio feature consistency function Φ(y i ) Determine whether the audio meets the predetermined quality standard:
[0147]
[0148] in, It is an indicator function. When the audio signal meets the standards of speech rate, pitch and timbre parameters, the value is 1, otherwise it is 0. If the detection result does not meet the standards, the TTS engine is called to re-synthesize the audio.
[0149] In this implementation, S6 includes the following steps:
[0150] S71, modularizing the generated Chinese audio and English audio according to the title, question stem, dialogue and monologue parts in the structure of the junior high school English listening test paper, so that the language label of each audio clip is consistent with the corresponding text part;
[0151] S72. According to the teaching objectives of the junior high school English listening test and the learning needs of students, configure the audio playback order to set appropriate interval time and transition effects for each module audio to simulate the real listening test scene;
[0152] S73. When splicing the audio of each module, adjust the audio connection method according to the content logic to ensure that the dialogue part remains smooth and coherent, and the monologue part is clear and complete;
[0153] S74. Generate multiple versions of audio junior high school English listening test papers according to different student needs, including standard speed and slow speed versions. At the same time, generate customized audio junior high school English listening test papers containing only specific dialogues or monologues according to teacher needs;
[0154] S75. Combine the metadata tags in the audio data to generate a playlist and control panel for the listening audio, supporting teachers and students to adjust the play order, repeat play and pause functions as needed;
[0155] S76. Store the complete organized listening audio junior high school English listening test paper in a database and generate a unique identifier for it.
[0156] Embodiment 1:
[0157] During the mid-term exam preparation stage from September to October 2024, the English teaching and research group of a city experimental middle school tried to use the junior high school English listening audio intelligent generation system of the present invention to generate personalized listening test content for students in four classes of the third grade of junior high school, so as to replace the traditional manual recording and fixed textbook audio. The teaching and research group hopes to shorten the preparation time and improve the quality of listening materials through the new system, and adjust the test content in real time according to students' feedback.
[0158] In actual application, the teacher first logged into the system and entered the three topics of "Travel Conversation", "Shopping Conversation" and "Campus Life" as the main content of the listening test. The system automatically called the annotated junior high school English listening sample data set and used the large language model for text generation and semantic optimization to ensure that the content meets the teaching objectives.
[0159] The system generates audio versions with different speaking speeds and accents to meet the different needs of students with standard level and students with learning difficulties. For example, a section about "shopping conversation" is conducted by two characters. The system selects standard British accent and American accent for synthesis, and sets two versions with standard speaking speed and slow speed. In addition, to ensure the naturalness and consistency of the content, the system automatically optimizes the intonation of the transition part of the conversation to make the generated audio smooth and coherent.
[0160] The generated audio is automatically matched to each part of the test paper:
[0161] Title audio: "A shopping conversation in London";
[0162] Audio of the question stem: introduces the scenario and question requirements;
[0163] Dialogue audio: simulates a shopping conversation between two characters, covering topics such as asking for directions and making payments;
[0164] Monologue Audio: Describing customer feedback after a purchase;
[0165] The system splices all module audios according to the order set by the teacher to form a complete listening test paper and provides teachers with online playback and download options. The test paper contains two versions: standard speed and slow speed, and automatically generates audio for question analysis for students to use when reviewing.
[0166] The method of the present invention was compared with the traditional method in multiple key links, and the following are the main data:
[0167]
[0168] Specifically, in the student feedback phase, the system automatically generated supplementary audio with a slightly lower difficulty level based on the students' performance in the first test (the average accuracy rate was 65%). After the test, students could use the slow version of the audio in after-class exercises and repeatedly listen to the analysis audio for incorrect questions. In the second round of tests, the students' average accuracy rate increased to 85%.
[0169] In addition, the system supports automatic generation of feedback reports after the exam. The reports include analysis of students' overall performance, statistics on the accuracy of each part, and individual listening weaknesses. Teachers use these reports to provide students with targeted review suggestions. The 20 students in Class 3 performed poorly in the "Shopping Conversation" section. Based on the system's suggestions, the teacher provided these students with special accent adaptation training and used American-language audio for intensive practice.
[0170] Through this example, it can be clearly seen that the method of the present invention significantly improves the efficiency and quality of generating listening materials and greatly reduces the workload of teachers compared to traditional manual recording and teaching material audio. Through adaptive learning technology, the system can adjust the content in real time according to students' feedback, truly realizing personalized teaching. The data shows that students' listening scores and learning satisfaction have been significantly improved.
[0171] The present invention introduces an adaptive learning mechanism into a large language model and realizes dynamic generation of listening content by fine-tuning the model in real time according to students' feedback data. Compared with traditional listening training materials with fixed content, the model of the present invention can dynamically adjust the content, speaking speed and pitch of the listening audio according to the students' learning progress, language level and weak links to ensure that students always train at a learning rhythm that suits them. The adaptive mechanism avoids the disadvantage that traditional audio materials cannot be flexibly adjusted, making the listening content generated by the model more targeted and personalized, thereby improving learning effects and students' sense of participation.
[0172] The present invention adopts a modular organization method in the audio generation process to modularly classify Chinese audio and English audio according to the title, question stem, dialogue and monologue parts, and splice them into a complete listening test paper in a logical order. By designing an audio splicing algorithm, the dialogue part remains fluent and coherent, and the monologue part is clear and complete, simulating the real test scenario. At the same time, the system supports the generation of multiple versions of audio test papers, and allows teachers to customize audio test papers with specific content according to needs, which not only improves the efficiency of audio management, but also enhances the flexibility of teachers in the classroom, enabling them to quickly generate the required listening materials according to different teaching situations.
[0173] The present invention combines text context embedding with TTS timbre configuration by introducing a conditional probability model, thereby achieving highly natural speech synthesis. During the generation process, the model selects the most appropriate timbre and speaking speed according to the context of the text, so that the generated audio has a consistent expression effect in a variety of situations. The system supports intelligent switching between Chinese and English timbre, ensuring the natural transition of bilingual audio content, avoiding the problems of stiff intonation and timbre mismatch in traditional speech synthesis. Through innovative context embedding and multimodal configuration algorithms, the listening audio generated by the model has achieved a higher level in speech naturalness, coherence and diversity, effectively improving students' listening comprehension ability.
[0174] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for intelligently generating audio for junior high school English listening based on a large model, characterized in that: The steps include: S1. Collect junior high school English listening test paper samples; S2. Clean the text of the junior high school English listening test paper sample, remove irrelevant characters and formats, annotate the Chinese text and the English text respectively, and obtain the annotated junior high school English listening test paper sample; S3. Based on the junior high school English listening test paper samples that have been cleaned and annotated, the large language model is trained to enable the large language model to distinguish between Chinese and English texts, and the large language model is tuned; S4. Use the trained large language model to automatically classify the text in the junior high school English listening test paper sample and distinguish the Chinese text from the English text; S5. Select different TTS timbres for Chinese text and English text respectively. The TTS timbres include different pronunciation timbres suitable for Chinese text and English text. Different speech speed, pitch and timbre parameters are configured for Chinese text and English text according to actual needs. S6, inputting the extracted Chinese text and English text into the corresponding TTS synthesis engine to generate Chinese audio and English audio respectively, and adjusting the speech speed and pitch parameters of the synthesized speech during the generation process to obtain the best auditory effect; S7. According to the structure of the junior high school English listening test paper sample, the synthesized Chinese audio and English audio are organized according to the test paper structure to generate a complete listening audio test paper; S8, storing the generated listening audio materials in a database, and updating the audio materials in real time according to the teaching progress and learner feedback data; The S5 comprises the following steps: S51, load the trained large language model and the junior high school English listening test sample dataset T final Input large language model; S52, using the conditional probability model with language context embedding for each text x i Calculate the joint probability that the text belongs to Chinese or English; S53, using the joint probability result in S52, assigning a language label to each text based on maximum likelihood estimation; S54, classify the language-classified text according to the title, topic, dialogue and monologue to generate a classified data set T class ; S55. Check the classified data set T class Whether it complies with the context consistency rule, if inconsistency is detected, it is based on the consistency detection function Ψ(T class ) for reclassification.
2. According to claim 1, a method for intelligently generating audio for junior high school English listening based on a large model is characterized in that: The S1 comprises the following steps: S11. Determine the sample range of English listening test papers for junior high schools, which includes official test papers provided by the education department, publicly published English listening textbook test papers, and mock test papers independently prepared by schools; S12. Screen the collected junior high school English listening test papers and only retain the sample data that is consistent with the junior high school English teaching syllabus. The sample data includes the complete test paper structure and content; S13. Perform structured storage on the collected junior high school English listening test paper samples to construct a junior high school English listening test paper sample set T: T={(D t ,Q s ,C d ,N m )∣D t ∈L,Q s ∈L,C d ∈L,N m ∈L}; Among them, D t Indicates the title part of the test paper, which is used to guide students to understand the question type. s Indicates the content of the question, clarifies the question or the task required of students. d Indicates the dialogue part in the test paper, simulating the communication between characters in real situations, N m It indicates the monologue part, which is used to train students' continuous listening comprehension ability. L indicates the language set, including Chinese L c and English L e , that is, L = L c ∪L e , used to identify the language type of each part.
3. According to the method of intelligent generation of junior high school English listening audio based on large model in claim 1, it is characterized in that: The S2 comprises the following steps: S21. Perform preliminary cleaning on the collected junior high school English listening test paper samples, remove non-text symbols, format codes and invisible characters in the junior high school English listening test paper samples, and retain the Chinese and English titles, questions, dialogues and monologues; S22, according to the language set L = L c ∪L e Perform language identification detection on sample data, where L c Indicates Chinese text, L e Represents English text, ensuring that the language type of each part of the content is correctly identified during subsequent processing; S23. Divide the preliminarily cleaned junior high school English listening test paper sample into headings D according to the content blocks t 、Question Q s 、Dialogue C d and monologue content N m , and use metadata tags to annotate the corresponding language type to construct a language annotation dataset T mark : T mark ={(D t ,L i ),(Q s ,L j ),(C d ,L k ),(N m ,L m )∣L i ,L j ,L k ,L m ∈{L c ,L e }}; Among them, L i ,L j ,L k ,L m Indicates the language type label of the corresponding content block, used to distinguish Chinese text from English text; S24, performing structure verification on the divided junior high school English listening test paper samples, wherein the structure verification rules include character length, sentence end identifier and language consistency, and excluding incomplete or incorrectly formatted data of the junior high school English listening test paper samples; S25. Fine-grained annotation is performed on the verified junior high school English listening test paper samples, using the tag set {tag1, tag2,…, tag n }Tag the key elements in the junior high school English listening test paper sample, where tag1 represents the name of a person or role, tag2 represents the time or date, tag3 represents the place name or place name, and other tags are n Used to identify key words in a specific context; S26, storing the marked junior high school English listening test paper samples according to language type and content block to generate the final junior high school English listening test paper sample T final : T final ={(D t ,Q s ,C d ,N m )∣D t ∈L i ,Q s ∈L j ,C d ∈L k ,N m ∈L m }。 4. The method for intelligently generating audio for junior high school English listening based on a large model according to claim 1 is characterized in that: The S3 comprises the following steps: S31, the junior high school English listening test sample data set T final Input the initial training module of the large language model to pre-train the large language model so that it has basic language understanding capabilities and learns the grammatical differences between Chinese and English texts; S32. Define the adaptive optimization loss function Loss(θ), which is adaptively adjusted according to student feedback: Among them, N is the number of samples, θ is the parameter set of the large language model, and y i represents the true label of the i-th sample, represents the predicted output of the model, α is the feedback adjustment coefficient, and F s Provide feedback data to students; S33, fine-tuning the large language model using an adaptive learning algorithm, inputting the student's feedback data into the large language model in real time after each round of training, updating the feedback adjustment coefficient α in the loss function, so that the parameter optimization process of the large language model gradually adapts to the student's learning rhythm and language level; S34. During the training of the large language model, the classification accuracy Acc of the large language model is monitored in real time: in, is an indicator function, which takes the value of 1 when the predicted result is consistent with the true label, otherwise it takes the value of 0; S35. Automatically adjust the difficulty and speech speed of the audio content generated by the large language model based on student feedback data and real-time monitoring results, so that the generated listening audio matches the student's actual ability level; S36, repeat the large language model training, feedback input, loss function update and fine-tuning process until the classification accuracy Acc of the model reaches the preset threshold T acc , and make the student's feedback index F s Meet the standards for adaptive learning.
5. The method for intelligently generating audio for junior high school English listening based on a large model according to claim 1 is characterized in that: The S5 specifically further comprises the following steps: The large language model is processed by initializing the parameter set θ0 to enter the inference state: in, is the initial predicted text and language type pairing set, x i For the i-th input text, L i is the Chinese or English language label, θ0 represents the initial large language model parameter set; For each text x, a conditional probability model with language context embedding is used i Calculate the joint probability that the text belongs to Chinese or English: Among them, P(L i ∣x i ,C(x i )) means in the context C(x i )Next text x i Belongs to language tag L i The probability of W Li and b Li For language tag L i The corresponding weight matrix and bias term, h(x i ) is the text x i The embedding representation of φ(C(x i )) is the text context C(x i ), M is the number of possible language types; Assign a language label to each text based on maximum likelihood estimation: Among them, L i For text x i The final language tag of The language-classified text is classified into titles, stems, dialogues, and monologues to generate a classified dataset T class ; Check the classified dataset T class Whether it complies with the context consistency rule, if inconsistency is detected, it is based on the consistency detection function Ψ(T class ) to reclassify: Among them, Ψ(T class ) is the consistency detection function, Indicates that under the context C(x), the text x should satisfy the corresponding language tag L, It is an indicator function. If the condition is met, the value is 1, otherwise it is 0. If the result is 0, the large language model is called for reclassification.
6. The method for intelligently generating audio for junior high school English listening based on a large model according to claim 1 is characterized in that: The S6 comprises the following steps: S61, select different TTS voice sets for Chinese text and English text, according to the language tag L i ∈{L c ,L e }Choose the appropriate pronunciation tone set S tts : S tts ={S c ,S e }Where S c Applicable to Chinese text, S e Applicable to English texts; Among them, S c and S e A collection of TTS voices for Chinese and English pronunciations; S62, configuring the speech speed, pitch and timbre parameters for the selected TTS timbre to establish an audio parameter set P; S63. Language label L based on input text i For each text segment x i Choose the appropriate sound set S i And the audio parameter P i : Among them, S i Represents text x i The corresponding timbre set, P i Represents text x i The speaking speed, pitch and timbre parameters; S64, input the selected TTS timbre and configured parameters into the TTS synthesis engine to generate the corresponding audio signal y i : y i =TTS(x i ,S i ,P i ); Among them, y i For text x i For the corresponding audio output, the TTS synthesis engine generates a natural and fluent speech signal based on the input text, timbre and parameters; S65, the generated audio signal y i Quality assessment is performed based on the audio feature consistency function Φ(y i ) Determine whether the audio meets the predetermined quality standard: in, It is an indicator function. When the audio signal meets the standards of speech rate, pitch and timbre parameters, the value is 1, otherwise it is 0. If the detection result does not meet the standards, the TTS engine is called to re-synthesize the audio.
7. The method for intelligently generating audio for junior high school English listening based on a large model according to claim 1 is characterized in that: The S7 comprises the following steps: S71, modularizing the generated Chinese audio and English audio according to the title, question stem, dialogue and monologue parts in the structure of the junior high school English listening test paper, so that the language label of each audio clip is consistent with the corresponding text part; S72. According to the teaching objectives of the junior high school English listening test and the learning needs of students, configure the audio playback order to set appropriate interval time and transition effects for each module audio to simulate the real listening test scene; S73. When splicing the audio of each module, adjust the audio connection method according to the content logic to ensure that the dialogue part remains smooth and coherent, and the monologue part is clear and complete; S74. Generate multiple versions of audio junior high school English listening test papers according to different student needs, including standard speed and slow speed versions. At the same time, generate customized audio junior high school English listening test papers containing only specific dialogues or monologues according to teacher needs; S75. Combine the metadata tags in the audio data to generate a playlist and control panel for the listening audio, supporting teachers and students to adjust the play order, repeat play and pause functions as needed; S76. Store the complete organized listening audio junior high school English listening test paper in a database and generate a unique identifier for it.
Citation Information
Patent Citations
Multi-language text automatic broadcast method and system
CN106856091A
An English test paper structuring method and device
CN109947836A