Oral language question and answer evaluation method and system based on large language model and intelligent voice
Through the oral question-and-answer evaluation method based on large language model and intelligent pronunciation, the integration of multiple answer modes and monitoring the dialogue process, the problem that the oral learning system in the existing technology cannot effectively evaluate and improve oral expression ability is solved, and more efficient oral proficiency evaluation and training is achieved.
Patent Information
- Application Number
- CN202510459914.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-01-09
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing oral learning system cannot effectively evaluate and improve oral expression skills, especially in the absence of an interactive learning environment after class, and the existing technology fails to comprehensively evaluate and targeted training of oral expression skills.
The spoken question-and-answer evaluation method based on large language model and intelligent pronunciation is adopted, and three spoken answer modes (standard answer mode, call answer mode and free answer mode) are integrated, and the corresponding answer content is generated through speech recognition and speech synthesis technology, and the entire dialogue process is monitored and evaluated.
It improves the efficiency and accuracy of oral evaluation, and can flexibly call the answer mode according to the oral level, guides oral reviewers to speak the corresponding content and improves their oral skills.
Smart Images

Figure CN120236608A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method and system for evaluating oral question - answering based on large - language models and intelligent speech. Background Art
[0002] Oral language is the most direct and commonly used way of human communication. Oral expression ability is very important both in work and in daily life. Good oral expression ability can effectively and accurately convey information, thereby improving communication efficiency. In addition, with the advent of globalization, it is also very necessary for professionals to master one or several foreign languages.
[0003] Mastering a language is not an easy task. The skills of reading and writing a language can generally achieve good results after a long - time study. However, for oral expression ability, the improvement process is particularly time - consuming and laborious. The reason is that the study of reading and writing a language can be completed independently by an individual, while the improvement of oral language ability requires an interactive learning environment, such as one - on - one oral guidance training.
[0004] Under the current learning environment and conditions, for the study of a language, especially a foreign language, the interactive learning environment is mainly realized in the classroom; after class, generally, an interactive learning environment cannot be provided. The lack of an interactive learning environment has led to the result that even after a long - time study, the oral expression ability is still difficult to improve substantially, and even basic daily communication cannot be carried out.
[0005] The existing technologies and methods related to language learning mostly focus on pronunciation evaluation, correction, scoring, etc., and there is no method and technology for comprehensively evaluating oral expression ability, nor is there a method and technology for targeted training and improvement based on the comprehensive evaluation results of oral expression ability.
[0006] After nearly a century of development, speech analysis and recognition technologies have become increasingly mature; with the rapid development of computer information technology and artificial intelligence, text file analysis technology and speech synthesis technology have also made great progress. These new technological breakthroughs have made it possible to develop interactive methods and technologies for evaluating and strengthening oral expression ability.
[0007] In existing oral language learning systems, some are dialogue-based oral language learning based on a standard answer library. According to the content of the learner's question, the answer content is directly matched in the standard answer library, and the matched answer content is output in voice. This dialogue method has poor intelligence. If no match is found in the standard answer library, the dialogue cannot continue. Some are free-style oral language learning. This implementation method cannot monitor the dialogue and cannot remind the user when the user cannot speak. The existing oral language learning systems are roughly designed and do not consider various situations to adopt different answer modes according to different situations.
[0008] There is a disclosed invention patent with the application number 2023105853137 and the name of "Oral Language Learning Method and Device Based on Large Language Model". The whole process adopts a free dialogue method based on the large language model, does not consider the integration of other dialogue methods, and cannot monitor the entire dialogue process to guide the user to say the corresponding dialogue. Summary of the Invention
[0009] In view of the problems and deficiencies existing in the prior art, the present invention provides an oral language question-and-answer evaluation method and system based on a large language model and intelligent voice.
[0010] The present invention solves the above technical problems through the following technical solutions:
[0011] The present invention provides an oral language question-and-answer evaluation method based on a large language model and intelligent voice, which is characterized in that it includes the following steps:
[0012] S1. The oral language evaluator inputs the target oral language dialogue scenario and the role played in the scenario, and calls the scenario large language model corresponding to the target oral language dialogue scenario as the target scenario large language model. Each oral language dialogue scenario corresponds to a scenario large language model, and the scenario large language model is a large language model constructed by deep learning using the oral language dialogues of the corresponding oral language dialogue scenarios;
[0013] S2. Convert the current input simulated oral language voice signal of the oral language evaluator into a digital format oral language voice signal, generate an original oral language voice file, and use speech recognition technology to generate an original text file;
[0014] S3. Determine whether the current original text file is a fixed question-and-answer sentence. If so, go to S4; otherwise, go to S5;
[0015] S4. Determine that the virtual robot answer mode is the standard answer mode, match the standard answer text corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for virtual robot voice output, and enter S8;
[0016] S5. Determine whether the current original text file is a first-time non-fixed Q&A sentence. If so, go to S6; otherwise, go to S7.
[0017] S6. Analyze the spoken language proficiency level of the spoken language evaluator based on the current original text file, and then go to S7.
[0018] S7. Determine the answering mode based on the original text file. When the determined answering mode is the standard answering mode, match the standard answering text corresponding to the spoken language proficiency level of the current original text file from the standard answering library, and use text-to-speech technology to generate a standard answering voice file for virtual robot voice output, then go to S8. When the determined answering mode is the retrieval answering mode, call the target scenario large language model to obtain real-time answering content and generate a retrieval answering text containing the real-time answering content corresponding to the spoken language proficiency level, and use text-to-speech technology to generate a retrieval answering voice file for virtual robot voice output, then go to S8. When the determined answering mode is the free answering mode, call the target scenario large language model to generate a free answering text corresponding to the spoken language proficiency level for the current original text file, and use text-to-speech technology to generate a free answering voice file for virtual robot voice output, then go to S8.
[0019] S8. Call the target scenario large language model to monitor all spoken language Q&As between the virtual robot and the spoken language evaluator, and analyze whether the spoken language Q&A is over. If not, go to S9; if so, go to S10.
[0020] S9. When it is monitored that the spoken language evaluator makes the next answer, go to S2.
[0021] S10. Conduct an evaluation for this spoken language Q&A and output the spoken language evaluation score of the spoken language evaluator.
[0022] The present invention also provides a spoken language Q&A evaluation system based on a large language model and intelligent speech, which is characterized in that it includes a spoken language input module, a file generation module, a first judgment module, a first determination module, a second judgment module, a spoken language proficiency analysis module, a second determination module, a spoken language end analysis module, a spoken language monitoring module, and a spoken language evaluation module.
[0023] The spoken language input module is used for the spoken language evaluator to input the target spoken language dialogue scenario and the scene role played, and call the scenario large language model corresponding to the target spoken language dialogue scenario as the target scenario large language model. Each spoken language dialogue scenario corresponds to a scenario large language model, and the scenario large language model is a large language model constructed by deep learning using the spoken language dialogues of the corresponding spoken language dialogue scenario.
[0024] The document generation module is used to convert the simulated spoken language voice signal currently input by the spoken language evaluator into a digital format spoken language voice signal, generate an original spoken language voice file, and generate an original text file using speech recognition technology;
[0025] The first judgment module is used to judge whether the current original text file is a fixed-form question-and-answer sentence. If it is, the first determination module is called; otherwise, the second judgment module is called;
[0026] The first determination module is used to determine that the virtual robot's answering mode is the standard answering mode, match the standard answering text corresponding to the current original text file from the standard answering library, and generate a standard answering voice file using speech synthesis technology for the virtual robot to output by voice, and call the spoken language end analysis module;
[0027] The second judgment module is used to judge whether the current original text file is a first-time non-fixed-form question-and-answer sentence. If it is, the spoken language proficiency analysis module is called; otherwise, the second determination module is called;
[0028] The spoken language proficiency analysis module is used to analyze the spoken language proficiency level of the spoken language evaluator based on the current original text file, and call the second determination module;
[0029] The second determination module is used to determine the answering mode based on the original text file. When the determined answering mode is the standard answering mode, match the standard answering text corresponding to the spoken language proficiency level of the current original text file from the standard answering library, and generate a standard answering voice file using speech synthesis technology for the virtual robot to output by voice, and call the spoken language end analysis module. When the determined answering mode is the retrieved answering mode, call the target scenario large language model retrieval system to obtain real-time answering content and generate a retrieved answering text containing the real-time answering content corresponding to the spoken language proficiency level, and generate a retrieved answering voice file using speech synthesis technology for the virtual robot to output by voice, and call the spoken language end analysis module. When the determined answering mode is the free answering mode, call the target scenario large language model to generate a free answering text corresponding to the current original text file, and generate a free answering voice file using speech synthesis technology for the virtual robot to output by voice, and call the spoken language end analysis module;
[0030] The spoken language end analysis module is used to call the target scenario large language model to monitor all spoken language questions and answers between the virtual robot and the spoken language evaluator, analyze whether the spoken language questions and answers are over. If not, the spoken language monitoring module is called; if so, the spoken language evaluation module is called;
[0031] The spoken language monitoring module is used to call the document generation module when it monitors the next answer from the spoken language evaluator;
[0032] The oral evaluation module is used to evaluate the oral response in this session and output the oral evaluation score of the oral evaluator.
[0033] The positive and progressive effects of the present invention are as follows:
[0034] The oral response evaluation method and system based on large language models and intelligent speech designed by the present invention integrate multiple oral response modes, which are divided into three modes: standard response mode, retrieval response mode, and free response mode. According to the oral content input by the oral evaluator, the corresponding mode is determined to be entered, and different response modes can be flexibly called to output the corresponding response speech content faster.
[0035] The present invention can monitor the entire conversation process, guide the oral evaluator to speak the corresponding oral conversation content, thereby improving the oral level of the oral evaluator.
[0036] The present invention can accurately analyze the oral level of the oral evaluator, making the entire conversation conform to the oral level of the oral evaluator, which is beneficial to improving the oral level of the oral evaluator. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the oral response evaluation method based on large language models and intelligent speech according to Embodiment 1 of the present invention.
[0038] Figure 2 It is a structural block diagram of the oral response evaluation system based on large language models and intelligent speech according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1
[0041] As Figure 1 shown, this embodiment provides an oral response evaluation method based on large language models and intelligent speech, which includes the following steps:
[0042] Step 101: The oral evaluation tester inputs the target oral dialogue scenario and the role to play in the scenario, and calls the scenario large language model corresponding to the target oral dialogue scenario as the target scenario large language model. Each oral dialogue scenario corresponds to a scenario large language model, and the scenario large language model is a large language model constructed by deep learning using the oral dialogues of the corresponding oral dialogue scenario.
[0043] For example, the oral dialogue scenario is a shopping oral dialogue scenario, a travel oral dialogue scenario, etc. The role to play in the scenario is a shopper, etc. In this step, a scenario large language model is constructed for each oral dialogue scenario, and the large language model is trained using the historical dialogues of the oral dialogue scenario to obtain the corresponding scenario large language model. Each oral dialogue scenario corresponds to a scenario large language model, which can improve the efficiency of oral interaction.
[0044] Step 102: Convert the current simulated oral speech signal input by the oral evaluation tester into a digital format oral speech signal, generate an original oral speech file, and use speech recognition technology to generate an original text file.
[0045] In this step, the corresponding original text file is obtained based on the original oral speech file for subsequent processing.
[0046] Step 103: Determine whether the current original text file is a fixed response sentence. If so, go to Step 104; otherwise, go to Step 105.
[0047] For example, Nice to meet you and Nice to meet you form a fixed response sentence, How are you and I’m fine form a fixed response sentence, and so on.
[0048] Step 104: Determine the virtual robot response mode as the standard response mode, match the standard response text corresponding to the current original text file from the standard response library, and use speech synthesis technology to generate a standard response speech file for virtual robot voice output, and then go to Step 109.
[0049] In this step, for the current original text file being a fixed response sentence, there is no need to call the target scenario large language model, and the corresponding standard response text can be directly and quickly matched from the standard response library, with high response efficiency.
[0050] Step 105: Determine whether the oral proficiency level of the oral evaluation tester is stored in the system. If so, go to Step 108; otherwise, go to Step 106.
[0051] In this step, if the oral proficiency level of the oral evaluator is pre-stored in the system, there is no need to further analyze the oral proficiency level of the oral evaluator, which improves the efficiency of oral interaction. If the oral proficiency level of the oral evaluator is not pre-stored, it is necessary to further analyze the oral proficiency level of the oral evaluator.
[0052] Step 106: Determine whether the current original text file is a first-time non-fixed question-and-answer sentence. If so, go to step 107; otherwise, go to step 108.
[0053] In this step, if the current original text file is a first-time non-fixed question-and-answer sentence, use this first-time non-fixed question-and-answer sentence to analyze the oral proficiency level of the oral evaluator. If not, it means that the oral proficiency level of the oral evaluator has been analyzed using the previous first-time non-fixed question-and-answer sentence and there is no need to analyze it again.
[0054] Step 107: Analyze the oral proficiency level of the oral evaluator based on the current original text file, and then go to step 108.
[0055] In step 107, analyze the oral proficiency level of the oral evaluator based on the grammar application, word application, or fixed phrase application in the current original text file: perform semantic analysis on the current original text file to analyze the grammar application and / or word application and / or fixed phrase application in the current original text file. For the analyzed grammar application and / or word application and / or fixed phrase application, eliminate the incorrectly applied grammar, words, and fixed phrases. For the remaining grammar application and / or word application and / or fixed phrase application, extract grammar features (such as the simple present tense, present continuous tense, present perfect tense, etc.) and / or word features and / or fixed phrase features in English to construct a text feature vector. Input the text feature vector into the trained oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the oral evaluator. The oral proficiency level analysis model is constructed based on a deep learning model, such as a convolutional neural network, etc., and the oral proficiency level analysis model is pre-trained with historical data. The extraction process of extracting grammar features, word features, and fixed phrase features is a prior art.
[0056] Further, in order to improve the accuracy of the analysis of the spoken language proficiency level, this embodiment analyzes the number of sub-features in the text feature vector. When the number of sub-features reaches the set number, it indicates that the feature samples are sufficient, and the text feature vector is input into the spoken language proficiency level analysis model corresponding to the target spoken language dialogue scenario for training and analysis, so as to analyze the spoken language proficiency level of the spoken language evaluator; when the number of sub-features does not reach the set number, it indicates that the feature samples are insufficient, and the original spoken language audio file is analyzed for speech, acoustic features, speech rate features, and fluency features are extracted to construct a speech feature vector, and the text feature vector and the speech feature vector are input into the spoken language proficiency level analysis model corresponding to the target spoken language dialogue scenario for training and analysis, so as to analyze the spoken language proficiency level of the spoken language evaluator. Among them, the extraction process of extracting acoustic features, speech rate features, and fluency features is a prior art. Making up for the insufficient number of feature samples can obtain a more accurate spoken language proficiency level.
[0057] Step 108: Determine the answering mode based on the original text file. When the determined answering mode is the standard answering mode, match the standard answering text corresponding to the spoken language proficiency level of the current original text file from the standard answering library, and use speech synthesis technology to generate a standard answering voice file for virtual robot voice output, and enter step 109; when the determined answering mode is the retrieval answering mode, call the large language model of the target scenario to retrieve the system to obtain real-time answering content and generate a retrieval answering text containing the real-time answering content corresponding to the spoken language proficiency level, and use speech synthesis technology to generate a retrieval answering voice file for virtual robot voice output, and enter step 109; when the determined answering mode is the free answering mode, call the large language model of the target scenario to generate a free answering text corresponding to the spoken language proficiency level for the current original text file, and use speech synthesis technology to generate a free answering voice file for virtual robot voice output, and enter step 109.
[0058] In this step, the retrieval answering mode is for answering that requires obtaining the current actual content. For example, when the spoken language evaluator asks what time it is now, what's the weather like today, what's the temperature now, etc., the large language model of the target scenario needs to obtain real-time answering content from the system and generate a retrieval answering text containing the real-time answering content corresponding to the spoken language proficiency level.
[0059] This step determines different response modes according to the content of different original text files. When there is a standard answer for the response content, there is no need to call the large language model of the target scenario. The response content corresponding to the spoken language level can be quickly matched from the standard response library. When the response content requires the cooperation of the current actual content, the large language model of the target scenario is called to obtain this real-time response content, and the response content containing the real-time response content corresponding to the spoken language level is generated. When the response content does not belong to the above two types, it is the free response mode, and the large language model of the target scenario is called to generate the response content corresponding to the spoken language level.
[0060] In step 108, further, when the determined response mode is the free response mode, the large language model of the target scenario is called to analyze whether the current original text file deviates from the target spoken language dialogue scenario. If it does not deviate, the free response text corresponding to the spoken language level is generated, and the free response voice file is generated using the speech synthesis technology and output by the virtual robot's voice. If it deviates, the free response text corresponding to the spoken language level and the scenario deviation prompt text are generated, and the free response voice file and the scenario deviation prompt voice file are generated using the speech synthesis technology and output by the virtual robot's voice.
[0061] The large language model of the target scenario can control the entire scenario dialogue process and give a reminder when the dialogue of the oral evaluation tester deviates from the target scenario, so that the entire oral dialogue process can always revolve around the target scenario.
[0062] Step 109: Call the large language model of the target scenario to monitor all oral questions and answers between the virtual robot and the oral evaluation tester, and analyze whether the oral questions and answers are over. If not, go to step 110; if so, go to step 111.
[0063] Step 110: When it is monitored that the oral evaluation tester gives the next complete answer, go to step 102. Or, if it is monitored that the oral evaluation tester has no oral input within the first set time (such as 20 seconds), the large language model of the target scenario gives a response prompt based on the previous oral questions and answers until the oral evaluation tester gives a complete answer and enters step 102. Or, if it is monitored that the oral evaluation tester has oral input within the second set time but does not form a complete sentence, the large language model of the target scenario gives a response prompt based on the previous oral questions and answers until the oral evaluation tester gives a complete answer and enters step 102. The second set time is greater than the first set time.
[0064] In this step, when the oral evaluation tester cannot give feedback on the dialogue with the virtual robot or gives feedback but the spoken language of the feedback is not a complete sentence (that is, a sentence cannot be spoken out completely), the large language model of the target scenario gives a response prompt based on the previous oral questions and answers until the oral evaluation tester gives a complete answer, which is beneficial to the smooth progress of the oral dialogue.
[0065] Step 111: Evaluate this oral Q&A and output the oral evaluation score of the oral evaluator.
[0066] In Step 111, for each original oral speech file of the oral evaluator, through speech synthesis technology, an original standard oral speech file is generated. For each original oral speech file of the oral evaluator, speech analysis is performed to extract acoustic features, speech rate features, and fluency features. For each original standard oral speech file of the oral evaluator, speech analysis is performed to extract acoustic features, speech rate features, and fluency features. The acoustic features (such as frequency, duration, etc.), speech rate features, and fluency features of each original oral speech file of the oral evaluator and the corresponding original standard oral speech file are compared respectively, and a single-speech evaluation score for each original oral speech file is obtained using a weighted calculation method. The average value of each single-speech evaluation score is calculated as the comprehensive speech evaluation score.
[0067] For example: For an original oral speech file and the corresponding original standard oral speech file, the ratio of their durations is used as the duration feature ratio, the ratio of their speech rates is used as the speech rate feature ratio, and the ratio of their fluencies is used as the fluency feature ratio. Certain weight values are assigned to the acoustic features, speech rate features, and fluency features respectively, and a single-speech evaluation score for this original oral speech file can be obtained using a weighted calculation method.
[0068] Semantic analysis is performed on each original text file of the oral evaluator to analyze whether there are errors in grammar application, word application, or fixed phrase application. The incorrect grammar application, word application, or fixed phrase application is corrected and a standard text file is generated. The grammar features, word features, and fixed phrase features extracted from each original text file of the oral evaluator and the corresponding standard text file are compared respectively, and a single-text evaluation score for each original text file is obtained using a weighted calculation method. The average value of each single-text evaluation score is calculated as the comprehensive text evaluation score.
[0069] For example: For a certain grammar feature, if it is applied correctly, it is 1; if it is applied incorrectly, it is 0. For a certain word feature, if it is applied correctly, it is 1; if it is applied incorrectly, it is 0. For a certain fixed phrase feature, if it is applied correctly, it is 1; if it is applied incorrectly, it is 0. Certain weight values are assigned to the grammar features, word features, and fixed phrase features respectively, and a single-text evaluation score for this original text file can be obtained using a weighted calculation method.
[0070] Calculate the sum of the comprehensive speech evaluation score and the comprehensive text evaluation score as the oral evaluation score of the oral evaluator and output it.
[0071] Determine whether the oral evaluation score of the oral evaluator reaches the set evaluation score. If it reaches, output that the oral proficiency level of the oral evaluator is the next higher oral proficiency level of the current oral proficiency level and store it. If it does not reach, output that the oral proficiency level of the oral evaluator is the current oral proficiency level and store it. After storing, when this oral evaluator conducts an oral dialogue in this scenario again, the pre-stored oral proficiency level can be directly utilized, without the need to analyze the oral proficiency level of the oral evaluator again, thereby improving the efficiency of oral dialogue.
[0072] In this embodiment, the oral Q&A dialogue composed of all the standard text files corresponding to the oral evaluator and all the response text files corresponding to the virtual robot is input into the large language model of the target scenario, and the large language model of the target scenario is fine-tuned to achieve the update and optimization of the large language model of the target scenario, making the large language model of the target scenario better and better.
[0073] As Figure 2 shown, this embodiment also provides an oral Q&A evaluation system based on a large language model and intelligent voice, which includes an oral input module 1, a file generation module 2, a first judgment module 3, a first determination module 4, a second judgment module 5, a third judgment module 6, an oral proficiency analysis module 7, a second determination module 8, an oral end analysis module 9, an oral monitoring module 10, and an oral evaluation module 11.
[0074] The oral input module 1 is used to allow the oral evaluator to input the target oral dialogue scenario and the scenario role played, and call the scenario large language model corresponding to the target oral dialogue scenario as the large language model of the target scenario. Each oral dialogue scenario corresponds to a scenario large language model, and the scenario large language model is a large language model constructed by deep learning using the oral dialogues of the corresponding oral dialogue scenario.
[0075] The file generation module 2 is used to convert the simulated oral speech signal currently input by the oral evaluator into a digital format oral speech signal, generate an original oral speech file, and generate an original text file using speech recognition technology.
[0076] The first judgment module 3 is used to judge whether the current original text file is a fixed Q&A sentence. If it is, call the first determination module 4; otherwise, call the third judgment module 6.
[0077] The first determination module 4 is used to determine that the virtual robot response mode is the standard response mode, match the standard response text corresponding to the current original text file from the standard response library, and generate a standard response speech file by voice synthesis for the virtual robot to output by voice, and call the oral end analysis module 9.
[0078] The third judgment module 6 is used to determine whether the oral proficiency level of the oral evaluator is stored in the system. If so, the second determination module 8 is called; otherwise, the second judgment module 5 is called.
[0079] The second judgment module 5 is used to determine whether the current original text file is a first-time non-fixed question-and-answer sentence. If so, the oral proficiency analysis module 7 is called; otherwise, the second determination module 8 is called.
[0080] The oral proficiency analysis module 7 is used to analyze the oral proficiency level of the oral evaluator based on the current original text file and call the second determination module 8.
[0081] The second determination module 8 is used to determine the response mode based on the original text file. When the determined response mode is the standard response mode, the standard response text corresponding to the oral proficiency level of the current original text file is matched from the standard response library, and the standard response voice file is generated using speech synthesis technology and output by the virtual robot's voice. The oral end analysis module 9 is called. When the determined response mode is the retrieval response mode, the target scenario large language model is called to retrieve the real-time response content of the system and generate a retrieval response text containing the real-time response content corresponding to the oral proficiency level, and the retrieval response voice file is generated using speech synthesis technology and output by the virtual robot's voice. The oral end analysis module 9 is called. When the determined response mode is the free response mode, the target scenario large language model is called to generate a free response text corresponding to the oral proficiency level for the current original text file, and the free response voice file is generated using speech synthesis technology and output by the virtual robot's voice. The oral end analysis module 9 is called.
[0082] The oral end analysis module 9 is used to call the target scenario large language model to monitor all oral questions and answers between the virtual robot and the oral evaluator, and analyze whether the oral questions and answers are over. If not, the oral monitoring module 10 is called; if so, the oral evaluation module 11 is called.
[0083] The oral monitoring module 10 is used to call the file generation module 2 when it monitors the next complete response of the oral evaluator, or, if it monitors that the oral evaluator has no oral input within the first set time, the target scenario large language model gives a response prompt based on the previous oral questions and answers until the oral evaluator gives a complete response, and then calls the file generation module 2, or, if it monitors that the oral evaluator has oral input but does not form a complete sentence within the second set time, the target scenario large language model gives a response prompt based on the previous oral questions and answers until the oral evaluator gives a complete response, and then calls the file generation module 2. The second set time is greater than the first set time.
[0084] The oral evaluation module 11 is used to evaluate this oral question and answer and output the oral evaluation score of the oral evaluator.
[0085] Example 2
[0086] Based on Example 1, this example optimizes steps 106 and 107, and can obtain a more accurate oral proficiency level of the oral evaluator.
[0087] In step 106, it is judged whether the current original text file is an initial non-fixed question-and-answer sentence. If so, go to step 107; otherwise, go to step 108. The initial stage refers to the non-fixed question-and-answer sentences output by the oral evaluator in the first N question-and-answer sessions, where N is a positive integer and 1 ≤ n ≤ N.
[0088] Step 107: Based on the current nth original text file, construct the nth text feature vector = the summary of all text feature vectors from the 1st to the nth. Analyze the number of sub-features in the nth text feature vector. When the number of sub-features reaches the set number, input the text feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the oral evaluator. When the number of sub-features does not reach the set number, based on the current nth original oral speech file, construct the nth speech feature vector = the summary of all speech feature vectors from the 1st to the nth, and input the nth text feature vector and the nth speech feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the oral evaluator, and then go to step 108.
[0089] In order to further improve the analysis accuracy of the oral proficiency level of the oral evaluator, this example uses the non-fixed question-and-answer sentences of the previous several times (such as the previous 3 times, one question and one answer is regarded as one time) to obtain the oral proficiency level of the oral evaluator, so as to improve the accuracy.
[0090] Although the specific implementation manners of the present invention have been described above, those skilled in the art should understand that these are only examples. The protection scope of the present invention is defined by the appended claims. Without departing from the principles and essence of the present invention, those skilled in the art can make various changes or modifications to these implementation manners, but these changes and modifications all fall within the protection scope of the present invention.
Claims
1. A spoken question-answering evaluation method based on a large language model and intelligent speech, characterized by comprising: S1. The oral evaluator inputs the target oral dialogue scene and the role played in the scene, and calls the scene large language model corresponding to the target oral dialogue scene as the target. Each oral dialogue scene corresponds to a scene large language model; S2, converting the analog spoken speech signal currently input by the oral evaluator into a digital format and generating an original spoken speech file, and generating an original text file through speech recognition; S3, determine whether the current original text file is a fixed question-answer sentence, if so, proceed to S4, otherwise proceed to S5; S4, determining that the virtual robot's answering mode is a standard answering mode, matching a standard answering text corresponding to the current original text file from a standard answering library, and generating a standard answering voice file for the virtual robot to output by voice, and entering S8; S5, determining whether the current original text file is the first non-fixed question-answer sentence, if so, proceed to S6, otherwise proceed to S7; S6, analyzing the oral proficiency level of the oral evaluator based on the current original text file, and entering S7; S7, determine the answer mode based on the original text file, when the answer mode is the standard answer mode, match the standard answer text of the oral level level corresponding to the current original text file from the standard answer library, and generate the standard answer voice file voice output, enter S8, when the answer mode is the retrieval answer mode, call the target scene large language model to retrieve the system to obtain the real-time answer content and generate the retrieval answer text of the corresponding oral level level, and generate the retrieval answer voice file voice output, enter S8, when the answer mode is the free answer mode, call the target scene large language model to generate the free answer text of the corresponding oral level level for the current original text file, and generate the free answer voice file voice output, enter S8; S8, calling the target scenario large language model to monitor all current spoken questions and answers, and analyzing whether the spoken question and answer is finished, if not, proceed to S9, if yes, proceed to S10; S9, when the next answer of the oral assessor is monitored, enter S2; S10. Output the oral assessment score of the oral assessor for this oral question-and-answer assessment.
2. The oral question-answering evaluation method based on a large language model and intelligent speech as claimed in claim 1, characterized in that In S9, if the oral evaluator's next complete answer is monitored, then the process goes to S2; or, if the oral evaluator is monitored without any oral input within the first set time, then the target scenario large language model gives answer prompts based on the previous oral questions and answers until the oral evaluator gives a complete answer, then the process goes to S2; or, if the oral evaluator is monitored with oral input within the second set time but it does not constitute a complete sentence, then the target scenario large language model gives answer prompts based on the previous oral questions and answers until the oral evaluator gives a complete answer, then the process goes to S2, and the second set time is greater than the first set time.
3. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 1, characterized in that: In S6, the oral proficiency level of the oral language assessor is analyzed based on the grammar application, word application or fixed phrase application in the current original text file: the current original text file is semantically analyzed to analyze the grammar application and / or word application and / or fixed phrase application therein, and incorrect grammar, words and fixed phrases are eliminated; for the retained grammar application and / or word application and / or fixed phrase application, grammatical features and / or word features and / or fixed phrase features are extracted to construct a text feature vector; the text feature vector is input into an oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the oral language assessor, wherein the oral proficiency level analysis model is constructed based on a deep learning model.
4. The oral question-answering evaluation method based on a large language model and intelligent speech as claimed in claim 3, characterized in that: The number of sub-features in the text feature vector is analyzed. When the number of sub-features reaches a set number, the text feature vector is input into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis, so as to analyze the oral proficiency level of the oral evaluator. When the number of sub-features does not reach the set number, the original oral speech file is subjected to speech analysis, the acoustic features, speech speed features and fluency features are extracted, and a speech feature vector is constructed. The text feature vector and the speech feature vector are input into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis, so as to analyze the oral proficiency level of the oral evaluator.
5. The oral question-answering evaluation method based on a large language model and intelligent speech as claimed in claim 4, characterized in that: In S5, it is determined whether the current original text file is an initial non-fixed question-answer sentence. If so, it proceeds to S6, otherwise it proceeds to S7; the initial stage refers to the non-fixed question-answer sentence output by the oral evaluator in the previous N questions and answers, where N is a positive integer; S6. Based on the current nth original text file, construct the nth text feature vector = the summary of all text feature vectors from the 1st to the nth time, analyze the number of sub-features in the nth text feature vector, and when the number of sub-features reaches the set number, input the text feature vector into the oral proficiency grade analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency grade of the oral language evaluator. When the number of sub-features does not reach the set number, construct the nth voice feature vector = the summary of all voice feature vectors from the 1st to the nth time based on the current nth original oral voice file, input the nth text feature vector and the nth voice feature vector into the oral proficiency grade analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency grade of the oral language evaluator, and enter S7.
6. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 1, characterized in that: In S7, when the determined answering mode is the free answering mode, the target scene large language model is called to analyze whether the current original text file deviates from the target oral dialogue scene. If it does not deviate, a free answering text corresponding to the oral level level is generated, and a free answering voice file is generated by speech synthesis technology for voice output by a virtual robot. If it deviates, a free answering text and a scene deviation prompting text corresponding to the oral level level are generated, and a free answering voice file and a scene deviation prompting voice file are generated by speech synthesis technology for voice output by a virtual robot.
7. The oral question-answering evaluation method based on a large language model and intelligent speech as claimed in claim 3, characterized in that In S10, for each original spoken speech file of the oral evaluator, an original standard spoken speech file is generated by speech synthesis technology, and speech analysis is performed on each original spoken speech file of the oral evaluator to extract acoustic features, speech speed features and fluency features, and speech analysis is performed on each original standard spoken speech file of the oral evaluator to extract acoustic features, speech speed features and fluency features, and the acoustic features, speech speed features and fluency features of each original spoken speech file of the oral evaluator and the corresponding original standard spoken speech file are compared respectively, and a single speech evaluation score of each original spoken speech file is obtained by a weighted calculation method, and the average value of each single speech evaluation score is calculated as the comprehensive speech evaluation score; Perform semantic analysis on each original text file of the oral evaluator to analyze whether there are errors in grammatical application, word application or fixed phrase application, correct the incorrect grammatical application, word application or fixed phrase application and generate a standard text file, compare the grammatical features, word features and fixed phrase features extracted from each original text file of the oral evaluator and the corresponding standard text file, and use a weighted calculation method to obtain a single text evaluation score for each original text file, and calculate the average of each single text evaluation score as the comprehensive text evaluation score; The sum of the comprehensive speech evaluation score and the comprehensive text evaluation score is calculated as the oral evaluation score of the oral evaluator and output.
8. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 7, characterized in that: Determine whether the oral assessment score of the oral assessor reaches the set assessment score, if so, output the oral level of the oral assessor as the oral level one level higher than the current oral level and store it, if not, output the oral level of the oral assessor as the current oral level and store it; The oral question-and-answer dialogue consisting of all standard text files corresponding to the oral evaluator and all answer texts corresponding to the virtual robot is input into the target scenario large language model, and the target scenario large language model is fine-tuned to achieve the update of the target scenario large language model.
9. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 8, characterized in that: S3, determine whether the current original text file is a fixed question-answer sentence, if so, enter S4, otherwise enter SA; SA, determine whether the oral proficiency level of the oral evaluator is stored in the system, if so, go to S7, otherwise go to S5.
10. A spoken question-answering evaluation system based on a large language model and intelligent speech, characterized by comprising: The oral input module is used for the oral evaluator to input the target oral dialogue scene and the role played in the scene, and call the scene large language model corresponding to the target oral dialogue scene as the target. Each oral dialogue scene corresponds to a scene large language model; The file generation module is used to convert the analog spoken speech signal currently input by the oral evaluator into a digital format and generate an original spoken speech file, and the speech recognition generates an original text file; The first judgment module is used to judge whether the current original text file is a fixed question-answer sentence, and if so, the first determination module is called, otherwise, the second judgment module is called; The first determination module is used to determine that the virtual robot's answering mode is a standard answering mode, match the standard answer text corresponding to the current original text file from the standard answer library, generate a standard answer voice file for the virtual robot to output by voice, and call the spoken end analysis module; The second judgment module is used to judge whether the current original text file is the first non-fixed question and answer sentence, and if so, the spoken language level analysis module is called, otherwise, the second determination module is called; The oral proficiency analysis module is used to analyze the oral proficiency level of the oral evaluator based on the current original text file and call the second determination module; The second determination module is used to determine the answer mode based on the original text file. If the answer mode is the standard answer mode, the standard answer text of the oral level level corresponding to the current original text file is matched from the standard answer library, and the standard answer voice file voice output is generated, and the oral end analysis module is called. If the answer mode is the retrieved answer mode, the target scene large language model is called to retrieve the system to obtain the real-time answer content and generate the retrieved answer text of the corresponding oral level level, and generate the retrieved answer voice file voice output, and call the oral end analysis module. If the answer mode is the free answer mode, the target scene large language model is called to generate the free answer text of the corresponding oral level level for the current original text file, and the free answer voice file voice output is generated, and the oral end analysis module is called; The spoken language end analysis module is used to call the target scene large language model, monitor all current spoken questions and answers, and analyze whether the spoken question and answer has ended. If not, the spoken language monitoring module is called, and if so, the spoken language evaluation module is called; The oral monitoring module is used to monitor the oral evaluator's next answer and then call the file generation module; The oral assessment module is used to output the oral assessment score of the oral assessor for this oral question-and-answer assessment.
Citation Information
Patent Citations
Artificial intelligent controlling method for discriminating robot speech
CN101075433A
Spoken language evaluation method through freely read topics and spoken language evaluation system thereof
CN105845134A
Inquiry processing method and system
CN109543020A
Method and platform for inputting voice during production operation process
CN109754805A
Joint modeling method, system and equipment for Mongolian dialogue model
CN113515952A