Oral question-answering evaluation method and system based on large language model and intelligent speech

Through an oral question and answer evaluation system based on large language model and intelligent pronunciation, a variety of answering modes are integrated and the dialogue process is monitored, the problem that the existing system cannot effectively evaluate and improve oral expression ability is solved, and the accurate analysis and training of oral proficiency is achieved.

CN120236608BActive Publication Date: 2025-09-02HEBEI UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510459914.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2025-01-09
Filing Date
2025-04-14
Publication Date
2025-09-02
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing oral learning system cannot effectively evaluate and improve oral expression ability, lacks an interactive learning environment, and cannot conduct targeted training and improvement based on oral expression ability.

Method used

A spoken question-answer evaluation method and system based on large language model and intelligent voice is designed, integrating three spoken answer modes (standard answer mode, call answer mode and free answer mode), and corresponding answer content is generated through speech recognition and speech synthesis technology, and the entire dialogue process is monitored and evaluated.

Benefits of technology

It improves the oral proficiency of oral reviewers, can accurately analyze oral proficiency levels and provide targeted training, which enhances the efficiency and effect of oral interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236608B_ABST
    Figure CN120236608B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for evaluating oral question and answer based on a large language model and intelligent voice. The method integrates multiple oral answer modes and divides the oral answer modes into three modes: standard answer mode, retrieved answer mode, and free answer mode. The method determines the corresponding mode to enter based on the oral content input by the oral evaluator. The method can flexibly call different answer modes and output the corresponding answer voice content more quickly. The method can monitor the entire conversation process and guide the oral evaluator to speak the corresponding oral conversation content, thereby improving the oral proficiency of the oral evaluator. The method can accurately analyze the oral proficiency level of the oral evaluator, so that the entire conversation conforms to the oral proficiency level of the oral evaluator, which is conducive to improving the oral proficiency of the oral evaluator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a spoken question-answering evaluation method and system based on a large language model and intelligent speech. Background Art

[0002] Spoken language is the most direct and common form of communication for humans. Oral communication skills are crucial both at work and in daily life. Good oral communication skills enable effective and accurate communication, thereby enhancing communication efficiency. Furthermore, with the advent of globalization, mastering one or more foreign languages ​​is essential for professionals.

[0003] Mastering a language is no easy task. While reading and writing skills generally achieve good results after a long period of study, improving oral communication skills is particularly time-consuming and laborious. This is because learning to read and write a language can be accomplished independently, while improving oral communication requires an interactive learning environment, such as one-on-one oral instruction.

[0004] In today's learning environment and conditions, language learning, especially foreign language learning, primarily relies on interactive learning environments in the classroom. After-school interactive learning environments are generally unavailable. This lack of an interactive learning environment results in a situation where, even after long periods of study, students struggle to make substantial improvements in their oral communication skills, and even struggle to engage in basic daily communication.

[0005] Existing technologies and methods related to language learning are mostly about pronunciation assessment, correction, and scoring. There is no method or technology for comprehensive assessment of oral expression ability, nor is there a method or technology for targeted training and improvement based on the comprehensive assessment results of oral expression ability.

[0006] After nearly a century of development, speech analysis and recognition technologies have matured. With the rapid advancement of computer information technology and artificial intelligence, text analysis and speech synthesis technologies have also made significant progress. These breakthroughs have made it possible to develop interactive methods and technologies for assessing and strengthening oral expression skills.

[0007] Some existing spoken language learning systems use a conversational approach based on a standard answer library. These systems directly match answers to learners' questions against the standard answer library and then output the matched answers as spoken words. This conversational approach lacks intelligence; if no answer is found in the standard answer library, the conversation cannot continue. Others use free-style spoken language learning, which lacks conversation monitoring and provides no notifications when users become unsuccessful. Existing spoken language learning systems are crudely designed and fail to account for diverse scenarios, requiring different response modes.

[0008] There is currently a patent for invention with application number 2023105853137, titled "Oral Learning Method and Device Based on Large Language Model." The full name adopts a free dialogue method based on a large language model, and does not consider the integration of other dialogue methods. It is also impossible to monitor the entire dialogue process and guide users to speak the corresponding dialogue. Summary of the Invention

[0009] In response to the problems and shortcomings of the prior art, the present invention provides a spoken question-answering evaluation method and system based on a large language model and intelligent speech.

[0010] The present invention solves the above technical problems through the following technical solutions:

[0011] The present invention provides a spoken question-answering evaluation method based on a large language model and intelligent speech, which is characterized in that it includes the following steps:

[0012] S1. The oral evaluator inputs the target oral dialogue scene and the role played in the scene, and calls the scene-based large language model corresponding to the target oral dialogue scene as the target scene-based large language model. Each oral dialogue scene corresponds to a scene-based large language model. The scene-based large language model is a large language model constructed by deep learning using the oral dialogue of the corresponding oral dialogue scene;

[0013] S2. Convert the analog spoken speech signal currently input by the oral evaluator into a digital spoken speech signal, generate an original spoken speech file, and generate an original text file using speech recognition technology;

[0014] S3, determine whether the current original text file is a fixed question and answer sentence, if so, go to S4, otherwise go to S5;

[0015] S4. Determine that the virtual robot's answer mode is the standard answer mode, match the standard answer text corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot to output, and then enter S8;

[0016] S5, determining whether the current original text file is the first non-fixed question-answer sentence, if so, proceeding to S6, otherwise proceeding to S7;

[0017] S6. Analyze the oral proficiency level of the oral evaluator based on the current original text file and proceed to S7;

[0018] S7. Determine the answer mode based on the original text file. When the determined answer mode is the standard answer mode, match the standard answer text of the oral proficiency level corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot voice output, and enter S8. When the determined answer mode is the retrieval answer mode, call the target scene large language model to retrieve the system to obtain real-time answer content and generate a retrieval answer text containing real-time answer content corresponding to the oral proficiency level, and use speech synthesis technology to generate a retrieval answer voice file for the virtual robot voice output, and enter S8. When the determined answer mode is the free answer mode, call the target scene large language model to generate a free answer text of the corresponding oral proficiency level for the current original text file, and use speech synthesis technology to generate a free answer voice file for the virtual robot voice output, and enter S8.

[0019] S8: Call the target scenario large language model to monitor all oral questions and answers between the virtual robot and the oral evaluator, and analyze whether the oral question and answer is completed. If not, go to S9; if yes, go to S10;

[0020] S9, when monitoring the oral evaluator's next answer, enter S2;

[0021] S10. Evaluate the oral question and answer session and output the oral evaluation score of the oral evaluator.

[0022] The present invention also provides a spoken question-answering evaluation system based on a large language model and intelligent speech, which is characterized in that it includes a spoken input module, a file generation module, a first judgment module, a first determination module, a second judgment module, a spoken level analysis module, a second determination module, a spoken end analysis module, a spoken monitoring module and a spoken evaluation module;

[0023] The spoken language input module is used for the spoken language evaluator to input the target spoken language dialogue scene and the role played in the scene, and call the scene large language model corresponding to the target spoken language dialogue scene as the target scene large language model. Each spoken language dialogue scene corresponds to a scene large language model. The scene large language model is a large language model constructed by deep learning using the spoken dialogue of the corresponding spoken language dialogue scene;

[0024] The file generation module is used to convert the analog spoken speech signal currently input by the oral evaluator into a spoken speech signal in digital format, generate an original spoken speech file, and generate an original text file using speech recognition technology;

[0025] The first judgment module is used to judge whether the current original text file is a fixed question-answer sentence, and if so, call the first determination module, otherwise call the second judgment module;

[0026] The first determination module is used to determine that the virtual robot's answer mode is the standard answer mode, match the standard answer text corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot to output, and call the spoken end analysis module;

[0027] The second judgment module is used to judge whether the current original text file is the first non-fixed question and answer sentence, and if so, call the spoken language level analysis module, otherwise call the second determination module;

[0028] The spoken language proficiency analysis module is used to analyze the spoken language proficiency level of the spoken language evaluator based on the current original text file and call the second determination module;

[0029] The second determination module is used to determine the answer mode based on the original text file. When the determined answer mode is the standard answer mode, a standard answer text of the oral proficiency level corresponding to the current original text file is matched from the standard answer library, and a standard answer voice file is generated by using speech synthesis technology for voice output by the virtual robot, and the spoken end analysis module is called. When the determined answer mode is the retrieved answer mode, the target scene large language model is called to retrieve the system to obtain real-time answer content and generate a retrieved answer text containing real-time answer content corresponding to the oral proficiency level, and a retrieved answer voice file is generated by using speech synthesis technology for voice output by the virtual robot, and the spoken end analysis module is called. When the determined answer mode is the free answer mode, the target scene large language model is called to generate a free answer text of the oral proficiency level for the current original text file, and a free answer voice file is generated by using speech synthesis technology for voice output by the virtual robot, and the spoken end analysis module is called.

[0030] The spoken language end analysis module is used to call the target scene large language model, monitor all spoken questions and answers of the virtual robot and the spoken language evaluator, and analyze whether the spoken question and answer is completed. If not, the spoken language monitoring module is called, and if so, the spoken language evaluation module is called;

[0031] The spoken language monitoring module is used to monitor the next answer of the spoken language evaluator and then call the file generation module;

[0032] The oral evaluation module is used to evaluate the oral question and answer and output the oral evaluation score of the oral evaluator.

[0033] The positive progress effect of the present invention is:

[0034] The oral question-and-answer evaluation method and system designed by the present invention, based on a large language model and intelligent voice, integrates multiple oral answering modes and divides the oral answering modes into three modes: standard answering mode, retrieved answering mode and free answering mode. The corresponding mode is determined according to the oral content input by the oral evaluator, and different answering modes can be flexibly called to output the corresponding answering voice content more quickly.

[0035] The present invention can monitor the entire conversation process and guide the oral evaluator to speak the corresponding oral conversation content, thereby improving the oral evaluator's oral level.

[0036] The present invention can accurately analyze the oral proficiency level of the oral evaluator, so that the entire conversation conforms to the oral proficiency level of the oral evaluator, which is conducive to improving the oral proficiency of the oral evaluator. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a flowchart of the oral question-answering evaluation method based on a large language model and intelligent speech according to Example 1 of the present invention.

[0038] Figure 2 This is a structural block diagram of the oral question-answering evaluation system based on a large language model and intelligent speech according to Example 1 of the present invention. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0040] Example 1

[0041] like Figure 1 As shown, this embodiment provides a spoken question-answering evaluation method based on a large language model and intelligent speech, which includes the following steps:

[0042] Step 101: The oral evaluator inputs the target oral dialogue scene and the role played in the scene, and calls the scene large language model corresponding to the target oral dialogue scene as the target scene large language model. Each oral dialogue scene corresponds to a scene large language model. The scene large language model is a large language model constructed by deep learning using the oral dialogue of the corresponding oral dialogue scene.

[0043] For example, the spoken dialogue scenarios include shopping and travel. The characters in the scenarios are shoppers. In this step, a scenario-based large language model is constructed for each spoken dialogue scenario. The large language model is trained using the historical conversations of the spoken dialogue scenario to obtain the corresponding scenario-based large language model. Each spoken dialogue scenario corresponds to a scenario-based large language model, which can improve the efficiency of spoken interaction.

[0044] Step 102: Convert the analog spoken speech signal currently input by the spoken language evaluator into a digital spoken speech signal, generate an original spoken speech file, and use speech recognition technology to generate an original text file.

[0045] In this step, the corresponding original text file is obtained based on the original spoken voice file for subsequent processing.

[0046] Step 103 , determine whether the current original text file is a fixed question-answer sentence. If so, proceed to step 104 ; otherwise, proceed to step 105 .

[0047] For example: Nice to meet you and Nice to meet you constitute a fixed question and answer sentence, How are you and I'm fine constitute a fixed question and answer sentence, and so on.

[0048] Step 104 , determine that the virtual robot's answering mode is the standard answering mode, match the standard answer text corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot to output, and then enter step 109 .

[0049] In this step, since the current original text file is a fixed question-answer sentence, there is no need to call the target scenario large language model. The corresponding standard answer text can be directly and quickly matched from the standard answer library, and the response is efficient.

[0050] Step 105 , determining whether the oral proficiency level of the oral evaluator is stored in the system, if so, proceeding to step 108 , otherwise proceeding to step 106 .

[0051] In this step, if the oral proficiency level of the oral evaluator is pre-stored in the system, there is no need to analyze the oral proficiency level of the oral evaluator, thereby improving the efficiency of oral interaction. If the oral proficiency level of the oral evaluator is not pre-stored, further analysis of the oral proficiency level of the oral evaluator is required.

[0052] Step 106 , determining whether the current original text file is the first non-fixed question-answer sentence, if so, proceeding to step 107 , otherwise proceeding to step 108 .

[0053] In this step, if the current original text file is the first non-fixed question and answer sentence, the first non-fixed question and answer sentence is used to analyze the oral proficiency level of the oral evaluator. If not, it indicates that the oral proficiency level of the oral evaluator has been analyzed using the previous first non-fixed question and answer sentence and there is no need to analyze it again.

[0054] Step 107 : Analyze the oral proficiency level of the oral evaluator based on the current original text file, and proceed to step 108 .

[0055] In step 107, the oral proficiency level of the spoken language assessor is analyzed based on the grammar application, word application, or fixed phrase application in the current original text file: the current original text file is semantically analyzed to analyze the grammar application and / or word application and / or fixed phrase application in the current original text file; grammar, words, and fixed phrases with incorrect application are eliminated from the analyzed grammar application and / or word application and / or fixed phrase application; grammatical features (such as English general tense, progressive tense, perfect tense, etc.) and / or word features and / or fixed phrase features are extracted from the retained grammar application and / or word application and / or fixed phrase application; a text feature vector is constructed; and the text feature vector is input into a trained oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the spoken language assessor. The oral proficiency level analysis model is constructed based on a deep learning model, such as a convolutional neural network, and is pre-trained using historical data. The extraction process of grammatical features, word features, and fixed phrase features is conventional.

[0056] Furthermore, in order to improve the accuracy of oral proficiency level analysis, this embodiment analyzes the number of sub-features in the text feature vector. When the number of sub-features reaches a set number, it indicates that the feature samples are sufficient. The text feature vector is input into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency level of the oral evaluator. When the number of sub-features does not reach the set number, it indicates that the feature samples are insufficient. The original oral speech file is subjected to speech analysis to extract acoustic features, speech rate features, and fluency features, and a speech feature vector is constructed. The text feature vector and the speech feature vector are input into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency level of the oral evaluator. Among them, the extraction process of extracting acoustic features, speech rate features, and fluency features is an existing technology. When the number of feature samples is insufficient, supplementing them can obtain a more accurate oral proficiency level.

[0057] Step 108: Determine the answer mode based on the original text file. When the determined answer mode is the standard answer mode, match the standard answer text of the oral proficiency level corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot voice output, and enter step 109; when the determined answer mode is the retrieval answer mode, call the target scene large language model to retrieve the system to obtain real-time answer content and generate a retrieval answer text containing real-time answer content corresponding to the oral proficiency level, and use speech synthesis technology to generate a retrieval answer voice file for the virtual robot voice output, and enter step 109; when the determined answer mode is the free answer mode, call the target scene large language model to generate a free answer text of the corresponding oral proficiency level for the current original text file, and use speech synthesis technology to generate a free answer voice file for the virtual robot voice output, and enter step 109.

[0058] In this step, the retrieval answer mode is aimed at those who need to obtain the current actual content before answering. For example, when the oral evaluator asks what time it is, what is the weather like today, what is the temperature now, etc., the target scenario large language model is required to obtain the real-time answer content from the system and generate a retrieval answer text containing the real-time answer content corresponding to the oral proficiency level.

[0059] This step determines to enter different answering modes according to the different contents of the original text files. When the answer content has a standard answer, there is no need to call the target scenario large language model, and the answer content of the corresponding oral level level can be quickly matched from the standard answer library. When the answer content requires the cooperation of the current actual content, the target scenario large language model is called to obtain this real-time answer content, and generate the answer content containing the real-time answer content of the corresponding oral level level. When the answer content does not belong to the above two modes, it is the free answer mode, and the target scenario large language model is called to generate the answer content of the corresponding oral level level.

[0060] In step 108, further, when the determined answering mode is the free answering mode, the target scene large language model is called to analyze whether the current original text file deviates from the target oral dialogue scene. If it does not deviate, a free answering text corresponding to the oral level level is generated, and a free answering voice file is generated by speech synthesis technology and output by the virtual robot voice. If it deviates, a free answering text and a scene deviation prompt text corresponding to the oral level level are generated, and a free answering voice file and a scene deviation prompt voice file are generated by speech synthesis technology and output by the virtual robot voice.

[0061] The target scene large language model can control the entire scene dialogue process and remind the oral evaluator when the dialogue deviates from the target scene, so that the entire oral dialogue process can always revolve around the target scene.

[0062] Step 109 : Call the target scenario large language model to monitor all spoken questions and answers between the virtual robot and the oral evaluator, and analyze whether the spoken question and answer is completed. If not, proceed to step 110 ; if so, proceed to step 111 .

[0063] Step 110: If the oral evaluator's next complete answer is monitored, then the process proceeds to step 102; or, if the oral evaluator is monitored to have no oral input within a first set time (e.g., 20 seconds), then the target scenario large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator gives a complete answer, then the process proceeds to step 102; or, if the oral evaluator is monitored to have oral input but does not constitute a complete sentence within a second set time, then the target scenario large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator gives a complete answer, then the process proceeds to step 102; the second set time is greater than the first set time.

[0064] In this step, when the oral evaluator is unable to provide feedback on the conversation with the virtual robot or provides feedback but the feedback is not a complete sentence (that is, the sentence cannot be completed), the target scene large language model gives answer prompts based on the previous oral questions and answers until the oral evaluator gives a complete answer, which is conducive to the smooth progress of the oral conversation.

[0065] Step 111: Evaluate the oral question and answer session, and output the oral evaluation score of the oral evaluator.

[0066] In step 111, an original standard spoken speech file is generated from each original spoken speech file of the oral language evaluator through speech synthesis technology, and speech analysis is performed on each original spoken speech file of the oral language evaluator to extract acoustic features, speaking rate features, and fluency features. The acoustic features (such as frequency, duration, etc.), speaking rate features, and fluency features of each original spoken speech file of the oral language evaluator and the corresponding original standard spoken speech file are compared respectively, and a single speech evaluation score of each original spoken speech file is obtained using a weighted calculation method, and the average of each single speech evaluation score is calculated as the comprehensive speech evaluation score.

[0067] For example, for an original spoken speech file and a corresponding original standard spoken speech file, the ratio of their durations is used as the duration feature ratio, the ratio of their speaking speeds is used as the speaking speed feature ratio, and the ratio of their fluency is used as the fluency feature ratio. The acoustic features, speaking speed features, and fluency features are respectively assigned certain weight values, and a weighted calculation method is used to obtain a single speech evaluation score for the original spoken speech file.

[0068] A semantic analysis is performed on each original text file of the oral evaluator to analyze whether there are errors in the application of grammar, words or fixed phrases. The incorrect grammar, words or fixed phrases are corrected and a standard text file is generated. The grammatical features, word features and fixed phrase features extracted from each original text file of the oral evaluator and the corresponding standard text file are compared respectively, and a weighted calculation method is used to obtain the individual text evaluation score of each original text file, and the average of each individual text evaluation score is calculated as the comprehensive text evaluation score.

[0069] For example, a grammatical feature is scored as 1 if used correctly and as 0 if used incorrectly; a word feature is scored as 1 if used correctly and as 0 if used incorrectly; and a fixed phrase feature is scored as 1 if used correctly and as 0 if used incorrectly. Grammatical features, word features, and fixed phrase features are each assigned a weighted value, and a weighted calculation method is used to obtain a single text evaluation score for the original text file.

[0070] The sum of the comprehensive voice evaluation score and the comprehensive text evaluation score is calculated as the oral evaluation score of the oral evaluator and output.

[0071] The system determines whether the speaker's oral proficiency score reaches the set score. If so, the system outputs and stores the speaker's oral proficiency level as the next higher than the current one. If not, the system outputs and stores the speaker's oral proficiency level as the current one. Once stored, the speaker can directly utilize this pre-stored oral proficiency level when engaging in oral dialogue in the same scenario, eliminating the need to analyze the speaker's oral proficiency level, thereby improving conversation efficiency.

[0072] In this embodiment, the oral question-and-answer dialogue consisting of all standard text files corresponding to the oral evaluator and all answer texts corresponding to the virtual robot is input into the target scenario large language model, and the target scenario large language model is fine-tuned to achieve update and optimization of the target scenario large language model, making the target scenario large language model better and better.

[0073] like Figure 2 As shown, this embodiment also provides a spoken question and answer evaluation system based on a large language model and intelligent voice, which includes a spoken input module 1, a file generation module 2, a first judgment module 3, a first determination module 4, a second judgment module 5, a third judgment module 6, a spoken level analysis module 7, a second determination module 8, a spoken end analysis module 9, a spoken monitoring module 10 and a spoken evaluation module 11.

[0074] The oral input module 1 is used for the oral evaluator to input the target oral dialogue scene and the role played in the scene, and call the scene large language model corresponding to the target oral dialogue scene as the target scene large language model. Each oral dialogue scene corresponds to a scene large language model. The scene large language model is a large language model constructed by deep learning of the oral dialogue of the corresponding oral dialogue scene.

[0075] The file generation module 2 is used to convert the analog spoken speech signal currently input by the oral evaluator into a spoken speech signal in digital format, generate an original spoken speech file, and generate an original text file using speech recognition technology.

[0076] The first judgment module 3 is used to judge whether the current original text file is a fixed question-answer sentence, and if so, calls the first determination module 4, otherwise calls the third judgment module 6.

[0077] The first determination module 4 is used to determine that the virtual robot's answer mode is the standard answer mode, match the standard answer text corresponding to the current original text file from the standard answer library, and use speech synthesis technology to generate a standard answer voice file for the virtual robot voice output, and call the spoken end analysis module 9.

[0078] The third judging module 6 is used to judge whether the oral proficiency level of the oral evaluator is stored in the system. If so, the second determining module 8 is called; otherwise, the second judging module 5 is called.

[0079] The second judgment module 5 is used to judge whether the current original text file is the first non-fixed question and answer sentence, and if so, call the spoken language level analysis module 7, otherwise call the second determination module 8.

[0080] The spoken language proficiency analysis module 7 is used to analyze the spoken language proficiency level of the spoken language evaluator based on the current original text file and call the second determination module 8.

[0081] The second determination module 8 is used to determine the answer mode based on the original text file. When the determined answer mode is the standard answer mode, the standard answer text of the oral level level corresponding to the current original text file is matched from the standard answer library, and the standard answer voice file is generated by the speech synthesis technology for voice output by the virtual robot, and the oral end analysis module 9 is called. When the determined answer mode is the retrieval answer mode, the target scene large language model is called to retrieve the system to obtain real-time answer content and generate a retrieval answer text containing real-time answer content corresponding to the oral level level, and the speech synthesis technology is used to generate the retrieval answer voice file for voice output by the virtual robot, and the oral end analysis module 9 is called. When the determined answer mode is the free answer mode, the target scene large language model is called to generate a free answer text corresponding to the oral level level for the current original text file, and the speech synthesis technology is used to generate a free answer voice file for voice output by the virtual robot, and the oral end analysis module 9 is called.

[0082] The spoken language end analysis module 9 is used to call the target scene large language model, monitor all spoken questions and answers between the virtual robot and the spoken language evaluator, and analyze whether the spoken question and answer is ended. If not, the spoken language monitoring module 10 is called, and if so, the spoken language evaluation module 11 is called.

[0083] The oral monitoring module 10 is used to call the file generation module 2 when it monitors the oral evaluator's next complete answer, or, if it monitors that the oral evaluator has no oral input within the first set time, the target scene large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator completes the answer, and calls the file generation module 2, or, if it monitors that the oral evaluator has oral input but does not constitute a complete sentence within the second set time, the target scene large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator completes the answer, and calls the file generation module 2, and the second set time is greater than the first set time.

[0084] The oral evaluation module 11 is used to evaluate the oral question and answer and output the oral evaluation score of the oral evaluator.

[0085] Example 2

[0086] Based on Example 1, this embodiment optimizes step 106 and step 107, and can obtain a more accurate oral proficiency level of the oral evaluator.

[0087] In step 106, it is determined whether the current original text file is an initial non-fixed question-answer sentence. If so, the process proceeds to step 107; otherwise, the process proceeds to step 108; the initial stage refers to the non-fixed question-answer sentence output by the oral evaluator in the first N questions and answers, where N is a positive integer, 1≤n≤N.

[0088] Step 107: Based on the current nth original text file, construct an nth text feature vector = a summary of all text feature vectors from the 1st to the nth time, analyze the number of sub-features in the nth text feature vector, and when the number of sub-features reaches a set number, input the text feature vector into a spoken language proficiency grade analysis model corresponding to the target spoken language dialogue scenario for training and analysis to analyze the spoken language proficiency grade of the spoken language evaluator; when the number of sub-features does not reach the set number, construct an nth speech feature vector = a summary of all speech feature vectors from the 1st to the nth time based on the current nth original spoken voice file, input the nth text feature vector and the nth speech feature vector into a spoken language proficiency grade analysis model corresponding to the target spoken language dialogue scenario for training and analysis to analyze the spoken language proficiency grade of the spoken language evaluator, and then proceed to step 108.

[0089] In order to further improve the analysis accuracy of the oral proficiency level of the oral evaluator, this embodiment uses the non-fixed question and answer sentences of the first few times (such as the first three times, one question and one answer as one time) to obtain the oral proficiency level of the oral evaluator, thereby improving accuracy.

[0090] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely illustrative and that the scope of the present invention is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present invention, and such changes and modifications are intended to fall within the scope of the present invention.

Claims

1. A method for evaluating spoken question-answering based on a large language model and intelligent speech, comprising: S1. The oral evaluator inputs the target oral dialogue scene and the role played in the scene, and calls the scene-based large language model corresponding to the target oral dialogue scene as the target. Each oral dialogue scene corresponds to a scene-based large language model; S2. Convert the analog spoken speech signal currently input by the oral evaluator into a digital format and generate an original spoken speech file, and generate an original text file through speech recognition; S3. Determine whether the current original text file is a fixed question-answer sentence. If so, proceed to S4. Otherwise, proceed to S5. S4. Determine that the virtual robot's answering mode is the standard answering mode, match the standard answer text corresponding to the current original text file from the standard answer library, and generate a standard answer voice file for the virtual robot to output, and enter S8; S5. Determine whether the current original text file is the first non-fixed question-answer sentence. If so, proceed to S6. Otherwise, proceed to S7. S6. Analyze the oral proficiency level of the oral evaluator based on the current original text file and proceed to S7; S7. Determine the answer mode based on the original text file. When the answer mode is the standard answer mode, match the standard answer text of the oral proficiency level corresponding to the current original text file from the standard answer library, and generate a standard answer voice file for voice output, and enter S8. When the answer mode is the retrieval answer mode, call the target scene large language model to retrieve the system to obtain real-time answer content and generate the retrieval answer text of the corresponding oral proficiency level, and generate the retrieval answer voice file for voice output, and enter S8. When the answer mode is the free answer mode, call the target scene large language model to analyze whether the current original text file deviates from the target oral dialogue scene. If not, generate the free answer text of the corresponding oral proficiency level, and use speech synthesis technology to generate the free answer voice file for virtual robot voice output. If deviated, generate the free answer text and scene deviation prompt text of the corresponding oral proficiency level, and use speech synthesis technology to generate the free answer voice file and scene deviation prompt voice file for virtual robot voice output, and enter S8. S8: Call the target scenario large language model to monitor all current spoken questions and answers, and analyze whether the spoken question and answer is completed. If not, proceed to S9; if so, proceed to S10; S9: If the oral evaluator's next complete answer is monitored, then the process proceeds to S2. Alternatively, if the oral evaluator has not provided any spoken input within the first set time, then the target scenario large language model provides answer prompts based on the previous spoken questions and answers until the oral evaluator provides a complete answer, then the process proceeds to S2. Alternatively, if the oral evaluator has provided spoken input but does not form a complete sentence within the second set time, then the target scenario large language model provides answer prompts based on the previous spoken questions and answers until the oral evaluator provides a complete answer, then the process proceeds to S2. The second set time is greater than the first set time. S10. Output the oral assessment score of the oral assessor for this oral question-and-answer assessment.

2. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 1, wherein: In S6, the oral proficiency level of the oral evaluator is analyzed based on the grammar application, word application or fixed phrase application in the current original text file: the current original text file is semantically analyzed to analyze the grammar application and / or word application and / or fixed phrase application therein, and incorrect grammar, words and fixed phrases are eliminated. For the retained grammar application and / or word application and / or fixed phrase application, grammatical features and / or word features and / or fixed phrase features are extracted to construct a text feature vector, and the text feature vector is input into the oral proficiency level analysis model corresponding to the target oral dialogue scenario for training and analysis to analyze the oral proficiency level of the oral evaluator, wherein the oral proficiency level analysis model is constructed based on a deep learning model.

3. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 2, characterized in that: Analyze the number of sub-features in the text feature vector. When the number of sub-features reaches a set number, input the text feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency level of the oral evaluator. When the number of sub-features does not reach the set number, perform speech analysis on the original oral speech file, extract acoustic features, speaking speed features and fluency features, construct a speech feature vector, input the text feature vector and the speech feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training analysis to analyze the oral proficiency level of the oral evaluator.

4. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 3, wherein: In S5, it is determined whether the current original text file is an initial non-fixed question-answer sentence. If so, the process proceeds to S6; otherwise, the process proceeds to S7. The initial stage refers to the non-fixed question-answer sentence output by the oral evaluator in the first N questions and answers, where N is a positive integer. S6. Based on the current nth original text file, construct the nth text feature vector = the summary of all text feature vectors from the 1st to the nth time, analyze the number of sub-features in the nth text feature vector, and when the number of sub-features reaches the set number, input the text feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training and analysis to analyze the oral proficiency level of the oral language evaluator. When the number of sub-features does not reach the set number, based on the current nth original oral voice file, construct the nth voice feature vector = the summary of all voice feature vectors from the 1st to the nth time, input the nth text feature vector and the nth voice feature vector into the oral proficiency level analysis model corresponding to the target oral dialogue scene for training and analysis to analyze the oral proficiency level of the oral language evaluator, and enter S7.

5. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 2, wherein In S10, for each original spoken speech file of the oral language evaluator, an original standard spoken speech file is generated by speech synthesis technology, each original spoken speech file of the oral language evaluator is subjected to speech analysis, and acoustic features, speech rate features, and fluency features are extracted; each original standard spoken speech file of the oral language evaluator is subjected to speech analysis, and acoustic features, speech rate features, and fluency features are extracted; the acoustic features, speech rate features, and fluency features of each original spoken speech file of the oral language evaluator and the corresponding original standard spoken speech file are respectively compared, and a single speech evaluation score of each original spoken speech file is obtained by using a weighted calculation method, and the average of each single speech evaluation score is calculated as the comprehensive speech evaluation score; Perform semantic analysis on each original text file of the oral evaluator to analyze whether there are errors in grammatical application, word application, or fixed phrase application, correct the incorrect grammatical application, word application, or fixed phrase application and generate a standard text file, compare the grammatical features, word features, and fixed phrase features extracted from each original text file of the oral evaluator with the corresponding standard text file, and use a weighted calculation method to obtain a single text evaluation score for each original text file, and calculate the average of each single text evaluation score as the comprehensive text evaluation score; The sum of the comprehensive voice evaluation score and the comprehensive text evaluation score is calculated as the oral evaluation score of the oral evaluator and output.

6. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 5, characterized in that: Determine whether the oral assessment score of the oral evaluator reaches the set assessment score. If so, output the oral proficiency level of the oral evaluator as the next higher oral proficiency level than the current oral proficiency level and store it; if not, output the oral proficiency level of the oral evaluator as the current oral proficiency level and store it; The oral question-and-answer dialogue consisting of all standard text files corresponding to the oral evaluator and all answer texts corresponding to the virtual robot is input into the target scenario large language model, and the target scenario large language model is fine-tuned to achieve the update of the target scenario large language model.

7. The oral question-answering evaluation method based on a large language model and intelligent speech according to claim 6, characterized in that: S3, determine whether the current original text file is a fixed question and answer sentence, if so, go to S4, otherwise go to SA; SA: Determine whether the oral proficiency level of the oral evaluator is stored in the system. If so, go to S7; otherwise, go to S5.

8. A spoken question-answering evaluation system based on a large language model and intelligent speech, comprising: The spoken language input module is used for the spoken language evaluator to input the target spoken dialogue scene and the role played in the scene, and call the scene large language model corresponding to the target spoken dialogue scene as the target. Each spoken dialogue scene corresponds to a scene large language model; The file generation module is used to convert the analog spoken speech signal currently input by the oral evaluator into a digital format and generate the original spoken speech file, and the speech recognition generates the original text file; The first judgment module is used to judge whether the current original text file is a fixed question and answer sentence. If so, the first determination module is called, otherwise the second judgment module is called; The first determination module is used to determine that the virtual robot's answer mode is the standard answer mode, match the standard answer text corresponding to the current original text file from the standard answer library, generate a standard answer voice file for the virtual robot to output, and call the spoken end analysis module; The second judgment module is used to judge whether the current original text file is the first non-fixed question and answer sentence. If so, the spoken language level analysis module is called, otherwise the second determination module is called; The oral proficiency analysis module is used to analyze the oral proficiency level of the oral evaluator based on the current original text file and call the second determination module; The second determination module is used to determine the answer mode based on the original text file. If the answer mode is the standard answer mode, the standard answer text of the oral level level corresponding to the current original text file is matched from the standard answer library, and the standard answer voice file is generated for voice output, and the oral end analysis module is called. If the answer mode is the retrieval answer mode, the target scene large language model is called to retrieve the system to obtain real-time answer content and generate the retrieval answer text corresponding to the oral level level, and generate the retrieval answer voice file for voice output, and the oral end analysis module is called. If the answer mode is the free answer mode, the target scene large language model is called to analyze whether the current original text file deviates from the target oral dialogue scene. If not, a free answer text corresponding to the oral level level is generated, and a free answer voice file is generated by using speech synthesis technology and output by the virtual robot voice. If deviated, a free answer text and a scene deviation prompt text corresponding to the oral level level are generated, and a free answer voice file and a scene deviation prompt voice file are generated by using speech synthesis technology and output by the virtual robot voice, and the oral end analysis module is called; The spoken word end analysis module is used to call the target scenario large language model, monitor all current spoken questions and answers, and analyze whether the spoken question and answer has ended. If not, it calls the spoken word monitoring module; if so, it calls the spoken word evaluation module; The oral monitoring module is used to call the file generation module when monitoring the oral evaluator's next complete answer, or, if it is monitored that the oral evaluator has no oral input within a first set time, the target scene large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator completes the answer, and calls the file generation module; or, if it is monitored that the oral evaluator has oral input but does not constitute a complete sentence within a second set time, the target scene large language model gives an answer prompt based on the previous oral question and answer until the oral evaluator completes the answer, and calls the file generation module, and the second set time is greater than the first set time; The oral assessment module is used to output the oral assessment score of the oral assessor for this oral question-and-answer assessment.

Citation Information

Patent Citations

  • Oral answer detection method, related device, equipment and storage medium

    CN118553246A

  • Real-time system for spoken natural stylistic conversations with large language models

    US20240169974A1