Dialogue methods, devices, computer equipment, and storage media for speech training.
The dialogue method and system enhance speech training flexibility and efficiency by using a large-scale language model to generate and refine questions and answers, addressing the limitations of fixed scenes and scripts in existing models.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- YOUMU TECHNOLOGY JAPAN CO LTD
- Filing Date
- 2025-06-20
- Publication Date
- 2026-06-02
AI Technical Summary
Existing speech training models for corporate customer service representatives and sales representatives lack flexibility due to fixed scenes and scripts, leading to increased workload and cost in model training, and reduced training effectiveness.
A dialogue method and system that utilizes a large-scale language model to generate and refine questions and answers based on user interaction, independent of the engine side data, reducing data volume while enhancing scene flexibility and training efficiency.
Improves the flexibility and efficiency of dialogue scene construction, reduces data requirements, and enhances the effectiveness of speech training by utilizing semantic analysis and feedback mechanisms.
Smart Images

Figure 0007869380000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and specifically, to a dialogue method, apparatus, computer device, and storage medium for speech training.
Background Art
[0002] When performing speech training for corporate customer service representatives, sales representatives, etc., in order to improve training efficiency and reduce labor costs, it is common to use a training model trained based on artificial intelligence for speech training. That is, the training model presents questions, the user answers the questions presented by the training model, the training model evaluates the user's answer and provides an evaluation result, and thereby the user improves the current speech based on the evaluation result to achieve the purpose of speech training. Since the scenes and scripts used for training are fixed, the flexibility of the entire trained training model is insufficient, and the effect of speech training is reduced. In order to improve the flexibility of the training model, it is necessary to provide more scenes and scripts for speech training so that the trained training model can handle various dialogue scenes. As a result, the workload of model training in the initial stage increases significantly, leading to an increase in the cost of the training model.
[0003] Therefore, how to reduce the data volume of scenes and scripts for speech training while improving the construction efficiency of dialogue scenes and the training effect of dialogue speech has become an urgent technical problem to be solved.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Based on the above situation, the main objective of the present invention is to provide a dialogue method, apparatus, computer equipment, and storage medium for speech training that improve the effectiveness of dialogue training by reducing the amount of data for scenes and scripts for speech training while improving the efficiency of constructing dialogue scenes. [Means for solving the problem]
[0005] To achieve the above objective, the present invention employs the following technical solutions. In a first embodiment, the embodiment of the present invention is applied to the engine side, where the engine side, user side, and large language model side constitute a speech training system, the engine side exchanges data with the user side and the large language model side respectively, the user side and the large language model side do not exchange data, and the data on the large language model side is independent of the engine side. The aforementioned dialogue method is Step S100 involves screening a dialogue corpus for question-answer sets that fit a dialogue scene selected by a user trigger, wherein the dialogue corpus includes question-answer sets for multiple types of dialogue scenes, and each type of question-answer set includes at least one presented question and its corresponding standard answer. Step S200 involves sending the target question from the target question-answer set to the large-scale language model, and the large-scale language model generating a first question sentence based on the target question. Step S300 includes receiving the first question sentence generated by the large-scale language model and sending the first question sentence to the user, A step of obtaining a first response text sent from the user, the first response text being generated by answering the first question text in step S400, Step S500 involves sending the standard answer and the first response sentence to the target question to the large-scale language model, and the large-scale language model determining whether the meaning of the first response sentence and the standard answer to the target question match. Step S610: If the feedback result from the large-scale language model shows that the meaning of the first answer sentence matches that of the standard answer to the target question, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question. The present invention discloses a dialogue method for speech training, which includes the step of sending feedback information to the large-scale language model if the feedback result from the large-scale language model indicates that the meaning of the first answer sentence does not match that of the standard answer to the target question, and the large-scale language model generates a feedback sentence for the target question, wherein the feedback sentence for the target question is used to present a revised answer to the target question to the user, S620.
[0006] The method is selectable, and before step S100, Step S110 to obtain dialogue materials, Step S120 further includes a step of analyzing dialogue data to obtain a dialogue corpus, wherein the dialogue corpus includes multiple types of question-answer sets, each type of question-answer set corresponds to one dialogue scene, and each type of question-answer set includes at least one presented question and its corresponding standard answer.
[0007] Selectable, after step S620, the method Step S621 further includes the step of sending the next question to the large-scale language model when the number of times the user has answered the target question exceeds a predetermined threshold, and the large-scale language model generates the next question sentence based on the next question, wherein the number of times the target question has been answered is the number of times the large-scale language model generates a feedback sentence for the target question.
[0008] The feedback sentences for the target question are selectable, include at least one, the feedback sentences are different from the first question sentence, and at least one feedback sentence for each target question is different from the other.
[0009] The method is selectable, Step S700 further includes a step in which the dialogue is terminated when the number of questions answered by the user meets a pre-set condition, wherein the pre-set condition is that the number of questions answered by the user reaches a pre-set ratio to the number of questions presented in the target question-answer set.
[0010] Selectable, prior to step S200, the method is Step S130 involves sending a dialogue start request to the large-scale language model based on a dialogue scene selected by a user trigger, and the large-scale language model generating a dialogue start sentence based on the dialogue start request. The process includes step S140, which involves sending the received dialogue start message to the user and receiving the reply message sent by the user.
[0011] Selectable, Step S700 is, Step S710: If the number of questions answered by the user meets a pre-set condition, a dialogue termination request is sent to the large-scale language model, and the large-scale language model generates a dialogue termination statement based on the dialogue termination request. The process includes step S720, which involves receiving a dialogue termination message and sending the dialogue termination message to the user to terminate the dialogue.
[0012] In a second embodiment, the embodiment of the present invention is applied to the engine side, where the engine side, user side, and large language model side constitute a speech training system, the engine side exchanges data with the user side and the large language model side respectively, the user side and the large language model side do not exchange data, and the data on the large language model side is independent of the engine side, and the dialogue device is A module that screens a dialogue corpus for question-answer sets that fit a dialogue scene selected by a user trigger, and selects a target question-answer set, wherein the dialogue corpus includes question-answer sets for multiple types of dialogue scenes, and each type of question-answer set includes a question-answer set selection module that includes at least one presented question and its corresponding standard answer, A question generation module that sends the target question from the target question-answer set to the large-scale language model, and the large-scale language model generates the first question sentence based on the target question, A question transmission module that receives the first question generated by the large-scale language model and sends the first question to the user, A module for receiving the first response text sent from the user, the first response text being generated by answering the first question text, A sentence response transmission module that sends a standard answer and a first response sentence to a target question, and the large-scale language model determines whether the meaning of the first response sentence matches that of the standard answer to the target question. If the feedback from the large-scale language model indicates that the meaning of the first answer sentence matches that of the standard answer to the target question, the next question is sent to the large-scale language model, and the next question generation module generates the next question sentence based on the next question. The present invention discloses a dialogue device for speech training, which includes a module that sends feedback information to the large-scale language model when the feedback result from the large-scale language model indicates that the meaning of the first answer sentence does not match that of the standard answer to the target question, and the large-scale language model generates a feedback sentence for the target question, and the feedback sentence for the target question is used to present a revised answer to the target question to the user.
[0013] In a third embodiment, an embodiment of the present invention discloses a computer device characterized by employing the speech training dialogue method described in the first embodiment or the apparatus described in the second embodiment. In a fourth embodiment, an embodiment of the present invention discloses a computer-readable storage medium that stores a computer program, characterized in that a processor executes the computer program stored in the storage medium to realize the method described in the first embodiment. [Effects of the Invention]
[0014] The speech training dialogue method, apparatus, computer equipment, and storage medium according to embodiments of the present invention are applied to the engine side, and the engine side, user side, and large language model side constitute a speech training system. The engine side exchanges data with the user side and the large language model side, respectively, but the user side and the large language model side do not exchange data, and the data on the large language model side is independent of the engine side. Based on the dialogue scene selected by the user side trigger, a set of questions and answers is screened to become a target question and answer set, and the target questions in the target question and answer set are sent to the large language model side, the large language model side generates a first question sentence, receives the first question sentence and sends it to the user side, and then receives the first answer sentence returned from the user side and sends both the first answer sentence and the standard answer to the large language model side. The large-scale language model determines whether the meaning of the first response sentence matches that of the standard response. If the meanings match, it sends the next question to the large-scale language model to generate the next question sentence. If the meanings do not match, it sends feedback information to the large-scale language model, which generates a feedback sentence and presents the user with a revised answer to the question. By exchanging data between the user and the engine, the engine can perceive the entire dialogue process throughout the entire dialogue process for speech skills training. Furthermore, by separating the engine and the large-scale language model, the amount of data for speech skills training can be reduced. In addition, by utilizing the semantic analysis and other functions of the large-scale language model, the flexibility of the generated dialogue scenes and scripts can be improved, the efficiency of constructing dialogue scenes can be increased, and the effectiveness of speech skills training can be improved.
[0015] Other beneficial effects of the present invention are as follows. These will be described in specific embodiments through the description of specific technical features and technical solutions. Those skilled in the art will be able to understand the beneficial technical effects brought about by these technical features and technical solutions through the description of these technical features and technical solutions. [Brief explanation of the drawing]
[0016] Hereinafter, preferred embodiments of the dialogue method, apparatus, computer device, and storage medium for speech training of the present invention will be described while referring to the drawings. [Figure 1] It is a configuration schematic diagram of a dialogue system for speech training according to this embodiment. [Figure 2] It is a flow schematic diagram of a dialogue method for speech training according to this embodiment. [Figure 3] It is a flow schematic diagram at the end of the dialogue according to this embodiment. [Figure 4] It is a flow schematic diagram at the start of the dialogue according to this embodiment. [Figure 5] It is a flow schematic diagram for obtaining a question-and-answer set according to this embodiment. [Figure 6] It is an exchange schematic diagram of three terminals in a dialogue system for speech training according to this embodiment. [Figure 7] It is a configuration schematic diagram of a dialogue apparatus for speech training according to this embodiment.
Embodiments for Implementing the Invention
[0017] Hereinafter, the present invention will be described based on embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail, but well-known methods, processes, flows, and components are not described in detail in order not to obscure the essence of the present invention.
[0018] Furthermore, those skilled in the art should understand that all the drawings provided in this specification are for illustrative purposes and are not necessarily drawn to scale.
[0019] Unless explicitly required by context, similar terms such as “includes” and “inclusive” in this specification and throughout the claims should be interpreted in an inclusive sense, i.e., “includes but not limited to,” rather than in an exclusive or exhaustive sense.
[0020] In this description of the present invention, terms such as "first," "second," etc., are used solely for illustrative purposes and should not be understood as indicating or implying relative importance. Furthermore, in this description of the present invention, unless otherwise specified, "plural" means two or more.
[0021] To improve the efficiency of constructing dialogue scenes and enhance the effectiveness of dialogue training while reducing the amount of data for scenes and scripts used in speech training, this embodiment discloses a dialogue method for speech training, as shown in Figure 1. Figure 1 is a schematic diagram of the configuration of the dialogue system for speech training according to this embodiment. This dialogue method for speech training is applied to the engine side, and the three terminals—the engine side, the user side, and the large-scale language model side—jointly constitute the speech training system. The engine side exchanges data with the user side and the large-scale language model side, respectively, but the user side and the large-scale language model side do not exchange data, and the data on the large-scale language model side is independent of the engine side.
[0022] Referring to Figure 2, Figure 2 is a schematic flowchart of the dialogue method for speech training according to this embodiment. As shown in Figure 2, the dialogue method for speech training includes the following steps. Step S100: Based on the dialogue scene selected by the user trigger, the dialogue corpus is screened for question-answer sets that fit the dialogue scene, and these are designated as target question-answer sets. The dialogue corpus includes question-answer sets for multiple types of dialogue scenes, and each type of question-answer set includes at least one presented question and its corresponding standard answer. In this embodiment, the dialogue corpus includes question-answer sets for various dialogue scenes, and each question-answer set for each scene includes at least one presented question and its corresponding standard answer. Dialogue scenes can refer to different dialogue topics, such as drug recommendation scenes, insurance product recommendation scenes, and retail product recommendation scenes. Since multiple types of dialogue scenes exist, the engine screens the dialogue corpus for question-answer sets that fit the dialogue scene specified by the user, and all subsequent speech training dialogues are conducted based on these screened target question-answer sets. In the specific implementation process, the dialogue corpus may be provided by the training executor, who is the company that uses the speech training dialogue method.
[0023] Step S200: The target question from the target question-answer set is sent to the large-scale language model, and the large-scale language model generates a first question sentence based on the target question. In this embodiment, the engine can select a target question from the target question-answer set according to a specific question selection strategy, and then send the target question to the large-scale language model, which can generate a first question sentence based on the target question. By utilizing the large-scale language model on the large-scale language model side, the large-scale language model can generate a first question sentence that is closer to the question-answer habits and question-answer style of a natural person based on the target question, making the entire dialogue process more natural and realistic. In the specific implementation process, the engine may, as a question selection strategy for selecting a question from the target question-answer set, determine the target question according to the order of the questions in the target question-answer set, randomly extract questions from the target question-answer set and use them as the target question, or select questions from easy to difficult according to the difficulty level of the questions in the target question-answer set.
[0024] Step S300: The large-scale language model receives the first question sentence generated by the large-scale language model and sends the first question sentence to the user. In this embodiment, after the large-scale language model generates the first question sentence, it sends the first question sentence to the engine, which receives the first question sentence generated by the large-scale language model and sends the first question sentence to the user, thereby enabling the user to answer based on the first question sentence sent from the engine.
[0025] Step S400: The first response text sent from the user is obtained, and the first response text is generated by answering the first question text. In this embodiment, the user answers based on the first question text and generates the first response text, the user sends the first response text to the engine side, the engine side receives the first response text sent from the user side and proceeds to the next dialogue based on the first response text.
[0026] Step S500: The standard answer to the target question and the first answer sentence are sent to the large-scale language model, which determines whether the meaning of the first answer sentence matches that of the standard answer to the target question. In this embodiment, the engine sends both the standard answer to the target question and the first answer sentence sent by the user to the large-scale language model. The large-scale language model then performs semantic analysis on the first answer sentence and the standard answer and determines whether the meaning of the first answer sentence matches that of the standard answer to the target question. If the meaning of the first answer sentence matches that of the standard answer to the target question, it indicates that the user's answer to the target question is correct, and the system can proceed to the next question and answer, or end the current dialogue. If the meaning of the first answer sentence does not match that of the standard answer to the target question, it indicates that the user's answer to the target question is incorrect, and the system needs to repeat the current question and answer, proceed to the next question and answer, or end the current dialogue.
[0027] Step S610: If the feedback result from the large-scale language model shows that the meaning of the first answer sentence matches that of the standard answer to the target question, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question. In this embodiment, the large-scale language model sends the feedback result of semantic analysis to the engine. If the feedback result from the large-scale language model shows that the meaning of the first answer sentence matches that of the standard answer to the target question, the target question is considered to have been answered correctly, and the next question can be asked. At this time, the engine selects the next question from the target question / answer set, sends the selected next question to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question, and steps S200 to S500 are repeated.
[0028] In selectable embodiments, the engine can send evaluation feedback to the large-scale language model, which then generates evaluation results based on that feedback. In this embodiment, this feedback result may be the merits and demerits of the first answer sentence determined by the large-scale language model through methods such as semantic analysis. The engine can receive the evaluation results, send them to the user, and then select the next question from the target question-answer set to execute step S610.
[0029] Step S620: If the feedback result from the large-scale language model does not show a semantic match between the first answer and the standard answer to the target question, feedback information is sent to the large-scale language model, which generates a feedback sentence for the target question. The feedback sentence for the target question is used to present the user with a revised answer to the target question. In this embodiment, the large-scale language model sends the semantic analysis feedback result to the engine, and if the feedback result from the large-scale language model does not show a semantic match between the first answer and the standard answer to the target question, the target question is considered not answered correctly, and the user needs to re-answer the target question. At this time, the engine sends feedback information to the large-scale language model, which allows the large-scale language model to continue generating a feedback sentence for the target question based on the target question. The feedback sentence is used to present the user with a revised answer to the target question, and steps S200 to S500 are repeated.
[0030] In selectable embodiments, the dialogue method for speech training further includes the following steps: Step S621: If the number of times the user answers a target question exceeds a preset threshold, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question. The number of times a target question is answered is the number of times the large-scale language model generates a feedback sentence for the target question. In this embodiment, if the number of times the user answers a target question exceeds a preset threshold, the user is deemed not to understand the target question, and in this case, the target question can be skipped and the system can proceed to the next question for the user. The preset threshold is set in advance and can be, for example, 2 times. If the large-scale language model generates two feedback sentences for the same target question, it indicates that three questions (including one first question sentence and two feedback sentences) have been asked for the target question. If the user does not return an answer that matches the standard answer even after multiple questions, the question can be temporarily skipped and the system can proceed directly to the next question, thereby preventing the dialogue from stalling at the same stage.
[0031] In the selectable embodiments, the target question includes at least one feedback sentence, the feedback sentences are different from the first question sentence, and at least one feedback sentence for each target question is different from each other. In this embodiment, when asking the same target question, the user may need to answer the question multiple times to answer it correctly, so there may be multiple feedback sentences generated for the target question. If there is at least one feedback sentence for the target question, each feedback sentence is different from the others, and each feedback sentence is different from the first question sentence. Different sentences refer to different question formats or question expressions, but it is understandable that they are essentially asking the same target question.
[0032] In selectable embodiments, the dialogue method for speech training further includes the following steps: Step S700: The dialogue ends when the number of questions answered by the user meets a pre-set condition. The pre-set condition is that the number of questions answered by the user reaches a pre-set percentage of the questions presented in the target question-answer set. In this embodiment, the pre-set percentage may be 100% or 90%. For example, if the number of questions presented in the target question-answer set is 5 and the pre-set percentage is 100%, the dialogue ends when the number of questions answered by the user reaches 5. Note that the current dialogue can be ended when the number of questions answered by the user meets the pre-set condition, and the user is not required to answer all questions correctly.
[0033] In one of the selectable embodiments, referring to Figure 3, which is a schematic flowchart of the interaction completion in this embodiment. As shown in Figure 3, step S700 includes steps S710 and S720. Step S710: When the number of questions answered by the user meets a pre-set condition, a dialogue termination request is sent to the large-scale language model, and the large-scale language model generates a dialogue termination statement based on the dialogue termination request. In this embodiment, the pre-set condition is that the number of questions answered by the user reaches a pre-set ratio to the number of questions presented in the target question-answer set. When the number of questions answered by the user meets the pre-set condition, the dialogue can be terminated. At this time, the engine sends a dialogue termination request to the large-scale language model, and the large-scale language model can generate a dialogue termination statement based on the dialogue termination request.
[0034] Step S720: The dialogue termination message is received and sent to the user to terminate the dialogue. In this embodiment, the engine receives the dialogue termination message sent from the large-scale language model and sends the dialogue termination message to the user to terminate the current dialogue.
[0035] In one of the selectable embodiments, referring to Figure 4, which is a schematic flowchart of the dialogue initiation according to this embodiment. As shown in Figure 4, prior to step S200, the dialogue method for speech training includes the following steps. Step S130: Based on the dialogue scene selected by the user trigger, the engine sends a dialogue start request to the large language model, and the large language model generates a dialogue start sentence based on the dialogue start request. In this embodiment, after the user triggers the selection of a dialogue scene, the engine can send a dialogue start request to the large language model based on the dialogue scene selected by the user trigger, and the large language model generates a dialogue start sentence, i.e., a greeting, based on the dialogue start request.
[0036] Step S140: The received dialogue start sentence is sent to the user side, and the reply sentence sent from the user side is received. In this embodiment, the engine side sends the dialogue start sentence generated by the received large-scale language model side to the user side, and receives the reply sentence sent from the user side, and then starts a formal question session.
[0037] In an optional embodiment, referring to Figure 5, which is a schematic flowchart for obtaining a question-answer set according to this embodiment. As shown in Figure 5, prior to step S100, the conversational method for speech training further includes the following steps. Step S110: Obtain the dialogue material. In this embodiment, the dialogue material is a knowledge base document provided by the company itself, and may be text material, image material, or audio material. There are no restrictions on the format and type of the dialogue material; it may be in doc. format, ppt. format, or various other formats such as excel, pdf, html, video, etc.
[0038] Step S120: Analyze the dialogue data to obtain a dialogue corpus. The dialogue corpus contains multiple types of question-answer sets, each corresponding to one dialogue scene, and each set contains at least one presented question and its corresponding standard answer. The engine analyzes the dialogue data and extracts the dialogue corpus from it. The dialogue corpus contains multiple types of question-answer sets, each corresponding to one dialogue scene, and the dialogue scenes may be on different dialogue topics. For example, dialogue scenes could be about cold medicine, immunological diseases, cardiovascular diseases, etc. In the specific implementation process, knowledge base documents are divided into blocks and saved as vectors, allowing for the identification of high-frequency and important words. Of course, it is understandable that important words can also be manually marked. Next, a vector matching algorithm is used to match relevant knowledge, and knowledge points are extracted and summarized via LLM to generate question-answers for glossary definitions. The system repeatedly searches for high-frequency knowledge points and questions within the text, creates a list of questions, finds knowledge related to the answers to those questions, summarizes and extracts core answers via LLM, extracts question-answer pairs, and obtains question-answer sets. Dialogue materials can be used to automatically generate question-answer sets, reducing the amount of data required to build scenes and scripts for public speaking training.
[0039] Referring to Figure 6, which is a schematic diagram of the exchange of three terminals in the speech training dialogue system according to this embodiment, the solution will be explained below with reference to Figure 6 to facilitate understanding of this solution.
[0040] As shown in Figure 6, the engine can send a dialogue start request to the large-scale language model based on the dialogue scene selected by the user trigger. Dialogue scenes include multiple types, such as introducing a biological drug product to a doctor, introducing a critical illness insurance product to a customer, or introducing an expensive classic handbag to a customer. The dialogue scene selected by the user trigger might be, for example, introducing a biological drug product to a doctor. The engine can receive touch selection actions or click selection actions performed by the user on the dialogue scene via the user interface, determine the dialogue scene selected by the user trigger, and send a dialogue start request to the large-scale language model based on that dialogue scene.
[0041] The large-scale language model generates a dialogue initiation sentence, i.e., a greeting, based on the dialogue initiation request, and sends the greeting to the engine. The engine sends the greeting to the user and receives the user's response to the greeting, the reply sent by the user.
[0042] The engine screens the dialogue corpus for question-and-answer sets that fit the dialogue scene selected by the user trigger, selects a target question from the question-and-answer set, and sends the target question to the large-scale language model. An example of a target question would be "mechanism of action of a novel biological agent for systemic lupus erythematosus."
[0043] The large-scale language model generates a first question sentence based on the target question. The large-scale language model uses the language model to generate an interactive question sentence based on the target question and sends the generated first question sentence to the engine. For example, a first question sentence might be, "Could you explain the mechanism of action of this biological drug?"
[0044] The engine receives the first question generated by the large-scale language model and sends the first question to the user. The user answers based on the first question, obtains the first answer, and sends the first answer to the engine.
[0045] After the engine receives the first response, it sends both the standard answer to the target question and the first response generated by the user based on the first question to the large-scale language model. After the large-scale language model receives the standard answer and the first response, it performs semantic analysis on the standard answer and the first response to determine whether the meaning of the first response matches the meaning of the standard answer, and sends the result to the engine.
[0046] If the engine determines that the meaning of the first answer matches the meaning of the standard answer, it can select the next question from the question-answer set and send the selected next question to the large-scale language model. The large-scale language model then generates the next question based on the first question. The user can then answer the next question based on the next question. The answer is sent by the engine to the large-scale language model until the user correctly answers the next question or the number of times the user answers the next question exceeds a preset threshold, and semantic analysis is performed again. At this point, the engine can decide whether to end the dialogue or ask the next question based on the number of questions the user has answered.
[0047] If the engine determines that the meaning of the first response does not match the meaning of the standard response, and the number of times the user has answered the question has not exceeded a predetermined threshold, the engine may request the user to continue supplementing their answer until they correctly answer the question or until the number of times the user has answered the question exceeds a predetermined threshold. Specifically, the engine sends feedback information to the large-scale language model, instructing the model to only provide feedback on the user's answers without asking new questions. After receiving the feedback information, the large-scale language model generates a feedback sentence for the question and feeds that feedback sentence back to the engine. For example, a possible feedback sentence might be, "Thank you for providing the information. However, could you please tell us more about the specific mechanism of action of this biological agent, particularly its mechanism of action within the immune system?" This feedback sentence is an additional question to the question, "Mechanism of action of the novel biological agent for systemic lupus erythematosus." The user can continue answering the question based on the feedback sentence. The response text is sent by the engine to the large-scale language model until the user correctly answers the question or the number of times the user has answered the question exceeds a predetermined threshold, at which point semantic analysis is performed again. At this point, based on the number of questions the user has answered so far, it can be decided whether to end the conversation or ask the next question.
[0048] If the number of questions answered by the user meets a pre-set condition, for example, if all questions in the dialogue scene have been answered, the engine can choose to end the dialogue. At this time, the engine can send a dialogue termination request to the large-scale language model, which generates a dialogue termination statement based on the request and sends it to the engine, and the engine ends the dialogue by sending this dialogue termination statement to the user.
[0049] Referring to Figure 7, Figure 7 is a schematic diagram of the configuration of a speech training dialogue device according to this embodiment. This speech training dialogue device is applied to the engine side, and the engine side, user side, and large language model side constitute a speech training system. The engine side exchanges data with the user side and the large language model side, respectively, but the user side and the large language model side do not exchange data, and the data on the large language model side is independent of the engine side. As shown in Figure 7, this speech training dialogue device includes the following modules. The question / answer set selection module 100 screens the dialogue corpus for question / answer sets that fit the dialogue scene based on the dialogue scene selected by the user trigger, and selects them as target question / answer sets. The dialogue corpus includes question / answer sets for multiple types of dialogue scenes, and each type of question / answer set includes at least one presented question and its corresponding standard answer.
[0050] Question generation module 200: The target question in the target question-answer set is sent to the large-scale language model, and the large-scale language model generates the first question sentence based on the target question.
[0051] Question transmission module 300: Receives the first question generated by the large-scale language model and sends the first question to the user.
[0052] Response message receiving module 400: Receives the first response message sent from the user, and the first response message is generated by answering the first question message.
[0053] Text Answer Transmission Module 500: Sends the standard answer and the first answer sentence to the target question to the large-scale language model, which then determines whether the meaning of the first answer sentence matches that of the standard answer to the target question.
[0054] The next question generation module 610: If the feedback result from the large-scale language model shows that the meaning of the first answer sentence matches that of the standard answer to the target question, it sends the next question to the large-scale language model, which then generates the next question sentence based on that question.
[0055] Feedback sentence generation module 620: If the feedback result from the large-scale language model does not match the meaning of the first answer sentence and the standard answer to the target question, feedback information is sent to the large-scale language model, which generates a feedback sentence for the target question. This feedback sentence is then used to present the user with a revised answer to the target question.
[0056] The speech training dialogue method, apparatus, computer equipment, and storage medium according to embodiments of the present invention are applied to the engine side, and the engine side, user side, and large language model side constitute a speech training system. The engine side exchanges data with the user side and the large language model side, respectively, but the user side and the large language model side do not exchange data, and the data on the large language model side is independent of the engine side. Based on the dialogue scene selected by the user side trigger, a set of questions and answers is screened to become a target question and answer set, and the target questions in the target question and answer set are sent to the large language model side, the large language model side generates a first question sentence, receives the first question sentence and sends it to the user side, and then receives the first answer sentence returned from the user side and sends both the first answer sentence and the standard answer to the large language model side. The large-scale language model determines whether the meaning of the first response sentence matches that of the standard response. If the meanings match, it sends the next question to the large-scale language model to generate the next question sentence. If the meanings do not match, it sends feedback information to the large-scale language model, which generates a feedback sentence and presents the user with a revised answer to the question. By exchanging data between the user and the engine, the engine can perceive the entire dialogue process throughout the entire dialogue process for speech skills training. Furthermore, by separating the engine and the large-scale language model, the amount of data for speech skills training can be reduced. In addition, by utilizing the semantic analysis and other functions of the large-scale language model, the flexibility of the generated dialogue scenes and scripts can be improved, the efficiency of constructing dialogue scenes can be increased, and the effectiveness of speech skills training can be improved.
[0057] Those skilled in the art will understand that, insofar as they do not contradict each other, the preferred solutions described above can be freely combined or superimposed. Here, the flowcharts and block diagrams in the drawings illustrate possible implementation architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, program segment, or part of code, which includes one or more executable instructions for implementing a given logical function. It should also be noted that in some alternative implementations, the functions shown in a block may be executed in an order different from the order shown in the drawings. For example, two blocks shown consecutively may actually be executed almost in parallel, or in reverse order, depending on the functions involved. Furthermore, each block in a block diagram and / or flowchart, and combinations thereof, may be implemented by a dedicated hardware-based system that performs a given function or operation, or by a combination of dedicated hardware and computer instructions. The numbering of each step in this specification is for illustrative and reference purposes only and does not limit the order. The specific execution order is determined by the technology itself, and a person skilled in the art can determine various acceptable and reasonable orders based on the technology itself.
[0058] Those skilled in the art will understand that, as long as they do not contradict each other, the above preferred solutions can be freely combined or superimposed.
[0059] It should be understood that the embodiments described above are merely illustrative and not limiting. Various obvious or equivalent modifications or substitutions made by those skilled in the art to the above details without departing from the basic principles of the present invention are included within the scope of the claims of the present invention.
Claims
1. A dialogue method for training speaking skills, Applied to the engine side, the engine side, the user side, and the large-scale language model side constitute a speech training system, the engine side exchanges data with the user side and the large-scale language model side respectively, the user side and the large-scale language model side do not exchange data, and the data of the large-scale language model side is independent of the engine side. The aforementioned dialogue method is Step S100 is a step in which a set of questions and answers suitable for a given dialogue scene is screened from a dialogue corpus based on a dialogue scene selected by the user-side trigger, and the set of questions and answers suitable for that dialogue scene is selected, wherein the dialogue corpus includes question and answer sets for multiple types of dialogue scenes, and each type of question and answer set includes at least one presented question and a corresponding standard answer. Step S200 includes sending the target question from the target question-answer set and a prompt to generate a first question sentence that is closer to the question-answering habits and question-answering style of a natural person based on the target question to the large-scale language model, and the large-scale language model generating the first question sentence based on the target question, Step S300 includes receiving the first question sentence generated by the large-scale language model and sending the first question sentence to the user, Step S400 of obtaining a first response sentence sent from the user, wherein the first response sentence is generated by answering the first question sentence, Step S500: The standard answer to the target question and the first answer sentence are sent to the large-scale language model, and the large-scale language model determines whether the meaning of the first answer sentence and the standard answer to the target question are consistent. Step S610: If the feedback result from the large-scale language model indicates that the meaning of the first answer sentence matches that of the standard answer to the target question, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question. If the feedback result from the large-scale language model indicates that the meaning of the first response sentence and the standard answer to the target question do not match, the large-scale language model should: In step S500, the large-scale language model determines whether the meaning of the first response sentence and the standard answer to the target question are consistent, and A prompt to generate a feedback statement for the aforementioned question, A dialogue method for speech training, characterized by comprising the step of sending a message, the large-scale language model generating a feedback sentence for the target question, the feedback sentence for the target question being used to present a revised answer to the target question to the user, and step S620.
2. Prior to step S100, the method is Step S110 to obtain dialogue materials, The dialogue method for speech training according to claim 1, further comprising step S120, which is a step of analyzing the aforementioned dialogue material to obtain a dialogue corpus, wherein the dialogue corpus includes multiple types of question-answer sets, each type of question-answer set corresponds to one dialogue scene, and each type of question-answer set includes at least one presented question and a corresponding standard answer.
3. After step S620, the method is performed as follows: The dialogue method for speech training according to claim 1, further comprising step S621, wherein if the number of times the user answers the target question exceeds a preset threshold, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question, the number of times the target question is answered is the number of times the large-scale language model generates a feedback sentence for the target question.
4. The aforementioned method, The dialogue method for speech skills training according to claim 1, further comprising step S700, which is a step to terminate the dialogue when the number of questions answered by the user satisfies a predetermined condition, wherein the predetermined condition is that the number of questions answered by the user reaches a predetermined ratio to the questions presented in the target question-answer set.
5. Before step S200, the method described above is Step S130 involves sending a dialogue start request to the large-scale language model based on the dialogue scene selected by the user-side trigger, and the large-scale language model generating a dialogue start sentence based on the dialogue start request. The dialogue method for speech skills training according to claim 1, characterized by including step S140 of sending the received dialogue start sentence to the user and receiving a reply sentence sent from the user.
6. The aforementioned step S700 is, Step S710: If the number of questions answered by the user meets a pre-set condition, a dialogue termination request is sent to the large-scale language model, and the large-scale language model generates a dialogue termination statement based on the dialogue termination request. The dialogue method for speech skills training according to claim 4, characterized by including step S720 of receiving the dialogue termination message and sending the dialogue termination message to the user to terminate the dialogue.
7. A dialogue device for training speaking skills, Applied to the engine side, the engine side, the user side, and the large-scale language model side constitute a speech training system, the engine side exchanges data with the user side and the large-scale language model side respectively, the user side and the large-scale language model side do not exchange data, and the data of the large-scale language model side is independent of the engine side. The dialogue device is A module that screens a set of questions and answers from a dialogue corpus to select a target question and answer set based on a dialogue scene selected by the user trigger, wherein the dialogue corpus includes question and answer sets for multiple types of dialogue scenes, and each type of question and answer set includes at least one presented question and a corresponding standard answer, comprising a question and answer set selection module (100), A question generation module (200) transmits a target question from the aforementioned target question-answer set and a prompt to generate a first question sentence that is closer to the question-answering habits and question-answering style of a natural person based on the target question to the large-scale language model, and the large-scale language model generates the first question sentence based on the target question. A question transmission module (300) that receives the first question generated by the large-scale language model and transmits the first question to the user, A module for acquiring a first response sent from the user, comprising a response receiving module (400) which is generated when the first response is an answer to the first question, A sentence answer transmission module (500) transmits the standard answer to the target question and the first answer sentence to the large-scale language model, and the large-scale language model determines whether the meaning of the first answer sentence and the standard answer to the target question are consistent. If the feedback result from the large-scale language model indicates that the meaning of the first answer sentence and the standard answer to the target question match, the next question is sent to the large-scale language model, and the large-scale language model generates the next question sentence based on the next question, and the next question generation module (610) If the feedback result from the large-scale language model indicates that the meaning of the first response sentence and the standard answer to the target question do not match, the large-scale language model should: In the aforementioned text response transmission module, the large-scale language model determines whether the meaning of the first response text and the standard answer to the target question are consistent, and A prompt to generate a feedback statement for the aforementioned question, A dialogue device for speech training, characterized in that it transmits a message, and the large-scale language model side generates a feedback message for the target question, the feedback message for the target question includes a feedback message generation module (620) used to present a revised answer to the target question to the user side.
8. Computer equipment, A computer device characterized by performing the dialogue method for speech training described in any one of claims 1 to 6.
9. A computer-readable storage medium that stores computer programs, A computer-readable storage medium characterized in that a processor executes a computer program stored in the storage medium to realize the method according to any one of claims 1 to 6.