A dialogue method, device, computer equipment and storage medium for speech training
By introducing semantic analysis functions at the large language model end in the speech training system, screening and evaluating user's answers, the problem of inflexible training models in the existing technology is solved, and more efficient and flexible speech training effects are achieved.
Patent Information
- Application Number
- CN202510081123.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
In the prior art, the training model used for speech training is not flexible overall, resulting in poor speech training results and increasing the workload and cost of model training.
By interacting data with the user and the large language model on the engine side, a collection of Q&A that matches the dialogue scenes selected by the user is selected is selected, and the semantic analysis function of the large language model side is used to generate and evaluate user's answers, thereby improving the construction efficiency of dialogue scenarios and the flexibility of speech training.
The data volume of speech training scenarios and scripts is simplified, the efficiency of dialogue scenarios is improved, the flexibility of training models is enhanced, and the effect of speech training is improved.
Smart Images

Figure CN119494409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a dialogue method, device, computer equipment and storage medium for speech training. Background Art
[0002] When training customer service staff, sales staff, etc., in order to improve training efficiency and reduce labor costs, training models based on artificial intelligence training are usually used for speech training. That is to say, questions are asked using the training model, and users answer the questions asked by the training model. The training model evaluates the user's answers and gives evaluation results, so that users can improve the current speech based on the evaluation results to achieve the purpose of speech training. However, due to the fixed scenes and scripts used for training, the training model obtained by training is not flexible enough as a whole, resulting in poor results in speech training. In order to improve the flexibility of the training model, more speech training scenes and scripts need to be provided so that the training model can adapt to different dialogue scenes. This greatly increases the workload of model training in the early stage, which in turn leads to an increase in the cost of training the model.
[0003] Therefore, how to improve the efficiency of dialogue scene construction while simplifying the data volume of dialogue training scenarios and scripts, and thus improve the training effect of dialogue speech has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] Based on the above situation, the main purpose of the present invention is to provide a dialogue method, device, computer equipment and storage medium for speech training, so as to improve the efficiency of constructing dialogue scenes while simplifying the data volume of speech training scenes and scripts, thereby improving the training effect of dialogue speech.
[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, an embodiment of the present invention discloses a dialogue method for speech training, which is applied to an engine end. The engine end, the user end and the large language model end constitute a speech training system. The engine end interacts with the user end and the large language model end respectively for data, the user end and the large language model end do not interact with data, and the data of the large language model end is independent of the engine end. The dialogue method includes:
[0007] Step S100, based on the dialogue scenario triggered and selected by the user end, a question and answer collection matching the dialogue scenario is selected from the dialogue material as the target question and answer collection, the dialogue material includes question and answer collections under several types of dialogue scenarios, and each type of question and answer collection includes at least one question and a corresponding standard answer;
[0008] Step S200, sending a target question in the target question and answer collection to the large language model end, so that the large language model end generates a first question sentence based on the target question;
[0009] Step S300, receiving a first question sentence generated by the large language model end, and sending the first question sentence to the user end;
[0010] Step S400, obtaining a first answer statement sent by the user terminal, where the first answer statement is generated in response to the first question statement;
[0011] Step S500, sending the standard answer to the target question and the first answer sentence to the large language model end, so that the large language model end determines whether the first answer sentence is semantically consistent with the standard answer to the target question;
[0012] Step S610: When the feedback result obtained from the large language model is that the first answer statement is semantically consistent with the standard answer to the target question, the next question is sent to the large language model so that the large language model generates a next question statement based on the next question.
[0013] Step S620: When the feedback result obtained from the large language model end is that the first answer statement is semantically inconsistent with the standard answer to the target question, feedback information is sent to the large language model end so that the large language model end generates a feedback statement for the target question. The feedback statement for the target question is used to prompt the user end to re-answer the target question.
[0014] Optionally, before step S100, the method further includes:
[0015] Step S110, obtaining conversation data;
[0016] Step S120, the dialogue data is parsed to obtain dialogue data, which includes several types of question and answer collections, each type of question and answer collection corresponds to a dialogue scene, and each type of question and answer collection includes at least one question and a corresponding standard answer.
[0017] Optionally, after step S620, the method further includes:
[0018] Step S621, when the number of answer rounds of the user end to the target question exceeds a preset threshold, the next question is sent to the large language model end, so that the large language model end generates the next question sentence based on the next question, wherein the number of answer rounds of the target question is the number of times the large language model end generates a feedback sentence for the target question.
[0019] Optionally, the feedback statement of the target question includes at least one, the feedback statement is different from the first question statement, and the feedback statements of at least one target question are different from each other.
[0020] Optionally, the method further comprises:
[0021] Step S700, when the number of questions answered by the user terminal meets a preset condition, the conversation is ended. The preset condition is that the number of questions answered by the user terminal reaches a preset proportion of the questions asked in the target question and answer collection.
[0022] Optionally, before step S200, the method includes:
[0023] Step S130, sending a dialog start request to the large language model end based on the dialog scenario triggered and selected by the user end, so that the large language model end generates a dialog start sentence based on the dialog start request;
[0024] Step S140, sending the received dialogue initiation statement to the user end, and receiving the reply statement sent by the user end.
[0025] Optionally, step S700 includes:
[0026] Step S710, when the number of questions answered by the user terminal meets the preset condition, a dialog end request is sent to the large language model terminal, so that the large language model terminal generates a dialog end statement based on the dialog end request;
[0027] Step S720: receiving a conversation ending statement, and sending the conversation ending statement to the user end to end the conversation.
[0028] In a second aspect, an embodiment of the present invention discloses a dialogue device for speech training, which is applied to an engine end. The engine end, the user end and the large language model end constitute a speech training system. The engine end interacts with the user end and the large language model end respectively, and the user end and the large language model end do not interact with each other, and the data of the large language model end is independent of the engine end. The dialogue device includes:
[0029] A question-answer collection selection module is used to select a question-answer collection matching the dialogue scene from the dialogue material based on the dialogue scene triggered and selected by the user end, as the target question-answer collection, the dialogue material includes question-answer collections under several types of dialogue scenes, and each type of question-answer collection includes at least one question and a corresponding standard answer;
[0030] A question statement generating module, configured to send a target question in a target question-answer collection to a large language model end, so that the large language model end generates a first question statement based on the target question;
[0031] A question sentence sending module, used for receiving a first question sentence generated by the large language model end, and sending the first question sentence to the user end;
[0032] An answer statement receiving module is used to obtain a first answer statement sent by the user end, where the first answer statement is generated to answer the first question statement;
[0033] A sentence answer sending module is used to send the standard answer to the target question and the first answer sentence to the large language model end, so that the large language model end determines whether the first answer sentence is semantically consistent with the standard answer to the target question;
[0034] A next question generation module is used to send the next question to the large language model end when the feedback result obtained from the large language model end is that the first answer sentence is semantically consistent with the standard answer to the target question, so that the large language model end generates the next question sentence based on the next question;
[0035] The feedback sentence generation module is used to send feedback information to the large language model end when the feedback result obtained from the large language model end is that the first answer sentence is semantically inconsistent with the standard answer to the target question, so that the large language model end generates a feedback sentence for the target question. The feedback sentence for the target question is used to prompt the user end to re-answer the target question.
[0036] In a third aspect, an embodiment of the present invention discloses a computer device, characterized in that it includes: a dialogue method for speech training as described in the first aspect; or, includes the apparatus as described in the second aspect.
[0037] In a fourth aspect, an embodiment of the present invention discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program stored in the storage medium is used to be executed by a processor to implement the method described in the first aspect. Beneficial Effects
[0038] According to the dialogue method, device, computer equipment and storage medium for speech training disclosed in the embodiment of the present invention, the engine end is applied, and the engine end, the user end and the large language model end constitute a speech training system. The engine end interacts with the user end and the large language model end respectively, and the user end and the large language model end do not interact with each other and the data of the large language model end is independent of the engine end. Based on the dialogue scene triggered by the user end, the question and answer collection is filtered as the target question and answer collection, and then the target question in the target question and answer collection is sent to the large language model end so that the large language model end generates a first question statement, receives the first question statement and sends the first question statement to the user end, and then receives the first answer statement returned by the user end, and sends the first answer statement and the standard answer to the large language model end together. The large language model end determines whether the semantics of the first answer statement and the standard answer are consistent. When the semantics of the two are consistent, the next question is sent to the large language model end to generate the next question statement. When the semantics of the two are inconsistent, feedback information is sent to the large language model end so that the large language model end generates a feedback statement to prompt the user to answer the target question again. The user side interacts with the engine side through data, so that the engine side can perceive the complete dialogue process during the entire dialogue process of speech training. In addition, separating the engine side from the large language model side can not only simplify the amount of data for speech training, but also utilize the semantic analysis and other functions of the large language model side to improve the flexibility of the generated dialogue scenes and scripts, improve the efficiency of constructing dialogue scenes, and thus improve the training effect of dialogue speech.
[0039] Other beneficial effects of the present invention will be explained in the specific implementation manner through the introduction of specific technical features and technical solutions. Through the introduction of these technical features and technical solutions, those skilled in the art should be able to understand the beneficial technical effects brought about by the technical features and technical solutions. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The following will describe a preferred embodiment of a conversation method, device, computer equipment and storage medium for speech training of the present invention with reference to the accompanying drawings.
[0041] Figure 1 A schematic diagram of the structure of a dialogue system for speech training disclosed in this embodiment;
[0042] Figure 2 A flowchart of a dialogue method for speech training disclosed in this embodiment;
[0043] Figure 3 This is a flow chart of the process of ending a conversation disclosed in this embodiment;
[0044] Figure 4 This is a flow chart of starting a conversation according to the present embodiment;
[0045] Figure 5A schematic diagram of the process of obtaining a question and answer collection disclosed in this embodiment;
[0046] Figure 6 A schematic diagram of the interaction between three terminals in the dialogue system for speech training disclosed in this embodiment;
[0047] Figure 7 A schematic diagram of the structure of a dialogue device for speech training disclosed in this embodiment. DETAILED DESCRIPTION
[0048] The present invention is described below based on embodiments, but the present invention is not limited to these embodiments. In the following detailed description of the present invention, some specific details are described in detail. In order to avoid confusing the essence of the present invention, known methods, processes, procedures, and components are not described in detail.
[0049] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.
[0050] Unless the context clearly requires otherwise, throughout the specification and claims, the words "include", "comprising" and similar words should be interpreted in an inclusive sense rather than an exclusive or exhaustive sense; that is, in the sense of "including but not limited to".
[0051] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0052] In order to improve the efficiency of dialogue scene construction while simplifying the data volume of dialogue training scenes and scripts, and thus improve the training effect of dialogue dialogue, this embodiment discloses a dialogue method for dialogue training, please refer to Figure 1 , Figure 1 This is a structural diagram of a dialogue system for speech training disclosed in this embodiment. The dialogue method for speech training is applied to the engine end, and the engine end, the user end and the large language model end together constitute the speech training system, wherein the engine end interacts with the user end and the large language model end respectively for data, the user end and the large language model end do not interact with data, and the data of the large language model end is independent of the engine end.
[0053] Please refer to Figure 2 , Figure 2 FIG. 1 is a flow chart of a dialogue method for speech training disclosed in this embodiment. Figure 2 As shown in , the dialogue methods for this speech training include:
[0054] Step S100, based on the dialogue scene triggered and selected by the user end, a question and answer collection matching the dialogue scene is selected from the dialogue material as the target question and answer collection. The dialogue material includes question and answer collections under several types of dialogue scenes, and each type of question and answer collection includes at least one question and a corresponding standard answer. In this embodiment, the dialogue material includes question and answer collections under different dialogue scenes, and the question and answer collection under each scene includes at least one question and a standard answer corresponding to the question. The dialogue scene can refer to different dialogue topics, for example, it can be a drug promotion scene, an insurance product promotion scene, a retail product promotion scene, etc. Since there are multiple dialogue scenes, the engine end selects a question and answer collection matching the dialogue scene from the dialogue material based on the dialogue scene started by the user end as the target question and answer collection, and the subsequent dialogues of the speech training are all based on the selected target question and answer collection. In the specific implementation process, the dialogue material can be provided by the training party, and the training party refers to the enterprise using the dialogue method of the speech training.
[0055] Step S200, the target question in the target question and answer collection is sent to the large language model end, so that the large language model end generates a first question statement based on the target question. In this embodiment, the engine end selects the target question from the target question and answer collection according to a certain topic selection strategy, and then sends the target question to the large language model end, so that the large language model end can generate a first question statement based on the target question, and uses the large language model in the large language model end to enable the large language model end to generate a first question statement based on the target question that is closer to the question and answer habits and question and answer methods of natural people, so that the entire dialogue process is more natural and real. In the specific implementation process, the topic selection strategy of the engine end from the target question and answer collection can be to determine the target question according to the arrangement order of the questions in the target question and answer collection, or to randomly extract questions from the target question and answer collection as the target question, or to select topics from easy to difficult according to the difficulty of the questions in the target question and answer collection.
[0056] Step S300, receiving the first question sentence generated by the large language model end, and sending the first question sentence to the user end. In this embodiment, after the large language model end generates the first question sentence, the first question sentence is sent to the engine end, and the engine end receives the first question sentence generated by the large language model end, and sends the first question sentence to the user end, so that the user end can answer based on the first question sentence sent by the engine end.
[0057] Step S400, obtaining a first answer sentence sent by the user end, the first answer sentence being generated in response to the first question sentence. In this embodiment, the user end generates a first answer sentence based on the first question sentence, and the user end sends the first answer sentence to the engine end, and the engine end receives the first answer sentence sent by the user end to proceed to the next step of the dialogue based on the first answer sentence.
[0058] Step S500, the standard answer to the target question and the first answer statement are sent to the large language model end, so that the large language model end determines whether the first answer statement is semantically consistent with the standard answer to the target question. In this embodiment, the engine end sends the standard answer to the target question and the first answer statement for the target question sent by the user end to the large language model end, so that the large language model end performs semantic analysis on the first answer statement and the standard answer to determine whether the first answer statement is semantically consistent with the standard answer to the target question. If the first answer statement is semantically consistent with the standard answer to the target question, it means that the user end's answer to the target question is correct, and the next question can be asked or the current conversation can be ended; if the first answer statement is semantically inconsistent with the standard answer to the target question, it means that the user end's answer to the target question is wrong, and it is necessary to repeat the question and answer of the current question, or to ask and answer the next question, or to end the current conversation.
[0059] Step S610, when the feedback result obtained from the large language model end is that the first answer statement is semantically consistent with the standard answer to the target question, the next question is sent to the large language model end so that the large language model end generates the next question statement based on the next question. In this embodiment, the large language model end sends the feedback result of the semantic analysis to the engine end. When the feedback result from the large language model end is that the first answer statement is semantically consistent with the standard answer to the target question, it is considered that the target question has been correctly answered at this time, and the next question can be asked. At this time, the engine end selects the next question from the target question and answer collection, and sends the selected next question to the large language model end, so that the large language model end generates the next question statement based on the next question, and repeats steps S200 to S500.
[0060] In an optional embodiment, the engine end may send evaluation feedback to the large language model end so that the large language model end generates an evaluation result based on the evaluation feedback. In this embodiment, this feedback result may be the advantages and disadvantages of the first answer sentence determined by the large language model end through semantic analysis or the like. The engine end may receive the evaluation result and send the evaluation result to the user end, and then select the next question from the target question and answer collection to execute step S610.
[0061] Step S620, when the feedback result obtained from the large language model end is that the first answer statement is semantically inconsistent with the standard answer to the target question, feedback information is sent to the large language model end so that the large language model end generates a feedback statement for the target question, and the feedback statement for the target question is used to prompt the user end to re-answer the target question. In this embodiment, the large language model end sends the feedback result of the semantic analysis to the engine end. When the feedback result of the large language model end is that the first answer statement is semantically inconsistent with the standard answer to the target question, it is considered that the target question is not correctly answered at this time, and the user end needs to re-answer the target question. At this time, the engine end will send feedback information to the large language model end, so that the large language model end can continue to generate a feedback statement for the target question based on the target question, wherein the feedback statement is used to prompt the user end to re-answer the target question, and steps S200 to S500 are repeated.
[0062] In an optional embodiment, the dialogue method for speech training further includes:
[0063] Step S621, when the number of rounds of answers to the target question by the user end exceeds the preset threshold, the next question is sent to the large language model end, so that the large language model end generates the next question statement based on the next question, wherein the number of rounds of answers to the target question is the number of times the large language model end generates the feedback statement of the target question. In this embodiment, if the number of rounds of answers to the target question by the user end exceeds the preset threshold, it is considered that the user end has not mastered the target question, and the target question can be skipped at this time, and the next question can be asked to the user end. Among them, the preset threshold can be pre-set, for example, it can be 2 times. When the large language model end generates two feedback statements for the same target question, it means that three questions have been asked for the target question (including a first question statement and two feedback statements). After asking multiple questions, if the user end still does not answer an answer that matches the standard answer, the question can be temporarily skipped and the next question can be asked directly to avoid the dialogue process being stuck in the same link and unable to proceed.
[0064] In an optional embodiment, the feedback statement of the target question includes at least one, the feedback statement is different from the first question statement, and the feedback statements of at least one target question are different from each other. In this embodiment, when asking questions to the same target question, since the user end has the possibility of needing multiple answers to answer the question correctly, this also makes it possible that more than one feedback statement is generated for the target question. When there is at least one feedback statement for the target question, each feedback statement is different from each other, and each feedback statement is different from the first question statement. It can be understood that different statements refer to different questioning methods or question expressions, but in essence, they are still asking questions to the same target question.
[0065] In an optional embodiment, the dialogue method for speech training further includes:
[0066] Step S700, when the number of questions answered by the user terminal meets the preset condition, the conversation is ended. The preset condition is that the number of questions answered by the user terminal reaches a preset proportion of the questions asked in the target question and answer collection. In this embodiment, the preset proportion can be 100% or 90%. For example, if the number of questions asked in the target question and answer collection is 5 and the preset proportion is 100%, the conversation is ended when the number of questions answered by the user terminal reaches 5. It should be noted that the current conversation can be ended as long as the number of questions answered by the user terminal reaches the preset condition, and the user terminal is not required to answer all these questions correctly.
[0067] In an alternative embodiment, see Figure 3 , Figure 3 This is a flow chart of the process of ending a conversation disclosed in this embodiment, such as Figure 3 As shown in , step S700 includes step S710 and step S720, wherein:
[0068] Step S710, when the number of questions answered by the user terminal meets the preset condition, a conversation end request is sent to the large language model terminal, so that the large language model terminal generates a conversation end statement based on the conversation end request. In this embodiment, the preset condition is that the number of questions answered by the user terminal reaches a preset proportion of questions asked in the target question and answer collection. When the number of questions answered by the user terminal meets the preset condition, the conversation can be ended. At this time, the engine terminal can send a conversation end request to the large language model terminal, so that the large language model terminal generates a conversation end statement based on the conversation end request.
[0069] Step S720, receiving a conversation ending statement, and sending the conversation ending statement to the user end to end the conversation. In this embodiment, the engine end receives the conversation ending statement sent by the large language model end, and sends the conversation ending statement to the user end to end the current conversation.
[0070] In an alternative embodiment, see Figure 4 , Figure 4 This is a flow chart of the process of starting a conversation disclosed in this embodiment, such as Figure 4 As shown in , before step S200, the dialogue method of the speech training includes:
[0071] Step S130, based on the dialogue scene selected by the user end triggering, a dialogue start request is sent to the large language model end, so that the large language model end generates a dialogue start sentence based on the dialogue start request. In this embodiment, after the user end triggers the selection of the dialogue scene, the engine end can send a dialogue start request to the large language model end based on the dialogue scene selected by the user end triggering, so that the large language model end generates a dialogue start sentence, i.e., an opening statement, based on the dialogue start request.
[0072] Step S140, sending the received conversation start sentence to the user end, and receiving the reply sentence sent by the user end. In this embodiment, the engine end sends the received conversation start sentence generated by the large language model end to the user end, and receives the reply sentence sent by the user end, and then starts the formal questioning session.
[0073] In an alternative embodiment, see Figure 5 , Figure 5 This is a schematic diagram of the process of obtaining a question and answer collection disclosed in this embodiment, such as Figure 5 As shown in , before step S100, the dialogue method for speech training also includes:
[0074] Step S110, obtaining the dialogue data. In this embodiment, the dialogue data can be a knowledge base document provided by the enterprise itself, which can be text data, image or voice data. The format and type of the dialogue data are not limited, and can be in doc. format, ppt. format, excel, pdf, html, vedio and other formats.
[0075] Step S120, the dialogue data is parsed to obtain the dialogue data, the dialogue data includes several types of question and answer collections, each type of question and answer collection corresponds to a dialogue scene, and each type of question and answer collection includes at least one question and a corresponding standard answer. The engine end parses the dialogue data and extracts the dialogue data therefrom, the dialogue data includes several types of question and answer collections, each question and answer collection corresponds to a dialogue scene, and the dialogue scene can be a different dialogue topic, for example, the dialogue scene can be a dialogue scene for cold medicine, a dialogue scene for immune diseases, a dialogue scene for cardiovascular diseases, etc. In the specific implementation process, the knowledge base document can be divided into blocks and then vectorized for storage, and high-frequency and important words can be identified. Of course, it is understandable that important words can also be manually marked. Then, the associated knowledge is matched through the vector matching algorithm, and then the knowledge points are refined and summarized through LLM to form a noun explanation question and answer. The high-frequency knowledge points and questions in the text are repeatedly retrieved, and then a question list is formed, and the knowledge associated with the question answer is found for the question list, and then the core answer is summarized and refined through LLM, and the question and answer pairs are extracted to obtain a question and answer collection. The conversation data can be used to automatically generate a collection of questions and answers, simplifying the amount of data required when building conversation training scenarios and scripts.
[0076] See also Figure 6 , Figure 6 This is a schematic diagram of the interaction between three terminals in the dialogue system for speech training disclosed in this embodiment. In order to facilitate the understanding of this solution, the following is combined with Figure 6 An example is given to illustrate this scheme.
[0077] like Figure 6As shown in , the engine end can send a conversation start request to the large language model end based on the conversation scenario triggered and selected by the user end. For example, the conversation scenario can include introducing biological products to doctors, introducing serious illness insurance products to customers, introducing high-value classic handbags to customers, and other scenarios, where the conversation scenario triggered and selected by the user end can be, for example, introducing biological products to doctors. The engine end can receive the user's touch selection action or click selection action on the conversation scene to determine the conversation scenario triggered and selected by the user end, and send a conversation start request to the large language model end based on the conversation scenario.
[0078] The large language model generates a conversation start statement based on the conversation start request, i.e., an opening statement, and sends the opening statement to the engine. The engine sends the opening statement to the user end, and receives the user end's reply to the opening statement, i.e., receives the reply statement sent by the user end.
[0079] Based on the dialogue scenario triggered by the user, the engine selects a set of questions and answers that match the dialogue scenario from the dialogue corpus, selects the target question from the set of questions and answers, and sends the target question to the large language model. For example, the target question can be "the mechanism of action of a new biological agent for systemic lupus erythematosus".
[0080] The large language model generates the first question sentence based on the target question. The large language model generates a conversational question sentence based on the target question based on the language model, and sends the generated first question sentence to the engine. For example, the first question sentence can be "Can you please introduce the mechanism of action of this biological agent?"
[0081] The engine side receives the first question sentence generated by the large language model side, and sends the first question sentence to the user side. The user side answers based on the first question sentence, obtains a first answer sentence, and sends the first answer sentence to the engine side.
[0082] After receiving the first answer statement, the engine sends the standard answer to the target question and the first answer statement generated by the user based on the first question statement to the large language model. After receiving the standard answer and the first answer statement, the large language model performs semantic analysis on the standard answer and the first answer statement to determine whether the semantics of the first answer statement matches the semantics of the standard answer, and sends the determination result to the engine.
[0083] When the judgment result received by the engine end is that the semantics of the first answer sentence matches the semantics of the standard answer, the next question can be selected from the question and answer collection, and the selected next question can be sent to the large language model end. The large language model end generates a next question sentence based on the next question, and the user end can answer the next question based on the next question sentence. The answer sentence will be sent to the large language model end through the engine end for semantic analysis again until the user end answers the next question correctly or the number of rounds of the user end answering the next question exceeds a preset threshold. At this time, the number of questions answered by the current user end can be combined to determine whether to end the conversation or ask the next question.
[0084] When the judgment result received by the engine end is that the semantics of the first answer sentence does not match the semantics of the standard answer and the number of rounds of the user end answering the target question does not exceed the preset threshold, the user can be asked to continue to supplement the answer until the user end answers the target question correctly or the number of rounds of the user end answering the target question exceeds the preset threshold. Specifically, the engine end sends feedback information to the large language model end, informing the large language model end to only provide feedback on the user's answer without asking new questions. After receiving the feedback information, the large language model end will generate a feedback statement for the target question and feed the feedback statement back to the engine end. For example, the feedback statement can be "Thank you for sharing the information. However, I would like to learn more about the specific mechanism of action of the biological agent, especially how it works in the immune system." This feedback statement is a continued question to the target question "The mechanism of action of new biological agents for systemic lupus erythematosus". The user end can continue to answer the target question based on the feedback statement. The answer statement will be sent to the large language model end through the engine end for semantic analysis again until the user end answers the target question correctly or the number of rounds of the user end answering the target question exceeds the preset threshold. At this time, it can be determined whether to end the conversation or ask the next question based on the number of questions answered by the current user end.
[0085] When the number of questions answered by the user meets the preset conditions, for example, all questions in the dialogue scenario are answered, the engine can choose to end the dialogue. At this time, the engine can send a dialogue end request to the large language model. The large language model will generate a dialogue end statement based on the dialogue end request and send it to the engine. The engine sends the dialogue end statement to the user to end the dialogue.
[0086] See also Figure 7 , Figure 7 The structure diagram of a dialogue device for speech training disclosed in this embodiment is shown in FIG. The dialogue device for speech training is applied to the engine end. The engine end, the user end and the large language model end constitute a speech training system. The engine end interacts with the user end and the large language model end respectively. The user end and the large language model end do not interact with each other, and the data of the large language model end is independent of the engine end. Figure 7 As shown in , the dialogue device for the speech training includes:
[0087] The question and answer collection selection module 100 is used to filter the question and answer collection that matches the dialogue scene from the dialogue material based on the dialogue scene triggered and selected by the user end. The dialogue material includes question and answer collections under several types of dialogue scenes as the target question and answer collection. Each type of question and answer collection includes at least one question and a corresponding standard answer.
[0088] The question sentence generating module 200 is used to send the target question in the target question and answer collection to the large language model end, so that the large language model end generates a first question sentence based on the target question.
[0089] The question sentence sending module 300 is used to receive a first question sentence generated by the large language model end, and send the first question sentence to the user end.
[0090] The answer statement receiving module 400 is used to obtain a first answer statement sent by the user terminal, where the first answer statement is generated in response to the first question statement.
[0091] The sentence answer sending module 500 is used to send the standard answer to the target question and the first answer sentence to the large language model end, so that the large language model end determines whether the first answer sentence is semantically consistent with the standard answer to the target question.
[0092] The next question generation module 610 is used to send the next question to the large language model end when the feedback result obtained from the large language model end is that the first answer sentence is semantically consistent with the standard answer to the target question, so that the large language model end generates the next question sentence based on the next question.
[0093] The feedback sentence generation module 620 is used to send feedback information to the large language model end when the feedback result obtained from the large language model end is that the first answer sentence is semantically inconsistent with the standard answer to the target question, so that the large language model end generates a feedback sentence for the target question. The feedback sentence for the target question is used to prompt the user end to re-answer the target question.
[0094] According to the dialogue method, device, computer equipment and storage medium for speech training disclosed in the embodiment of the present invention, the engine end is applied, and the engine end, the user end and the large language model end constitute a speech training system. The engine end interacts with the user end and the large language model end respectively, and the user end and the large language model end do not interact with each other and the data of the large language model end is independent of the engine end. Based on the dialogue scene triggered by the user end, the question and answer collection is filtered as the target question and answer collection, and then the target question in the target question and answer collection is sent to the large language model end so that the large language model end generates a first question statement, receives the first question statement and sends the first question statement to the user end, and then receives the first answer statement returned by the user end, and sends the first answer statement and the standard answer to the large language model end together. The large language model end determines whether the semantics of the first answer statement and the standard answer are consistent. When the semantics of the two are consistent, the next question is sent to the large language model end to generate the next question statement. When the semantics of the two are inconsistent, feedback information is sent to the large language model end so that the large language model end generates a feedback statement to prompt the user to answer the target question again. The user side interacts with the engine side through data, so that the engine side can perceive the complete dialogue process during the entire dialogue process of speech training. In addition, separating the engine side from the large language model side can not only simplify the amount of data for speech training, but also utilize the semantic analysis and other functions of the large language model side to improve the flexibility of the generated dialogue scenes and scripts, improve the efficiency of constructing dialogue scenes, and thus improve the training effect of dialogue speech.
[0095] It will be appreciated by those skilled in the art that, under the premise of no conflict, the above-mentioned preferred solutions can be freely combined and superimposed. Among them, the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings, for example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions. The numbering of each step in this article is only for the convenience of explanation and reference, and is not used to limit the order of execution. The specific execution order is determined by the technology itself, and those skilled in the art can determine various allowable and reasonable orders based on the technology itself.
[0096] Those skilled in the art will appreciate that, without conflict, the above-mentioned preferred solutions can be freely combined and superimposed.
[0097] It should be understood that the above-mentioned embodiments are merely illustrative and not restrictive. Without departing from the basic principles of the present invention, various obvious or equivalent modifications or substitutions that can be made by those skilled in the art to the above-mentioned details will all be included in the scope of the claims of the present invention.
Claims
1. A dialogue method for speech training, characterized in that: Applied to the engine end, the engine end, the user end and the large language model end constitute a speech training system, the engine end respectively interacts with the user end and the large language model end for data, the user end and the large language model end do not interact with data, and the data of the large language model end is independent of the engine end, wherein the engine end is used to select questions to ask and control the dialogue process, and the large language model end is used to perform semantic analysis, and the dialogue method includes: Step S100, based on the dialogue scenario triggered and selected by the user terminal, a question and answer collection matching the dialogue scenario is selected from the dialogue material as a target question and answer collection, wherein the dialogue material includes question and answer collections under several types of dialogue scenarios, and each type of question and answer collection includes at least one question and a corresponding standard answer; Step S200, sending a target question in the target question and answer collection to the large language model end, so that the large language model end generates a first question sentence based on the target question; Step S300, receiving a first question sentence generated by the large language model end, and sending the first question sentence to the user end; Step S400, obtaining a first answer statement sent by the user terminal, where the first answer statement is generated in response to the first question statement; Step S500, sending the standard answer to the target question and the first answer statement to the large language model end, so that the large language model end determines whether the first answer statement is semantically consistent with the standard answer to the target question; Step S610: When the feedback result obtained from the large language model end is that the first answer sentence is semantically consistent with the standard answer to the target question, send the next question to the large language model end, so that the large language model end generates a next question sentence based on the next question; Step S620: When the feedback result obtained from the large language model end is that the first answer statement is semantically inconsistent with the standard answer to the target question, feedback information is sent to the large language model end so that the large language model end generates a feedback statement for the target question, and the feedback statement for the target question is used to prompt the user end to re-answer the target question.
2. The dialogue method for speech training according to claim 1 is characterized in that: Before step S100, the method further includes: Step S110, obtaining conversation data; Step S120, parsing the dialogue data to obtain dialogue data, the dialogue data includes several types of question and answer collections, each type of question and answer collection corresponds to a dialogue scene, and each type of question and answer collection includes at least one question and a corresponding standard answer.
3. The dialogue method for speech training according to claim 1 is characterized in that: After step S620, the method further includes: Step S621, when the number of answer rounds of the user end to the target question exceeds a preset threshold, the next question is sent to the large language model end, so that the large language model end generates a next question statement based on the next question, wherein the number of answer rounds of the target question is the number of times the large language model end generates a feedback statement for the target question.
4. The dialogue method for speech training according to claim 1 or 3, characterized in that: The feedback sentence of the target question includes at least one, the feedback sentence is different from the first question sentence, and the feedback sentences of at least one target question are different from each other.
5. The dialogue method for speech training according to claim 1, characterized in that: The method further comprises: Step S700, when the number of questions answered by the user terminal meets a preset condition, the conversation is ended. The preset condition is that the number of questions answered by the user terminal reaches a preset proportion of questions asked in the target question and answer collection.
6. The dialogue method for speech training according to claim 1, characterized in that: Before step S200, the method includes: Step S130, sending a dialog start request to the large language model end based on the dialog scenario triggered and selected by the user end, so that the large language model end generates a dialog start sentence based on the dialog start request; Step S140, sending the received dialogue initiation statement to the user terminal, and receiving a reply statement sent by the user terminal.
7. The dialogue method for speech training according to claim 5, characterized in that: The step S700 includes: Step S710, when the number of questions answered by the user terminal meets a preset condition, sending a dialog ending request to the large language model terminal, so that the large language model terminal generates a dialog ending statement based on the dialog ending request; Step S720, receiving the conversation ending statement, and sending the conversation ending statement to the user terminal to end the conversation.
8. A dialogue device for speech training, characterized in that: Applied to the engine end, the engine end, the user end and the large language model end constitute a speech training system, the engine end respectively interacts with the user end and the large language model end for data, the user end and the large language model end do not interact with data, and the data of the large language model end is independent of the engine end, wherein the engine end is used to select questions to ask and control the dialogue process, the large language model end is used to perform semantic analysis, and the dialogue device includes: A question and answer collection selection module (100) is used to select a question and answer collection matching the dialogue scene from the dialogue material based on the dialogue scene triggered and selected by the user terminal, as a target question and answer collection, wherein the dialogue material includes question and answer collections under several types of dialogue scenes, and each type of question and answer collection includes at least one question and a corresponding standard answer; A question statement generating module (200), configured to send a target question in the target question and answer collection to the large language model end, so that the large language model end generates a first question statement based on the target question; A question sentence sending module (300), configured to receive a first question sentence generated by the large language model end, and send the first question sentence to the user end; An answer statement receiving module (400) is used to obtain a first answer statement sent by the user terminal, wherein the first answer statement is generated in response to the first question statement; A sentence answer sending module (500) is used to send the standard answer to the target question and the first answer sentence to the large language model end, so that the large language model end determines whether the first answer sentence is semantically consistent with the standard answer to the target question; A next question generation module (610) is used to send the next question to the large language model end when the feedback result obtained from the large language model end is that the first answer sentence is semantically consistent with the standard answer to the target question, so that the large language model end generates a next question sentence based on the next question; A feedback sentence generating module (620) is used to send feedback information to the large language model end so that the large language model end generates a feedback sentence for the target question when the feedback result obtained from the large language model end is that the first answer sentence is semantically inconsistent with the standard answer to the target question. The feedback sentence for the target question is used to prompt the user end to re-answer the target question.
9. A computer device, characterized in that: include: A conversation method for speech training as described in any one of claims 1 to 7; or, including a device as described in claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program stored in the storage medium is used to be executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Information processing method and device, equipment and readable storage medium
CN111444729A
Human-computer interaction method and device, related equipment and computer program product
CN118364086A