Question and answer method and device based on large language model and electronic equipment
By using a question-and-answer method based on a large language model, and leveraging historical dialogue records and expert systems for logical deduction, accurate answer texts are generated. This addresses the shortcomings of traditional question-and-answer methods in handling complex questions, achieving coherence and accuracy in dialogue, and improving mental health.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU SHULINGJI TECH CO LTD
- Filing Date
- 2023-10-12
- Publication Date
- 2026-04-17
AI Technical Summary
Traditional question-and-answer methods are not accurate enough when faced with complex questions, and cannot support effective responses to complex questions, resulting in a poor experience for visitors.
This paper adopts a question-answering method based on a large language model. By acquiring the current visitor's statements and historical dialogue records, it uses the first large language model for analysis and feature extraction, and combines it with a pre-set expert system for logical deduction to generate accurate answer text, which is then added to the historical dialogue records to achieve the coherence and accuracy of the dialogue.
It improved the accuracy of answers to complex questions and the coherence of conversations, making visitors feel understood, relieving psychological stress, and enhancing mental health.
Smart Images

Figure CN121880489A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, specifically to a question-answering method, apparatus, and electronic device based on a large language model. Background Technology
[0002] Various pressures and troubles in life can lead to psychological problems such as low mood, anxiety, and depression. Psychological counseling can help people understand and alleviate these emotional stresses, improving their psychological comfort. Human-computer interaction is a discipline that studies the interactive relationship between systems and users. Systems can be various types of machines, as well as computerized systems and software. Through question-and-answer interactions with psychological counseling robots, psychological distress can be overcome and mental health improved. In related technologies, traditional question-and-answer methods typically use speech recognition technology to obtain the question text, and then use data retrieval and databases to determine the answer. This approach has low accuracy and is insufficient for answering complex questions. Summary of the Invention
[0003] This application provides a question-answering method, apparatus, and electronic device based on a large language model, which provides high accuracy in answering questions and can support answers to sufficiently complex questions.
[0004] The technical solution of this application embodiment is as follows: In a first aspect, embodiments of this application provide a question-answering method based on a large language model, the method comprising: Obtain the current visitor's statement; if the current visitor's statement is not the first statement, obtain the historical dialogue record and the first mental map. The current visitor's statement, the first mental map, and the historical dialogue record are input into the first large language model to obtain the second mental map and the first cause summary text. The first large language model is a statement analysis model. Using a pre-set expert system, logical deduction is performed on the second mental map and the first cause summary text to obtain the first action suggestion; The first action suggestion is input into the preset second language model to obtain the first answer text. The second language model is a text generation model. The first response text and the current visitor's statement are added to the historical dialogue record. The second mental map is used as the first mental map. The steps of obtaining the current visitor's statement are executed. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained until the dialogue ends.
[0005] In the above technical solution, the current visitor's statement is first obtained. If the current visitor's statement is not their first statement, the historical dialogue record and the first mental map are obtained. By obtaining the historical dialogue record and the first mental map, the coherence of the question and answer can be improved, supporting answers to more complex questions. The current visitor's statement, the first mental map, and the historical dialogue record are input into the first large language model to obtain the second mental map and the first cause summary text. The first large language model is a statement analysis model that analyzes the current visitor's expectations, which is beneficial for obtaining the subsequent response text. A preset expert system is used to logically deduce the first action suggestion from the second mental map and the first cause summary text. Guided by a pre-set expert system, more accurate answers can be obtained. The first action suggestion is input into the pre-set second language model to obtain the first answer text. The second language model is a text generation model, and the first answer text conforms to the dialogue logic and has high accuracy. The first answer text and the current visitor's statement are added to the historical dialogue record. The second mental map is used as the first mental map. The process of obtaining the current visitor's statement is repeated. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained. This process continues until the dialogue ends. The answer to the current visitor's statement is based on the historical dialogue record. This can support answers to sufficiently complex questions, and the dialogue is coherent with high accuracy.
[0006] In some embodiments of this application, when the current visitor's statement is their first statement, the method further includes: The first large language model is used to extract features from the current visitor's statements to obtain the third psychological map and the second reason summary text; By using a pre-set expert system to logically deduce the third psychological map and the second cause summary text, a second action suggestion is obtained; The second action suggestion is input into the preset second language model to obtain the second response text; The current visitor's statement corresponding to the first statement and the second response text are used as historical dialogue records, and the third mental map is used as the first mental map.
[0007] In the above technical solution, when the current visitor's statement is the first statement, since the first and second language models are pre-trained, the first language model, the expert system, and the second language model are used for processing, which can obtain a relatively accurate answer text and support more complex question responses.
[0008] In some embodiments of this application, the step of using a preset expert system to logically deduce the second psychological map and the first cause summary text to obtain the first action suggestion includes: A dialogue process framework is constructed based on the second mental map and the summary text of the first cause, and multiple inference combinations are obtained based on the dialogue process framework; Based on the various deduction combinations and the guidance rules of the expert system, the first action suggestion is obtained.
[0009] In the above technical solution, a variety of deduction combinations are obtained by constructing a process dialogue framework, and the first action suggestion is obtained according to the guidance rules of the expert system, which can answer more accurately.
[0010] In some embodiments of this application, the method further includes a training step for the first large language model and the second large oracle model, the training step including: Obtain a historical corpus dataset, which includes multiple visitor corpus data and first mental map labels and first cause summary labels corresponding to the multiple visitor expected data; Input the corpus data of each visitor, the first mental map label, and the first reason summary label into the preset third language model, and train the third language model to obtain the first language model. The historical corpus dataset also includes counselor corpus corresponding to each of the client's expected data, and a first action label corresponding to the counselor corpus; The counselor corpora and the first action labels are input into a preset fourth language model, and the fourth language model is trained to obtain the second language model.
[0011] In the above technical solution, by training the third and fourth language models respectively, the first and second language models are obtained, which is beneficial for using the first language model for sentence analysis and the second language model for text generation.
[0012] In some embodiments of this application, after inputting the first action suggestion into a preset second language model to obtain the first response text, the method further includes: Based on the preset time, collect the current visitor's statement and the first response text; The current visitor's statement and the first response text are labeled to obtain the second psychological map label and the second cause summary label corresponding to the current visitor's statement, as well as the second action label corresponding to the first response text; The current visitor's statement, the second mental map label, the second reason summary label, the first answer text, and the second action label are added to the historical corpus dataset to obtain an updated corpus dataset; The parameters of the first and second largest language models are adjusted using the updated corpus dataset.
[0013] In the above technical solution, the current visitor's statement and the first response text are labeled, and the current visitor's statement, the second mental map label, the second reason summary label, the first response text and the second action label are added to the historical corpus dataset. Then, the parameters of the model are adjusted using the updated corpus dataset, so that the generalization of the first and second language models is better and more accurate response text is generated.
[0014] In some embodiments of this application, the first mental map label is obtained through the following steps: The manual annotation of the speech data of each of the visitors is obtained, and the manual annotation includes events, emotions, and causal relationships; By treating the events and emotions as entities and causal relationship features as connecting edges, the first psychological graph label is constructed.
[0015] In the above technical solution, the connections between various statements can be obtained by constructing a first mental map to support responses to complex questions.
[0016] In some embodiments of this application, the step of inputting the various counselor corpora and the first action labels into a preset fourth language model, and training the fourth language model to obtain the second language model, includes: The consultant's corpus and the first action label are input into the preset fourth language model to obtain the predicted response statement; The value of the loss function is calculated based on the predicted response statement and the first action label; The parameters of the fourth language model are adjusted using the value of the loss function to obtain the second language model.
[0017] In the above technical solution, the second language model is obtained by training the fourth language model with a large amount of data. This makes the answer text generated by the second language model more accurate and can support more complex question responses.
[0018] Secondly, embodiments of this application provide a question-answering device based on a large language model, the device comprising: The data acquisition module is used to acquire the current visitor's statement, and if the current visitor's statement is not the first statement, to acquire the historical dialogue record and the first mental map; The first data processing module is used to input the current visitor's statement, the first mental map, and the historical dialogue record into the first large language model to obtain the second mental map and the first cause summary text. The first large language model is a statement analysis model. The second data processing module is used to perform logical deduction on the second psychological map and the first cause summary text using a preset expert system to obtain the first action suggestion; The third data processing module is used to input the first action suggestion into a preset second language model to obtain the first answer text. The second language model is a text generation model. The looping question-and-answer module is used to add the first answer text and the current visitor's statement to the historical dialogue record, use the second mental map as the first mental map, and execute the steps of obtaining the current visitor's statement. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained until the dialogue ends.
[0019] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform any of the methods provided in the first aspect above.
[0020] Fourthly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed, perform the method described in any one of the methods provided in the first aspect above.
[0021] In summary, one or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: 1. By employing the technical means of acquiring historical dialogue records and a first mental map, analyzing the current visitor's statements, historical dialogue records, and the first mental map using a first major language model, deriving a first action suggestion using a preset expert system, generating a first response text using a second major language model, and adding the current visitor's statements and the first response text to the historical dialogue records for later use, the problem of insufficient support for answers to complex questions is effectively solved. This application embodiment accurately captures information from the current visitor's statements, historical dialogue records, and the first mental map using the first major language model, then derives action suggestions and generates response text, and adds the current visitor's statements and the first response text to the historical dialogue records, thus achieving a relatively accurate response to complex questions.
[0022] 2. By constructing the first psychological map and summarizing the causes, we can analyze the emotions in the visitor's language data, making the answers more natural, creating the experience of the visitor communicating with someone, and enabling the visitor to feel understood and relieve psychological stress. Attached Figure Description
[0023] Figure 1 This is one of the flowcharts illustrating a question-answering method based on a large language model provided in one embodiment of this application; Figure 2 yes Figure 1 A flowchart illustrating a sub-step of step S130; Figure 3 This is a second flowchart illustrating a question-answering method based on a large language model provided in one embodiment of this application; Figure 4 This is a schematic diagram of model training for a question-answering method based on a large language model provided in one embodiment of this application; Figure 5 This is the third flowchart of a question-answering method based on a large language model provided in one embodiment of this application; Figure 6 This is a schematic diagram of the first mental map of a question-answering method based on a large language model provided in one embodiment of this application; Figure 7 This is a schematic diagram of the overall process of a question-answering method based on a large language model provided in one embodiment of this application; Figure 8 This is a schematic diagram of the structure of a question-answering device based on a large language model provided in one embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0025] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.
[0026] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0027] Various pressures and troubles in life can lead to psychological problems such as low mood, anxiety, and depression. Psychological counseling can help people understand and alleviate these emotional stresses, improving their psychological comfort. Human-computer interaction is a discipline that studies the interactive relationship between systems and users. Systems can be various types of machines, as well as computerized systems and software. By interacting with psychological counseling robots through question-and-answer sessions, psychological distress can be overcome and mental health can be improved. However, traditional question-and-answer methods can only achieve simple conversations with clients. When clients ask complex questions, they cannot provide appropriate answers or their answers are rather curt, resulting in a poor experience for the client.
[0028] Based on this, embodiments of this application provide a question-answering method, apparatus, and electronic device based on a large language model. This question-answering method first obtains the current visitor's statement. If the current visitor's statement is not their first statement, it obtains historical dialogue records and a first mental map. By obtaining historical dialogue records and the first mental map, the coherence of subsequent questions and answers can be improved to support answers to more complex questions. The current visitor's statement, the first mental map, and the historical dialogue records are input into a first large language model to obtain a second mental map and a first cause summary text. The first large language model is a statement analysis model that analyzes the current visitor's expectations, which is beneficial for obtaining subsequent answer text. A preset expert system is used to analyze the second mental map and the first cause summary. Logical deduction is performed on the text to obtain the first action suggestion. Guided by a preset expert system, the answer is more accurate. The first action suggestion is input into a preset second language model to obtain the first answer text. The second language model is a text generation model. The first answer text conforms to the dialogue logic and has high accuracy. The first answer text and the current visitor's statement are added to the historical dialogue record. The second mental map is used as the first mental map. The process of obtaining the current visitor's statement is repeated. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained. This process continues until the dialogue ends. The answer to the current visitor's statement is based on the historical dialogue record. This method can support answers to sufficiently complex questions, and the dialogue is coherent with high accuracy.
[0029] It should be noted that the question-answering method based on the large language model is applied to question-answering robots or question-answering electronic devices. By accurately capturing the information of the current visitor's statement, historical dialogue records, and the first mental map through the first large language model, action suggestions are deduced and answer text is generated. The current visitor's statement and the first answer text are added to the historical dialogue records, thus achieving a relatively accurate response to complex questions.
[0030] The technical solutions provided in the embodiments of this application will be further described below with reference to the accompanying drawings.
[0031] Reference Figure 1 , Figure 1 This is a flowchart illustrating the question-answering method based on a large language model provided in this application embodiment. The question-answering method based on a large language model is applied to a question-answering device based on a large language model, and is executed by a question-answering robot, electronic device, or computer-readable storage medium. The question-answering method based on a large language model includes steps S110, S120, S130, S140, and S150.
[0032] Step S110: Obtain the current visitor's statement. If the current visitor's statement is not the first statement made, obtain the historical dialogue record and the first mental map.
[0033] In one embodiment, taking a question-and-answer robot as an example, a visitor inputs a question into the robot. This question is the visitor's current statement. The current visitor's statement is obtained through a preset reading method, and it includes at least one segment of corpus data. Obtaining the current visitor's statement provides data support for accurately answering the visitor's question based on it. The preset reading method can obtain the current visitor's statement through an API interface or a text reading method, such as calling the `read()` function. It should be noted that a current visitor's statement can be simple or complex. It can include not only a linguistic description of the communication problem but also emotional vocabulary, which helps in responding better and making the visitor feel understood, thus alleviating psychological pressure. For example, a visitor inputs, "I feel wronged after being criticized by my homeroom teacher and believe I have been wronged. What should I do?" This statement contains both a problem description and a description of the feeling of being wronged. Even when dealing with complex statements, the embodiments of this application can provide relatively accurate responses.
[0034] In another embodiment, taking a question-and-answer robot as an example, the visitor can input voice data, which is first converted into text data, representing the visitor's current statement. This conversion can be achieved using a preset voice conversion algorithm, such as an LSTM algorithm or a neural network algorithm with an attention mechanism, as long as it accurately converts to text data; details will not be elaborated here.
[0035] In another embodiment, when the current visitor's statement is not their first statement, the historical dialogue record and the first mental map are obtained. Using a pre-defined text reading method, such as calling the `read()` function, the historical dialogue record and the first mental map can be retrieved, improving the coherence of the question-and-answer sessions and supporting responses to more complex questions. The historical dialogue record consists of statements from visitors other than the current visitor and their responses. The first mental map includes events, emotions, and the mapping relationship between events and emotions. The historical dialogue record and the first mental map are obtained through the visitor's dialogue, which will be explained in detail later and will not be elaborated upon here.
[0036] Step S120: Input the current visitor's statement, the first mental map, and the historical dialogue record into the first major language model to obtain the second mental map and the first cause summary text. The first major language model is a statement analysis model.
[0037] In one embodiment, the current visitor's statement, the first mental map, and the historical dialogue record are first converted into feature vectors. These converted features are then processed using a first major language model for feature extraction and analysis to obtain a second mental map and a summary text of the first cause. This analysis of the current visitor's expectations is beneficial for obtaining the subsequent response text. The first major language model is a statement analysis model, which can be a pre-trained model such as GPT and BERT, or a deep learning-based neural network model, such as a recurrent neural network or a variant of a long short-term memory network and the Transformer model.
[0038] Step S130: Use a preset expert system to perform logical deduction on the second mental map and the first cause summary text to obtain the first action suggestion.
[0039] In one embodiment, the preset expert system provides guidance rules developed by experienced professionals, enabling logical deduction. By using the preset expert system to logically deduce the second mental map and the first cause summary text, a first action suggestion is obtained. Guided by the preset expert system, the answer can be more accurate. The second mental map includes events, emotions, and the mapping relationship between events and emotions; the first cause summary is a summary of the current visitor's statements.
[0040] like Figure 2As shown, a pre-set expert system is used to logically deduce the first action suggestion from the second mental map and the summary text of the first cause, including but not limited to the following steps: Step S131: Construct a dialogue process framework based on the second mental map and the summary text of the first cause, and obtain various inference combinations based on the dialogue process framework.
[0041] In one possible embodiment of this application, the second mental map includes events, emotions, and the mapping relationship between events and emotions. A question-and-answer dialogue framework is established based on the events, emotions, and the mapping relationship. Each event corresponds to multiple emotions, and one emotion may correspond to multiple events. Using mathematical combinations, the emotions corresponding to each event are listed one by one, resulting in multiple deductive combinations. Based on the obtained deductive combinations, it is beneficial to subsequently obtain suggestions for the first action.
[0042] Step S132: Based on various deduction combinations and the guidance rules of the expert system, the first action suggestion is obtained.
[0043] In one possible embodiment of this application, based on the multiple deduction combinations obtained in step S131, a deduction combination that conforms to the guidance rules of the expert system is selected to obtain a first action suggestion. With guidance from a preset expert system, the answer can be more accurate.
[0044] For example, if a visitor is criticized by the homeroom teacher for talking in class and has an argument with the teacher, this corresponds to a communication problem. This can lead to feelings of anger and dissatisfaction, or even feelings of being wronged. This can create multiple sets of inference combinations. According to the guidance rules, the visitor can choose to respond with feelings of being wronged to comfort the visitor and obtain the first action suggestion.
[0045] Step S140: Input the first action suggestion into the preset second language model to obtain the first answer text. The second language model is a text generation model.
[0046] In one embodiment, the preset second language model and the first language model can have the same model structure. The second language model is a text generation model, which can be a pre-trained model, such as GPT and BERT, or a deep learning-based neural network model, such as a recurrent neural network or a variant of a long short-term memory network and a Transformer model. The preset second language model is used to process the first action suggestion into text, resulting in a first response text. This first response text conforms to the dialogue logic and has high accuracy.
[0047] Step S150: Add the first response text and the current visitor's statement to the historical dialogue record, use the second mental map as the first mental map, and execute the steps of obtaining the current visitor's statement. If the current visitor's statement is not the first statement, obtain the historical dialogue record and the first mental map until the dialogue ends.
[0048] In one embodiment, the first response text and the current visitor's statement are added to the historical dialogue record. At this time, the current visitor's statement and the first response text are part of the historical dialogue. Adding the current visitor's statement and the first response text can improve the logic of the previous and subsequent dialogues. The second mental map is used as the first mental map, which is the updated mental map. Then, the steps of obtaining the current visitor's statement are performed. If the current visitor's statement is not the first statement, the steps of obtaining the historical dialogue record and the first mental map are performed. By referring to the historical dialogue record and the first mental map, the visitor's psychological changes can be analyzed, thereby supporting the answering of complex questions.
[0049] like Figure 3 As shown, when the current visitor's statement is their first statement, the question-answering method based on the large language model also includes, but is not limited to, the following steps: Step S160: Use the first language model to extract features from the current visitor's statements to obtain the third psychological map and the second cause summary text.
[0050] In one embodiment, when the current visitor's statement is the first statement made, there is no historical dialogue record at this time. First, the first language model is used to perform feature extraction and analysis on the current visitor's statement to obtain the third psychological map and the second reason summary text. Analyzing the current visitor's expectations is beneficial for obtaining the second response text later.
[0051] Step S170: Use a preset expert system to logically deduce the third mental map and the second cause summary text to obtain the second action suggestion.
[0052] In one embodiment, based on the third mental map and the second cause summary text obtained in step S160, a preset expert system is used to logically deduce the third mental map and the second cause summary text to obtain a second action suggestion. Guided by the preset expert system, the answer can be more accurate. The third mental map includes events, emotions, and the mapping relationship between events and emotions; the second cause summary is a summary of the current visitor's statements.
[0053] Step S180: Input the second action suggestion into the preset second language model to obtain the second response text.
[0054] In one embodiment, the second action suggestion is processed by the second language model to generate a second response text. Although the second response text is the response text of the first dialogue and does not refer to the historical dialogue record, it is processed by the first language model, the expert system and the second language model to generate a relatively accurate response text and can also support the response to complex questions.
[0055] Step S190: The current visitor's statement corresponding to the first statement and the second response text are recorded as historical dialogue records, and the third mental map is used as the first mental map.
[0056] In one embodiment, the current visitor's statement corresponding to the initial statement and the second response text are added to the historical dialogue record. In this case, the current visitor's statement and the second response text constitute the historical dialogue. Referring to historical responses may improve the logical coherence of the dialogue. The second mental map is used as the first mental map, representing the updated mental map. By referencing the historical dialogue record and the first mental map, the visitor's psychological changes can be analyzed, thereby supporting the answering of complex questions.
[0057] like Figure 4 As shown, the question-answering method based on the large language model also includes training steps for the first large language model and the second large oracle model. The training steps include, but are not limited to, the following steps: Step S210: Obtain the historical corpus dataset, which includes multiple visitor corpus data and first mental map labels and first cause summary labels corresponding to multiple visitor expected data.
[0058] In one embodiment, the historical corpus dataset includes multiple visitor speech data sets and corresponding first mental graph labels and first cause summary labels. The historical corpus dataset can be obtained by crawling web pages, or by collecting dialogue data between visitors and counselors. The visitor speech data is then manually labeled and summarized to obtain the first mental graph labels and first cause summary labels, which are then saved. The historical corpus dataset is then read using a preset file reading method. For example, the preset file reading method is to call the `read()` function or the `txt()` function. By obtaining the historical corpus dataset, a large amount of visitor speech data and corresponding first mental graph labels and first cause summary labels can be obtained, which is beneficial for subsequent training of large language models and avoids overfitting of the large language model.
[0059] Step S220: Input the corpus data of each visitor, the first mental map label and the first cause summary label into the preset third language model, and train the third language model to obtain the first language model.
[0060] In one embodiment, the visitor corpus data, first mental map labels, and first cause summary labels are first converted into feature vectors. These converted features are then processed using a third language model for feature extraction and analysis. The third language model is then trained to obtain a first language model. For example, the third language model is a transform model. Training the model according to the transform model's training method yields the first language model, which is beneficial for subsequently using the first language model to generate the second mental map and the first language summary text.
[0061] Step S230: The historical corpus dataset also includes counselor corpus corresponding to the expected data of each visitor, and first action labels corresponding to the counselor corpus.
[0062] In one embodiment, the historical corpus dataset obtained in step S210 also includes counselor corpus corresponding to the expected data of each visitor, and a first action label corresponding to the counselor corpus. The first action label is obtained by manually annotating and saving the counselor corpus. The counselor corpus and the first action label are obtained through a preset file reading method, which is beneficial for subsequent model training to obtain the second language model. For example, the preset file reading method is calling the `read()` function.
[0063] In step S240, the corpus of each counselor and the first action label are input into the preset fourth language model, and the fourth language model is trained to obtain the second language model.
[0064] In one embodiment, the corpora of each consultant and the first action label are converted into feature vectors. A fourth language model is then used to perform text generation and analysis on the converted feature vectors, and the fourth language model is trained to obtain a second language model. The second language model, trained on a large amount of data, exhibits good generalization ability and can produce relatively accurate answers to complex questions.
[0065] In one specific embodiment, the corpora of each consultant and the first action label are input into a preset fourth language model, and the fourth language model is trained to obtain a second language model. Specifically: First, the corpora of each consultant and the first action label are transformed into feature vectors. Then, the preset fourth language model is used to perform text generation processing on the transformed feature vectors to obtain the predicted answer statement. A loss function is calculated on the predicted answer statement and the first action label according to a preset loss function to obtain the value of the loss function. The preset loss function can be the cross-entropy loss function or other loss functions commonly used in transform models, which will not be elaborated here. Then, the value of the loss function is used for back-training to adjust the parameters of the fourth language model to obtain the second language model. This facilitates the generation of the first answer text by the second language model, enabling answers to complex questions.
[0066] like Figure 5 As shown, after inputting the first action suggestion into the preset second large language model and obtaining the first answer text, the question-answering method based on the large language model also includes, but is not limited to, the following steps: Step S310: Based on a preset time, collect the current visitor's statements and the text of the first response.
[0067] In one embodiment, the preset time can be 0.2 seconds or 1 second, continuously polling for whether a visitor has asked a question, and obtaining collected data through a preset API, collecting and storing the current visitor's statement and the first response text, which is beneficial for obtaining updated corpus datasets later.
[0068] Step S320: Annotate the current visitor's statement and the first response text to obtain the second psychological map label and the second cause summary label corresponding to the current visitor's statement, as well as the second action label corresponding to the first response text.
[0069] In one embodiment, the current visitor's statement and the first response text obtained in step S310 are manually annotated to obtain the second mental map and the second cause summary label corresponding to the current visitor's statement, as well as the second action label corresponding to the first response text, and stored, which is beneficial for obtaining updated corpus datasets in the future.
[0070] Step S330: Add the current visitor's statement, the second mental map label, the second reason summary label, the first answer text, and the second action label to the historical corpus dataset to obtain the updated corpus dataset.
[0071] In one embodiment, according to a preset addition time, the stored current visitor statement, second mental map label, second cause summary label, first response text, and second action label are read out. These data are then converted into the format of the historical corpus dataset, and the converted data is added to the historical corpus dataset to obtain an updated corpus dataset. The preset addition time can be one month or three months, adjusted according to requirements, and will not be elaborated here.
[0072] Step S340: Adjust the parameters of the first and second language models using the updated corpus dataset.
[0073] In one embodiment, the updated corpus dataset obtained in step S330 is used to adjust the parameters of the first and second language models, as described above and will not be repeated here. Parameter adjustment enhances the generalization ability of both the first and second language models, resulting in more accurate response texts.
[0074] like Figure 6 As shown, constructing the first mental graph label specifically involves: manually annotating the corpus data of each visitor. The annotations include events, emotions, and causal relationships. Events and emotions are treated as entities, and causal relationship features are treated as connecting edges, thus constructing the first mental graph label. For example, events include: the visitor being criticized by the homeroom teacher for talking in class and arguing with the teacher, corresponding to a communication problem; the visitor's parents are divorced, and the visitor is raised by their father and grandmother, but the father is busy with work and rarely spends time with the visitor, corresponding to a relationship imbalance problem; the visitor is inattentive in class, likes to talk to classmates, seriously disrupting classroom order and affecting their own and others' learning, corresponding to a learning problem. Emotions include: the visitor feeling unfairly criticized by the homeroom teacher, corresponding to anger and dissatisfaction; the visitor feeling wronged and believing they have been wronged after being criticized by the homeroom teacher, corresponding to feeling wronged; emotionalization, corresponding to emotionalization, etc. Connecting edges are formed based on causal relationships, thus constructing the first mental graph label. By constructing the first mental graph, the connections between various statements can be obtained, supporting responses to complex questions.
[0075] like Figure 7 As shown, Figure 7The diagram illustrates the overall process provided in this application embodiment. The process involves obtaining the current visitor's statement. If the current visitor's statement is their first statement, features are extracted from the current visitor's statement using a first large language model to obtain a third mental map and a second reason summary text. A preset expert system is used to logically deduce the third mental map and the second reason summary text to obtain a second action suggestion. The second action suggestion is input into a preset second large language model to obtain a second response text. The current visitor's statement corresponding to the first statement and the second response text are used as historical dialogue records, and the third mental map is used as the first mental map. When the current visitor's statement is not their first statement, the system retrieves historical dialogue records and a first mental map. It then inputs the current visitor's statement, the first mental map, and the historical dialogue records into a first large language model to obtain a second mental map and a first cause summary text. A preset expert system is used to logically deduce the second mental map and the first cause summary text to obtain a first action suggestion. This first action suggestion is then input into a preset second large language model, which is a text generation model. The first response text and the current visitor's statement are added to the historical dialogue records. The second mental map is used as the first mental map. The process of retrieving the current visitor's statement, and retrieving historical dialogue records and the first mental map when the current visitor's statement is not their first statement, continues until the dialogue ends. This embodiment accurately captures information from the current visitor's statement, historical dialogue records, and the first mental map using a first large language model, then performs action suggestion deduction and response text generation, and adds the current visitor's statement and the first response text to the historical dialogue records. This allows the current question response to refer to historical dialogues, achieving a more accurate response to complex questions.
[0076] like Figure 8As shown, this application embodiment provides a question-and-answer device 100 based on a large language model. The device 100 first acquires the current visitor's statement through a data acquisition module 110. If the current visitor's statement is not their first statement, it acquires historical dialogue records and a first mental map. Acquiring historical dialogue records and the first mental map improves the coherence of the question-and-answer process, supporting answers to more complex questions. Then, the first data processing module 120 inputs the current visitor's statement, the first mental map, and the historical dialogue records into a first large language model to obtain a second mental map and a first cause summary text. The first large language model is a statement analysis model that analyzes the current visitor's expectations, which is beneficial for obtaining subsequent answer text. Finally, the second data processing module 130 uses a preset expert system to process the second mental map and the first cause summary text. Logical deduction is performed on the text to obtain the first action suggestion. Guided by a preset expert system, the answer can be more accurate. Then, the first action suggestion is input into the preset second language model using the third data processing module 140 to obtain the first answer text. The second language model is a text generation model. The first answer text conforms to the dialogue logic and has high accuracy. Finally, the first answer text and the current visitor's statement are added to the historical dialogue record by the loop question-and-answer module 150. The second mental map is used as the first mental map. The steps of obtaining the current visitor's statement are executed. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained until the dialogue ends. The answer to the current visitor's statement is based on the historical dialogue record. It can support answers to sufficiently complex questions, the dialogue is coherent, and the answer is highly accurate.
[0077] It should be noted that the data acquisition module 110 is connected to the first data processing module 120, the first data processing module 120 is connected to the second data processing module 130, the second data processing module 130 is connected to the third data processing module 140, and the third data processing module 140 is connected to the loop question-and-answer module 150. The above-mentioned question-and-answer method based on a large language model is applied to a question-and-answer device 100 based on a large language model. The question-and-answer device 100 accurately captures information from the current visitor's statement, historical dialogue records, and the first mental map through the first large language model, then performs action suggestion deduction and response text generation, and adds the current visitor's statement and the first response text to the historical dialogue records, thus achieving relatively accurate responses to complex questions.
[0078] It should also be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0079] This application also discloses an electronic device. (See reference...) Figure 9 , Figure 9 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. The electronic device 500 may include: at least one processor 501, at least one network interface 504, a user interface 503, a memory 505, and at least one communication bus 502.
[0080] The communication bus 502 is used to enable communication between these components.
[0081] The user interface 503 may include a display screen and a camera. Optionally, the user interface 503 may also include a standard wired interface and a wireless interface.
[0082] The network interface 504 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).
[0083] The processor 501 may include one or more processing cores. The processor 501 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 505, and by calling data stored in memory 505. Optionally, the processor 501 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 501 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 501 and may be implemented as a separate chip.
[0084] The memory 505 may include random access memory (RAM) or read-only memory. Optionally, the memory 505 may include a non-transitory computer-readable storage medium. The memory 505 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 505 may also be at least one storage device located remotely from the aforementioned processor 501. (Refer to...) Figure 9 The memory 505, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program based on a question-and-answer method using a large language model.
[0085] exist Figure 9In the illustrated electronic device 500, the user interface 503 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 501 can be used to call an application program based on a large language model stored in the memory 505. When executed by one or more processors 501, the electronic device 500 performs one or more methods as described in the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simplicity, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0086] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0087] In the various embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.
[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0089] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0090] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks, or optical disks.
[0091] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will readily conceive of those skilled in the art upon consideration of the specification and the disclosure of practical truths.
[0092] This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for question answering based on a large language model, characterized in that, The method includes: Obtain the current visitor's statement; if the current visitor's statement is not the first statement, obtain the historical dialogue record and the first mental map. The current visitor's statement, the first mental map, and the historical dialogue record are input into the first large language model to obtain the second mental map and the first cause summary text. The first large language model is a statement analysis model. Using a pre-set expert system, logical deduction is performed on the second mental map and the first cause summary text to obtain the first action suggestion; The first action suggestion is input into the preset second language model to obtain the first answer text. The second language model is a text generation model. The first response text and the current visitor's statement are added to the historical dialogue record. The second mental map is used as the first mental map. The steps of obtaining the current visitor's statement are executed. If the current visitor's statement is not the first statement, the historical dialogue record and the first mental map are obtained until the dialogue ends.
2. The method of claim 1, wherein, If the current visitor's statement is their first statement, the method further includes: The first large language model is used to extract features from the current visitor's statements to obtain the third psychological map and the second reason summary text; By using a pre-set expert system to logically deduce the third psychological map and the second cause summary text, a second action suggestion is obtained; The second action suggestion is input into the preset second language model to obtain the second response text; The current visitor's statement corresponding to the first statement and the second response text are used as historical dialogue records, and the third mental map is used as the first mental map.
3. The method of claim 1, wherein, The step of using a preset expert system to logically deduce the first action suggestion from the second psychological map and the first cause summary text includes: A dialogue process framework is constructed based on the second mental map and the summary text of the first cause, and multiple inference combinations are obtained based on the dialogue process framework; Based on the various deduction combinations and the guidance rules of the expert system, the first action suggestion is obtained.
4. The method of claim 1, wherein, The method further includes training steps for the first large language model and the second large oracle model, the training steps including: Obtain a historical corpus dataset, which includes multiple visitor corpus data and first mental map labels and first cause summary labels corresponding to the multiple visitor expected data; Input the corpus data of each visitor, the first mental map label, and the first reason summary label into the preset third language model, and train the third language model to obtain the first language model. The historical corpus dataset also includes counselor corpus corresponding to each of the client's expected data, and a first action label corresponding to the counselor corpus; The counselor corpora and the first action labels are input into a preset fourth language model, and the fourth language model is trained to obtain the second language model.
5. The method of claim 4, wherein, After inputting the first action suggestion into a preset second language model to obtain the first response text, the method further includes: Based on the preset time, collect the current visitor's statement and the first response text; The current visitor's statement and the first response text are labeled to obtain the second psychological map label and the second cause summary label corresponding to the current visitor's statement, as well as the second action label corresponding to the first response text; The current visitor's statement, the second mental map label, the second reason summary label, the first answer text, and the second action label are added to the historical corpus dataset to obtain an updated corpus dataset; The parameters of the first and second largest language models are adjusted using the updated corpus dataset.
6. The method of claim 4, wherein, The first mental map label is obtained through the following steps: The manual annotation of the speech data of each of the visitors is obtained, and the manual annotation includes events, emotions, and causal relationships; By treating the events and emotions as entities and causal relationship features as connecting edges, the first psychological graph label is constructed.
7. The method of claim 4, wherein, The step of inputting the various counselor corpora and the first action labels into a preset fourth language model, and training the fourth language model to obtain the second language model, includes: The consultant's corpus and the first action label are input into the preset fourth language model to obtain the predicted response statement; The value of the loss function is calculated based on the predicted response statement and the first action label; The parameters of the fourth language model are adjusted using the value of the loss function to obtain the second language model.
8. A large language model-based question answering device, characterized by, The device includes: The data acquisition module (110) is used to acquire the current visitor's statement, and if the current visitor's statement is not the first statement, to acquire the historical dialogue record and the first mental map; The first data processing module (120) is used to input the current visitor's statement, the first mental map and the historical dialogue record into the first large language model to obtain the second mental map and the first cause summary text. The first large language model is a statement analysis model. The second data processing module (130) is used to perform logical deduction on the second psychological map and the first cause summary text using a preset expert system to obtain the first action suggestion; The third data processing module (140) is used to input the first action suggestion into a preset second language model to obtain the first answer text, wherein the second language model is a text generation model; The loop question-and-answer module (150) is used to add the first answer text and the current visitor's statement to the historical dialogue record, use the second mental map as the first mental map, and execute the steps of obtaining the current visitor's statement, and obtaining the historical dialogue record and the first mental map when the current visitor's statement is not the first statement, until the dialogue ends.
9. An electronic device, comprising: The device includes a processor (501), a memory (505), a user interface (503), a communication bus (502), and a network interface (504). The processor (501), the memory (505), the user interface (503), and the network interface (504) are respectively connected to the communication bus (502). The memory (505) is used to store instructions. The user interface (503) and the network interface (504) are used to communicate with other devices. The processor (501) is used to execute the instructions stored in the memory (505) so that the electronic device (500) performs the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.