Conversation generation method and device, electronic equipment, storage medium and computer program product
By combining knowledge graphs and self-coding large language models, using multi-layer instruction templates and confidence scoring mechanisms, the problems of the authenticity and reliability of language models and insufficient language capabilities of the knowledge graph are solved, and efficient and natural dialogue generation is achieved.
Patent Information
- Application Number
- CN202510020944.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-02
AI Technical Summary
In the prior art, the knowledge authenticity and reliability of language models are difficult to guarantee, and the text generated by the knowledge graph in dialogue questions and answers is insufficient, so it is impossible to generate natural and smooth text in real time.
By combining knowledge graphs and self-coded large language model (LLM), a multi-layer instruction template and confidence scoring mechanism are adopted to generate entity recognition, disambiguation and relationship recognition to generate the most relevant entities and relationship types, thereby building triple information in natural language for generating smooth dialogue.
It improves the authenticity and reliability of knowledge in the dialogue generation system, enhances language communication skills, and can generate natural and smooth texts in real time, solving the problems of inconsistent knowledge and insufficient language skills.
Smart Images

Figure CN120045658A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a dialogue generation method, apparatus, electronic device, storage medium, and computer program product. Background Art
[0002] In order to generate a dialogue for user input text, when using a related large language model, the authenticity and reliability of the knowledge cited by the large language model cannot be guaranteed, so there is a problem that the knowledge in the large language model is inconsistent with the knowledge in the real world; when using a knowledge graph, the knowledge graph returns static text in dialogue answering and cannot generate natural and fluent text in real time, resulting in a problem of insufficient language ability. Summary of the Invention
[0003] To solve the related technical problems, embodiments of this application provide a dialogue generation method, apparatus, electronic device, storage medium, and computer program product.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] Embodiments of this application provide a dialogue generation method, the method including:
[0006] Based on a set first instruction template, convert first input information into a first language instruction, and call a first autoencoder large language model (LLM) to process the first language instruction to obtain first output information; the first input information includes user input text; the first output information includes one or more first entities and corresponding entity types in the user input text;
[0007] Based on a set second instruction template, convert second input information into a second language instruction, and call a second autoencoder LLM to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information, and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained by performing single-hop retrieval on the first output information in a knowledge graph; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities;
[0008] Based on a set third instruction template, convert third input information into a third language instruction, and call a third autoencoder LLM to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to the one or more second entities in the second output information;
[0009] Generate a response statement regarding the user input text based on the second output information and the third output information.
[0010] In the above solution, the first input information further includes entity type description information; and / or,
[0011] The third input information further includes the entity type description information and relationship type description information; where
[0012] The entity type description information and / or the relationship type description information are determined based on the prior knowledge of the knowledge graph.
[0013] In the above solution, the entity type description information and / or the relationship type description information are updated as the knowledge graph is updated.
[0014] In the above solution, generating a response statement regarding the user input text based on the second output information and the third output information includes:
[0015] Invoke the knowledge graph to process the second output information and the third output information to obtain triple information;
[0016] Construct a prior knowledge statement based on the triple information;
[0017] Generate a response statement regarding the user input text based on the prior knowledge statement and the user input text.
[0018] In the above solution, before constructing a prior knowledge statement based on the triple information, the method further includes:
[0019] Input the triple information into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information; update the triple information based on the relevant knowledge information.
[0020] In the above solution, the first output information further includes a first confidence score for each of the first entities; the first confidence score is used to judge the credibility of the entity type corresponding to the first entity;
[0021] The second output information further includes a second confidence score for each of the second entities; the second confidence score is used to judge the relevance between each of the second entities in the candidate information and the corresponding first entity;
[0022] The third output information further includes a third confidence score; the third confidence score is used to judge the relevance between each of the second entities and different relationship types.
[0023] In the above solution, after obtaining the triple information, the method further includes:
[0024] Calculating the average of the second confidence score and the corresponding third confidence score for each triple in the triple information, and sorting all the triples in the triple information according to the average to obtain a sorting result;
[0025] Based on the sorting result, determining whether one or more triples in the triple information can be used to construct the prior knowledge statement.
[0026] In the above solution, constructing the prior knowledge statement based on the triple information includes:
[0027] Processing the triple information based on a set fourth instruction template to obtain a fourth language instruction, and calling a first autoregressive LLM to process the fourth language instruction to obtain the prior knowledge statement; the prior knowledge statement represents a natural language with smooth sentences based on the triple information.
[0028] In the above solution, generating the answer statement about the user input text based on the prior knowledge statement and the user input text includes:
[0029] Calling a second autoregressive LLM, taking the prior knowledge statement and the user input text as inputs, and generating the answer statement.
[0030] An embodiment of the present application further provides a dialogue generation device, including:
[0031] A first processing unit, configured to convert first input information into a first language instruction based on a set first instruction template, and call a first autoencoding large language model LLM to process the first language instruction to obtain first output information; the first input information includes user input text; the first output information includes one or more first entities and corresponding entity types in the user input text;
[0032] A second processing unit, configured to convert second input information into a second language instruction based on a set second instruction template, and call a second autoencoding LLM to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information, and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained by performing single-hop retrieval on the first output information in a knowledge graph; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities;
[0033] A third processing unit, configured to convert third input information into a third language instruction based on a set third instruction template, and call a third autoencoding large language model (LLM) to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to one or more second entities in the second output information.
[0034] A first generation unit, configured to generate a response statement regarding the user input text based on the second output information and the third output information.
[0035] An embodiment of this application further provides an electronic device, including a memory and a processor, where the memory stores a computer program that can run on the processor, and the processor is configured to execute the method for dialogue generation according to any one of the above.
[0036] An embodiment of this application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of the above are implemented.
[0037] An embodiment of this application further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method according to any one of the above are implemented.
[0038] In the dialogue generation method, device, electronic device, storage medium, and computer program product provided by the embodiments of the present application, first, based on a set first instruction template, the first input information is converted into a first language instruction, and a first auto-encoding large language model (LLM) is called to process the first language instruction to obtain first output information. The first input information includes user input text, and the first output information includes one or more first entities and corresponding entity types in the user input text. Secondly, based on a set second instruction template, the second input information is converted into a second language instruction, and a second auto-encoding LLM is called to process the second language instruction to obtain second output information. The second input information includes the user input text, the first output information, and candidate information. Thirdly, based on a set third instruction template, the third input information is converted into a third language instruction, and a third auto-encoding LLM is called to process the third language instruction to obtain third output information. The third input information includes the second output information and the candidate information. Since the candidate information includes one or more second entities and corresponding relationship types obtained by performing single-hop retrieval on the first output information in a knowledge graph, one or more second entities most relevant to the one or more first entities in the candidate information can be included in the second output information, and one or more relationship types most relevant to the one or more second entities in the second output information are included in the third output information. Therefore, based on the second output information and the third output information, that is, based on the most relevant one or more second entities and the most relevant one or more relationship types, the most accurate and reliable response statement regarding the user input text can be generated. Among them, by combining the method of retrieving in a knowledge graph and the method of constructing an instruction template and feeding it into an auto-encoding LLM, not only can entity recognition, entity disambiguation, and relationship recognition be completed, improving the authenticity and reliability of cited knowledge, but also it helps to subsequently construct triples for generating fluent conversations. Therefore, the problems of knowledge inconsistency and insufficient language communication ability in the dialogue generation method can be solved simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a schematic flowchart of a dialogue generation method according to an embodiment of the present application;
[0040] Figure 2 It is a schematic flowchart of another dialogue generation method according to an embodiment of the present application;
[0041] Figure 3 It is a schematic flowchart of a third dialogue generation method according to an embodiment of the present application;
[0042] Figure 4 It is a schematic structural diagram of a dialogue generation device according to an embodiment of the present application;
[0043] Figure 5 Schematic diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0044] Large language models are neural network models that use large-scale unlabeled text data for self-supervised pre-training. They have the ability to understand context and generate text, and can adapt to a variety of natural language tasks after fine-tuning training. However, there are problems with the inconsistency of knowledge in the real world in related dialogue large models, resulting in factual errors in the generated text. For example, some dialogue large models introduce a search engine or a file upload interface to introduce more context as prior knowledge when the model generates answers. However, the authenticity of the retrieved content information of the search engine cannot be guaranteed, and synonyms or ambiguities in the text can also lead to uncontrollable search results. Moreover, users need to consult documents in advance and ensure the quality of the knowledge in the documents. There are problems with the untrustworthy knowledge sources of the documents uploaded by users who are not familiar with the fields involved in the documents.
[0045] A knowledge graph is an information model that describes concepts, entities, and the relationships between them in the objective world in a structured form. However, the auto-encoding pre-training model used by the dialogue system using the knowledge graph cannot perform text generation tasks. The pipeline of the system finally returns static text stored in the corpus or knowledge graph and cannot generate natural and fluent text in real time during dialogue answering. Further, the application scenarios of the dialogue answering system based on knowledge retrieval ranking are restricted by the corpus or knowledge graph.
[0046] In summary, the related dialogue answering systems cannot solve the problems of knowledge inconsistency in language large models and the lack of language ability in knowledge graphs at the same time.
[0047] In the dialogue generation method, device, electronic device, storage medium, and computer program product provided by the embodiments of the present application, first, based on a set first instruction template, the first input information is converted into a first language instruction, and a first auto-encoding large language model (LLM) is called to process the first language instruction to obtain a first output information. The first input information includes user input text, and the first output information includes one or more first entities and corresponding entity types in the user input text. Secondly, based on a set second instruction template, the second input information is converted into a second language instruction, and a second auto-encoding LLM is called to process the second language instruction to obtain a second output information. The second input information includes the user input text, the first output information, and candidate information. Thirdly, based on a set third instruction template, the third input information is converted into a third language instruction, and a third auto-encoding LLM is called to process the third language instruction to obtain a third output information. The third input information includes the second output information and the candidate information. Since the candidate information includes one or more second entities and corresponding relationship types obtained by performing a single-hop retrieval on the first output information in a knowledge graph, one or more second entities most relevant to the one or more first entities in the candidate information can be included in the second output information, and one or more relationship types most relevant to the one or more second entities in the second output information can be included in the third output information. Therefore, based on the second output information and the third output information, that is, based on the most relevant one or more second entities and the most relevant one or more relationship types, the most accurate and credible response statement regarding the user input text can be generated. Among them, by combining the method of retrieving in a knowledge graph and the method of constructing an instruction template and feeding it into an auto-encoding LLM, not only can entity recognition, entity disambiguation, and relationship recognition be completed, improving the authenticity and reliability of cited knowledge, but also it helps to construct triples for generating fluent conversations subsequently. Therefore, the problems of knowledge inconsistency and insufficient language communication ability in the dialogue generation method can be solved simultaneously.
[0048] An embodiment of the present application provides a dialogue generation method. Refer to Figure 1 , this method includes:
[0049] Step 101: Based on a set first instruction template, convert the first input information into a first language instruction, and call a first auto-encoding LLM to process the first language instruction to obtain a first output information.
[0050] Here, the first input information includes user input text; the first output information includes one or more first entities and corresponding entity types in the user input text.
[0051] Step 102: Based on the set second instruction template, convert the second input information into a second-language instruction, and call the second auto-encoding LLM to process the second-language instruction to obtain second output information.
[0052] Here, the second input information includes the user input text, the first output information, and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained by performing single-hop retrieval on the first output information in a knowledge graph; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities.
[0053] Step 103: Based on the set third instruction template, convert the third input information into a third-language instruction, and call the third auto-encoding LLM to process the third-language instruction to obtain third output information.
[0054] Here, the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to the one or more second entities in the second output information.
[0055] Step 104: Generate a response statement regarding the user input text based on the second output information and the third output information.
[0056] In the embodiments of the present application, an instruction template refers to a preset natural sentence with words or sentences to be filled in. Exemplarily, the first instruction template can be "Find all named entities that may exist in the $ to-be-recognized text", where it can be stipulated that "$ to-be-recognized text = $ first input information", that is, the "first input information" can be used to replace the "to-be-recognized text", and "$" is used to mark the position information to be filled.
[0057] In the embodiments of the present application, exemplarily, when using the "first input information" to replace the "to-be-recognized text" in the first instruction template, when the first input information is "Are Lu Xun and Zhou Shuren the same person?", the first-language instruction can be obtained as "Find all named entities that may exist in 'Are Lu Xun and Zhou Shuren the same person?'".
[0058] In the embodiments of the present application, the first auto-encoding LLM, the second auto-encoding LLM, and the third auto-encoding LLM can be the same language model or different language models. In actual applications, selection can be made according to the input concurrency of the language model. Before any of the above auto-encoding LLMs is put into use, the language model can be trained by fine-tuning the parameters based on labeled data.
[0059] In the embodiments of the present application, the first auto-encoding large language model (LLM) can be invoked to process the first language instruction, and first output information is obtained. The first output information includes one or more first entities and corresponding entity types in the user input text. Exemplarily, when the first language instruction is "find all possible named entities in 'Is Lu Xun and Zhou Shuren the same person?'", invoking the first auto-encoding LLM to process this first language instruction, the first output information obtained can be "Lu Xun - person name; Zhou Shuren - person name", that is, the first entity "Lu Xun" and the first entity "Zhou Shuren" are obtained, and the entity type corresponding to "Lu Xun" is "person name", and the entity type corresponding to "Zhou Shuren" is "person name". Therefore, the purpose of obtaining the first output information is to complete entity recognition in the user input text.
[0060] In the embodiments of the present application, single-hop retrieval refers to a retrieval method in a knowledge graph where, starting from a starting node (entity), the target node (entity) can be found through one edge (relationship). Exemplarily, when the starting node (entity) is "Yao Ming", the relationship is "born in", and the target node (entity) is "Shanghai", that is, starting from the starting node "Yao Ming" and through the relationship "born in", the target node "Shanghai" can be directly found, which is the single-hop retrieval process.
[0061] In the embodiments of the present application, during the dialogue generation process, relevant retrieval or processing can also be combined with the knowledge graph. The above first output information can be input into the knowledge graph for single-hop retrieval to obtain candidate information. The candidate information includes one or more second entities and corresponding relationship types. Exemplarily, when the first output information is "Lu Xun - person name; Zhou Shuren - person name", after performing single-hop retrieval of this first output information in the knowledge graph, the candidate information obtained can be "1. Lu Xun (original name, native place, birthplace, occupation); 2. 'Lu Xun' (author, publisher, publication time); 3. 'Lu Xun' (director, starring actors, release time); 4. Zhou Shuren (pen name, original name, native place, birthplace, occupation)". Among them, for the second candidate information item, the second entity is "'Lu Xun'", and the corresponding relationship type is "author, publisher, publication time", that is, the author, publisher, and publication time of the book 'Lu Xun'.
[0062] In the embodiments of the present application, after entity recognition is completed, in order to perform entity disambiguation, that is, to find a closest true mapping for each entity in the user input text in the knowledge graph, not only can the candidate information retrieved from the knowledge graph be used, but also a second language instruction can be constructed by combining the first output information and the user input text. Specifically, based on the set second instruction template, the second input information is converted into a second language instruction, that is, the user input text, the first output information, and the candidate information in the second input information can be brought into the second instruction template to obtain the second language instruction.
[0063] In the embodiments of the present application, after the second language instruction is obtained, the second self-encoding LLM is called to process the second language instruction to obtain a second output information. The second output information includes one or more second entities in the candidate information that are most relevant to one or more first entities. Exemplarily, when the candidate information is "1. Lu Xun (original name, native place, birthplace, occupation); 2. "Lu Xun" (author, publisher, publication time); 3. "Lu Xun" (director, starring actors, release time); 4. Zhou Shuren (pen name, original name, native place, birthplace, occupation)", the second entities that are most relevant to the first entity "Lu Xun" and the first entity "Zhou Shuren" are the "Lu Xun" in item 1 of the candidate information and the "Zhou Shuren" in item 4.
[0064] In the embodiments of the present application, in order to complete relationship recognition, that is, to find the relationship type most relevant to the second entity, a third language instruction can be converted based on the set third instruction template with the third input information, and the third self-encoding LLM is called to process the third language instruction to obtain a third output information. Among them, the third input information includes the second output information and the candidate information. In practical applications, the third input information may further include the replaced input text, which can be obtained by replacing the relevant entities in the user input text with the second output information. Exemplarily, when the user input text is "Is Lu Xun and Zhou Shuren the same person?", the replaced input text can be "<Lu Xun-person name> and <Zhou Shuren-person name> the same person?".
[0065] In an embodiment of the present application, after obtaining the third language instruction, the third auto-encoding large language model (LLM) is called to process the third language instruction to obtain the third output information. The third output information includes one or more relationship types that are most relevant to one or more second entities in the second output information. Exemplarily, when the second output information is the second entities of the 1st and 4th items of the candidate information "1. Lu Xun (original name, native place, birthplace, occupation); 2. *Lu Xun* (author, publisher, publication time); 3. *Lu Xun* (director, lead actor, release time); 4. Zhou Shuren (pen name, original name, native place, birthplace, occupation)", then the relationship type most relevant to "Lu Xun" in the 1st item is "original name", and the relationship types most relevant to "Zhou Shuren" in the 4th item are "pen name" and "original name". Therefore, the third output information can be "Lu Xun - original name; Zhou Shuren - pen name; Zhou Shuren - original name".
[0066] In an embodiment of the present application, after obtaining the second output information and the third output information, since the second output information includes one or more second entities that are most relevant to one or more first entities in the candidate information, and the third output information includes one or more relationship types that are most relevant to one or more second entities in the second output information, an answer statement regarding the user input text can be generated based on the second output information and the third output information, using the most relevant entities obtained after entity disambiguation and the most relevant relationship types obtained after relationship recognition.
[0067] In one embodiment, the first input information further includes entity type description information; and / or,
[0068] the third input information further includes the entity type description information and relationship type description information; wherein,
[0069] the entity type description information and / or the relationship type description information are determined based on the prior knowledge of the knowledge graph.
[0070] In the embodiments of the present application, the entity type description information and / or the relationship type description information are determined based on the prior knowledge of the knowledge graph. The entity type description information refers to the detailed description or explanatory information of different entities; the relationship type description information refers to the detailed description or explanatory information of different relationships. Exemplarily, the entity type description information may be "Person Name: The name of a person. Location: The geographical location in the real world", that is, "Person Name" can be described as the name of a person, and "Location" can be described as the geographical location in the real world; the relationship type description information may be "Original Name: The name that a person entity has ever used, defaulting to Chinese name or translated name. Birth Place: The birth place of a person entity", that is, "Original Name" can be described as the name that a person entity has ever used, defaulting to Chinese name or translated name, and "Birth Place" can be described as the birth place of a person entity. The entity type description information and / or the relationship type description information are determined based on the prior knowledge of the knowledge graph.
[0071] In the embodiments of the present application, the entity type description information is further included in the first input information, which can make full use of the prior knowledge of the knowledge graph to facilitate the first auto-encoding LLM to better complete the recognition of the first entity; the entity type description information and the relationship type description information are further included in the third input information, which can make full use of the prior knowledge of the knowledge graph to facilitate the third auto-encoding LLM to more accurately complete the recognition or judgment of the relationship type.
[0072] In one embodiment, the entity type description information and / or the relationship type description information are updated as the knowledge graph is updated.
[0073] In the embodiments of the present application, when the knowledge graph is updated, the entity type description information and / or the relationship type description information obtained based on the prior knowledge of the knowledge graph will also be updated, so as to more timely and effectively help the auto-encoding model understand the tasks of entity recognition or relationship recognition.
[0074] In one embodiment, generating a response statement about the user input text based on the second output information and the third output information includes:
[0075] Invoking the knowledge graph to process the second output information and the third output information to obtain triple information;
[0076] Constructing a prior knowledge statement based on the triple information;
[0077] Generating a response statement about the user input text based on the prior knowledge statement and the user input text.
[0078] In the embodiments of the present application, in order to more accurately generate a response statement for the user input text, after obtaining the second output information and the third output information, a knowledge graph may also be invoked, that is, the knowledge graph is invoked to process the second output information and the third output information to obtain triple information. In a knowledge graph, a triple is an ordered set composed of three elements, namely a subject, a predicate, and an object. Exemplarily, the triple of "an apple is a fruit" can be represented as (apple, is, fruit).
[0079] In the embodiments of the present application, since the second output information includes one or more second entities that are most relevant to one or more first entities in the candidate information, and the third output information includes one or more relationship types that are most relevant to the one or more second entities in the second output information, one or more pairs can be formed by combining the second output information and the corresponding third output information. After entity retrieval in the knowledge graph, one or more triples are obtained. Exemplarily, after entity retrieval of "Lu Xun - original name - {}" in the knowledge graph, the triple "Lu Xun - original name - {Zhou Shuren}" can be obtained, where "{}" represents the entity that needs to be retrieved and filled.
[0080] In the embodiments of the present application, constructing a prior knowledge statement based on the triple information means converting the triple information into a language or statement in a natural process as the prior knowledge of the user input text, which can thus help generate a response statement. Exemplarily, when the triple is "Lu Xun - original name - {Zhou Shuren}", the prior knowledge statement can be "The original name of Lu Xun is Zhou Shuren." Then, based on this prior knowledge statement and the user input text, a response statement about the user input text is generated. Exemplarily, the response statement can be "They are the same person because the original name of Lu Xun is Zhou Shuren."
[0081] In the embodiments of the present application, by invoking the knowledge graph to process the second output information and the third output information to obtain triple information, constructing a prior knowledge statement based on the triple information, and finally generating a response statement about the user input text based on the prior knowledge statement and the user input text, all the structured knowledge involved in the user input text can be processed into a smooth unstructured natural language text, that is, an unstructured prior knowledge text, which is beneficial to knowledge alignment with the real world when generating dialogue text.
[0082] In one embodiment, before constructing the prior knowledge statement based on the triple information, the method further includes:
[0083] Inputting the triple information into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information; updating the triple information based on the relevant knowledge information.
[0084] In the embodiments of the present application, after processing the second output information and the third output information through the knowledge graph to obtain triple information, the obtained triple information can be input into the knowledge graph again for single-hop retrieval to obtain relevant knowledge information of the triple information. In practical applications, the obtained triple information can also be input into the knowledge graph for two-hop retrieval to obtain relevant knowledge information of the triple information. Among them, the relevant knowledge information can be one or more relevant triples. Exemplarily, when the triple is "Lu Xun - original name - Zhou Shuren", the relevant knowledge information obtained by single-hop retrieval can be "Lu Xun - ancestral home - Zhengyang County, Henan Province" and / or "Lu Xun - date of birth - September 25, 1881" and / or "Lu Xun - mother - Lu Rui". Based on the relevant knowledge information, the triple information can be updated. In practical applications, the relevant knowledge information can be added to the triple information.
[0085] In the embodiments of the present application, after inputting the triple information into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information, since the relevant knowledge information is retrieved from the knowledge graph, the relevant knowledge information is consistent with the knowledge in the real world and the source is credible. Moreover, based on the relevant knowledge information, updating the triple information can also make the triple information contain richer information, so as to facilitate the generation of richer answer statements.
[0086] In one embodiment, the first output information further includes a first confidence score for each of the first entities; the first confidence score is used to judge the credibility of the entity type corresponding to the first entity;
[0087] The second output information further includes a second confidence score for each of the second entities; the second confidence score is used to judge the relevance between each of the second entities in the candidate information and the corresponding first entity;
[0088] The third output information further includes a third confidence score; the third confidence score is used to judge the relevance between each of the second entities and different relationship types.
[0089] In the embodiments of the present application, the first output information may further include a first confidence score for each first entity. Exemplarily, the first output information includes "Lu Xun - personal name - 1.00", where the score "1.00" is used to judge the credibility score of the entity type "personal name" corresponding to the entity "Lu Xun" in the user input text, and "1.00" can represent 100% credibility.
[0090] In the embodiments of the present application, the second output information further includes the second confidence score of each second entity. Exemplarily, when the candidate information is "1. Lu Xun (original name, native place, birthplace, occupation); 2. *Lu Xun* (author, publisher, publication time); 3. *Lu Xun* (director, leading actor, release time); 4. Zhou Shuren (pen name, original name, native place, birthplace, occupation)", and the second entity is the entity "Zhou Shuren" in the 4th item, then the second output information may include "4. Zhou Shuren - 1.00", where the score "1.00" is used to judge the relevance between the second entity "4. Zhou Shuren" and the corresponding first entity "Zhou Shuren", and "1.00" may represent 100% relevance. Among them, the first entity "Zhou Shuren" is the entity of "Zhou Shuren - personal name" in the first output information.
[0091] In the embodiments of the present application, the third output information further includes the third confidence score. Exemplarily, the third output may include "Zhou Shuren - original name - 0.96", where the score "0.98" is used to judge the relevance between the second entity "Zhou Shuren" and the relationship type "original name"; the third output may also include "Zhou Shuren - pen name - 0.97", where the score "0.97" is used to judge the relevance between the second entity "Zhou Shuren" and the relationship type "pen name". Therefore, the third confidence score is used to judge the relevance between each second entity and different relationship types.
[0092] In the embodiments of the present application, since the first output information further includes the first confidence score of each first entity, the second output information further includes the second confidence score of each second entity, and the third output information further includes the third confidence score, the most accurate entity corresponding to the entity in the user input text and the most relevant relationship type can be found more precisely in the knowledge graph by calculating the scores.
[0093] In one embodiment, after obtaining the triple information, the method further includes:
[0094] Calculating the average value of the second confidence score and the corresponding third confidence score corresponding to each triple in the triple information, and sorting all the triples in the triple information according to the average value to obtain a sorting result;
[0095] Based on the sorting result, determining whether one or more triples in the triple information can be used to construct the prior knowledge statement.
[0096] In the embodiments of the present application, after obtaining the triple information, each triple in the triple information can also be scored. Since the triple information is obtained by processing the second output information and the third output information by invoking the knowledge graph, and the second output information includes a second confidence score, and the third output information includes a third confidence score, the second confidence score corresponding to each triple and the corresponding third confidence score can be averaged to obtain an average value. Exemplarily, if the second output information is "Lu Xun - 0.99" and the third output information is "Lu Xun - original name - 0.98", then the average value of the corresponding second confidence score (0.99) and the corresponding third confidence score (0.98) of the obtained triple "Lu Xun - original name - Zhou Shuren" is 0.985. After that, similar score calculations are also performed for each triple, and all triples are sorted according to the average value calculated for each triple to obtain a sorting result.
[0097] In the embodiments of the present application, based on the sorting result of the above triples, it can be determined whether one or more triples in the triple information can be used to construct the prior knowledge statement. In practical applications, a set number of triples with the highest average value can be selected to construct the prior knowledge statement; or triples with an average value higher than 0.9 can be used to construct the prior knowledge statement. When the number of triples is insufficient, a batch of triples with an average score not lower than 0.7 can also be supplemented according to the average score as an additional supplement.
[0098] In the embodiments of the present application, by calculating the average value of the second confidence score corresponding to each triple in the triple information and the corresponding third confidence score, and sorting all the triples in the triple information according to this average value to obtain a sorting result, and based on this sorting result, determining whether one or more triples in the triple information can be used to construct the prior knowledge statement, the triple information most suitable for constructing the prior knowledge statement can be screened out, which helps to generate a more accurate answer statement.
[0099] In one embodiment, constructing a prior knowledge statement based on the triple information includes:
[0100] Processing the triple information based on a set fourth instruction template to obtain a fourth language instruction, and invoking a first autoregressive LLM to process the fourth language instruction to obtain the prior knowledge statement; the prior knowledge statement represents a natural language with smooth sentences based on the triple information.
[0101] In the embodiments of the present application, since the triple information is structured knowledge, it needs to be converted into unstructured natural language, that is, an unstructured prior knowledge statement. Therefore, the triple information can be processed based on a set fourth instruction template, that is, the triple information is brought into the fourth instruction template to obtain a fourth language instruction, and the first autoregressive LLM is called to process the fourth language instruction. The first autoregressive LLM can perform few-shot or even zero-shot prompt learning for processing, and splice the obtained structured knowledge into a text instruction to obtain the prior knowledge statement.
[0102] In the embodiments of the present application, by processing the triple information based on a set fourth instruction template to obtain a fourth language instruction, and calling the first autoregressive LLM to process the fourth language instruction to obtain the prior knowledge statement, the structured knowledge can be better converted into a fluent unstructured natural language text. This can not only align the first autoregressive LLM with the real world in dialogue text generation, but also make the obtained prior knowledge statement and the input text modality-consistent, so as to better perform dialogue generation.
[0103] In one embodiment, generating the response statement about the user input text based on the prior knowledge statement and the user input text includes:
[0104] Calling a second autoregressive LLM, using the prior knowledge statement and the user input text as inputs, to generate the response statement.
[0105] In the embodiments of the present application, the second autoregressive LLM and the first autoregressive LLM can be the same autoregressive language model or not, depending on actual needs. After obtaining the prior knowledge statement, the second autoregressive LLM can be called, using the prior knowledge statement and the user input text as inputs, to generate the response statement.
[0106] In the embodiments of the present application, since the prior knowledge statement and the user input text are modality-consistent and both are natural language statements, the second autoregressive LLM can be called, using the prior knowledge statement and the user input text as inputs, to generate the response statement. Through this method, the second autoregressive LLM can better generalize and summarize prior knowledge, avoiding semantic loss when directly using structured knowledge as prompt information to the second autoregressive LLM, and making the relevant domain knowledge of the model consistent with the real world while ensuring the dialogue communication ability of the model.
[0107] The following further describes the present application in detail with application embodiments.
[0108] This application proposes an application embodiment, which proposes a dialogue question-answering system based on a knowledge graph and a large language model.
[0109] The system provided by this application embodiment is mainly divided into three major modules, namely: 1. Information extraction module; 2. Knowledge text generation module; 3. Dialogue generation module based on knowledge text. The above three modules will be described in detail below.
[0110] 1. Information extraction module
[0111] The purpose of the information extraction module is to extract the head entity and candidate relationships involved in the user input text, mine the tail entity in the knowledge graph, obtain structured knowledge, and send it to the dialogue generation module as the background knowledge of the input text. The information extraction module can be further divided into an entity recognition sub-module, an entity disambiguation sub-module, and a relationship recognition sub-module. As Figure 2 shown below, these 3 sub-modules will be described in detail.
[0112] (1) Entity recognition sub-module
[0113] In this sub-module, the user input text is concatenated with the entity recognition instruction template and the prior knowledge of entity types as a task prompt, and sent to the auto-encoding LLM model to complete the named entity recognition task, and output the entity position and type. Among them, the $user input text can be "Are Lu Xun and Zhou Shuren the same person?"; the first instruction template is "$task name = named entity recognition task; $task description text = please find all possible named entities in the $text to be recognized, and according to the $entity type description, judge the entity types they belong to and give the corresponding confidence scores; $text to be recognized = $user input text"; the $entity type description can be "person name: task name; location: geographical location in the real world", and the first output information is "Lu Xun - person name - 1.00; Zhou Shuren - person name - 1.00". The auto-encoding LLM represents the input text combined with the prompt template as a feature vector, and decodes and classifies it based on a downstream sequence labeling model specific to the named entity recognition task to predict the corresponding positions and types of entities in the text. The auto-encoding LLM needs to be fine-tuned based on labeled data before being put into use. The fine-tuning algorithm uses the P-Tuning V2 algorithm to calculate the loss and optimize the model parameters corresponding to the partial positions of the prompt text based on the loss function and the correct labeling results of the data.
[0114] During the instruction prompt and construction of the first instruction template, the entity recognition sub-module will maintain a prior knowledge template that combines common entity types and corresponding descriptions in the knowledge graph. The common entity types are obtained by counting the frequencies of the corresponding types of important entity nodes in the knowledge graph. The important nodes can be given based on knowledge graph node importance evaluation models such as GENI. The entity type descriptions can be directly retrieved from the ontology construction information. When the knowledge graph is updated, this prior knowledge template will also change. In addition, the opinions of experts in various fields can be referred to, and some common entity types can be specified for some special fields. These prior knowledge can better help the auto-encoder model understand the named entity recognition task. For the above user input text, the first output information can be "Lu Xun - person name - 1.00; Zhou Shuren - person name - 1.00".
[0115] (2) Entity disambiguation sub-module
[0116] According to the first output information, that is, according to the obtained entities and types, the knowledge graph retrieval module is called to search for entities in the graph that may be the same as or close to the text. For each entity to be disambiguated, a fixed type and number of single-hop relationship edges are selected to obtain $ candidate information. This candidate information is the candidate entity and relationship. For the above user input text, this candidate information is "1. Lu Xun (original name, native place, birthplace, occupation); 2. *Lu Xun* (author, publisher, publication time); 3. *Lu Xun* (director, starring actors, release time); 4. Zhou Shuren (pen name, original name, native place, birthplace, occupation)". The candidate entity and relationship need to be sent to the entity disambiguation sub-module together. $ The entity and entity type are "Lu Xun - person name; Zhou Shuren - person name".
[0117] The candidate information will be combined with the entity and type to be disambiguated through a constructed instruction template prompt template, that is, the second instruction template, to form an entity disambiguation instruction, and then sent to the auto-encoder LLM together with the user input text to predict the correct entity.
[0118] For the above user input text, the second instruction template is "$ task name = entity disambiguation task; $ task description text = Please find the most accurate entity corresponding to the entity in the text in the knowledge graph based on the provided text, the entity and entity type in the text, and the list of candidate entities and relationships in the knowledge graph, and give the corresponding confidence score. The entity and entity type in the text are $ entity and entity type, and the list of candidate entities and relationships in the knowledge graph is $ candidate information; $ text to be recognized = $ user input text".
[0119] The fine-tuning training process of the auto-encoding LLM is similar to entity recognition. Only the downstream task model needs to be modified to a classification model to perform supervised training on the corresponding dataset. Further, the relationship edges of the candidate entities are used in this embodiment to assist in explaining the differences between this entity and other entities. The selected relationship types are obtained through the prior knowledge templates maintained in the relationship recognition sub-module.
[0120] For the above user input text, the second output information is "1. Lu Xun - 0.99; 4. Zhou Shuren - 1.00".
[0121] (3) Relationship recognition sub-module
[0122] The second output information and the candidate information can replace the entities in the user input text, and the $ user input text in the relationship recognition sub-module can be obtained as "<Lu Xun: person's name> and <Zhou Shuren: person's name> are the same person?". The third instruction template can be "$ task name = relationship recognition task; $ task description text = Please find the relationship most relevant to the entities in the text to be recognized among the candidate relationships. The entities in the text have been processed as <entity: entity type> format according to the $ entity type description. According to the $ relationship type description, judge the relationship type and give the corresponding confidence score; $ text to be recognized = $ replaced user input text; $ candidate relationship = $ candidate information".
[0123] The disambiguated entities will be used together with the user input text, the third instruction template, the prior knowledge of the relationship type description, and / or the prior knowledge of the entity type description to be concatenated as a task prompt and sent to the auto-encoding LLM model to complete the relationship recognition task and predict the relationship type most relevant to each entity in the input text.
[0124] The auto-encoding LLM represents the input text combined with the prompt template as a feature vector and performs decoding classification based on a downstream model specific to the relationship recognition task to predict the relationship type most likely to be mined for the entities in the text. The parameter fine-tuning training of the auto-encoding LLM is similar to the named entity recognition task. Only the downstream task model needs to be modified to a classification model to perform supervised training on the corresponding dataset.
[0125] For the above user input text, the third output information is: "Lu Xun - original name - 0.98; Zhou Shuren - pen name - 0.97; Zhou Shuren - original name - 0.96".
[0126] II. Knowledge text generation module
[0127] In this knowledge text generation module, after combining each piece of information in the second output information and the third output information in a corresponding manner, the average value of the confidence scores is calculated and then sorted. That is, by calculating the average value of the entity and relationship confidence scores in entity disambiguation and relationship recognition and sorting, the path that each entity needs to jump to for query is obtained. Specifically, after combining "1. Lu Xun - 0.99" in the second output information and "Lu Xun - original name - 0.98" in the third output information and calculating the average value of the confidence scores, "Lu Xun - original name - {} - 0.985" is obtained; after combining "4. Zhou Shuren - 1.00" in the second output information and "Zhou Shuren - pen name - 0.97" in the third output information and calculating the average value of the confidence scores, "Zhou Shuren - pen name - {} - 0.985" is obtained; after combining "4. Zhou Shuren - 1.00" in the second output information and "Zhou Shuren - original name - 0.96" in the third output information and calculating the average value of the confidence scores, "Zhou Shuren - original name - {} - 0.98" is obtained. The sorted result is: "Lu Xun - original name - {} - 0.985; Zhou Shuren - pen name - {} - 0.985; Zhou Shuren - original name - {} - 0.98"
[0128] The sorted result is sent to the knowledge graph retrieval module, and the triple information composed of all the structured knowledge involved in the user input text can be obtained, that is, the triple list. Specifically, the obtained triple list is "Lu Xun - original name - {Zhou Shuren} - 0.985; Zhou Shuren - pen name - {Lu Xun} - 0.985; Zhou Shuren - original name - {Zhou Zhangshou} - 0.98". In actual application, the triple list needs to be sorted according to the average value of the confidence scores, and those higher than 0.9 are taken for subsequent text generation. However, when the number of triples is insufficient, a batch of triples not lower than 0.7 can also be supplemented according to the scores.
[0129] Optionally, Figure 2 The other relevant knowledge information in can be the relevant triple information obtained by performing single-hop or two-hop retrieval of the triple information in the knowledge graph. Specifically, the other relevant knowledge information can be "Lu Xun - ancestral home - Zhengyang County, Henan Province; Lu Xun - date of birth - September 25, 1881; Lu Xun - date of death - October 19, 1936; Lu Xun - occupation - writer; Lu Xun - occupation - thinker; Lu Xun - introduction - one of the founders of modern Chinese literature; Lu Xun - courtesy name - Yu Cai; Lu Xun - mother - Lu Rui; Lu Xun - father - Zhou Boyi; ……". After obtaining the other relevant knowledge information, the triple information can be updated.
[0130] After obtaining the triple information, according to Figure 3The fourth instruction template in the example splices the obtained structured knowledge into text generation instructions, which are sent to the autoregressive LLM for few-shot or even zero-shot prompt learning, and processes all the structured knowledge involved in the text input by the user into fluent unstructured natural language text, i.e., the fourth output information. The fourth output information is an unstructured prior knowledge text, which is used for the autoregressive LLM to align knowledge with the real world when generating dialogue text. Specifically, the fourth instruction template is "$task name = triple-to-text task; $task description text = please combine these triples into a natural, fluent, and grammatically correct text paragraph based on the triple information in the provided knowledge graph, in the format of (X, Y, Z), where X and Z represent entities and Y represents relationships. The triple information in the knowledge graph is $triple information; $text to be recognized = none". Specifically, the fourth output information is "Lu Xun, whose original name was Zhou Shuren, was from Zhengyang County, Henan Province. He was born on September 25, 1881 and died on October 19, 1936. Lu Xun was an outstanding writer, thinker and revolutionary. He was known as a "pioneer of culture and thought" and was one of the founders of modern Chinese literature. His name was Yucai, his mother was Lu Rui, and his father was Zhou Boyi."
[0131] Structured knowledge is all textual vocabulary in terms of specific content, but it is still different from conventional natural semantic text in terms of modality. Converting structured knowledge into unstructured text allows the structured knowledge of the knowledge graph to be consistent with the input text of the autoregressive LLM in terms of modality, so that the autoregressive LLM can better summarize and generalize prior knowledge, avoiding semantic loss when structured knowledge is directly sent to the autoregressive LLM as prompt information, and making the relevant domain knowledge of the model consistent with the real world while ensuring the dialogue and communication capabilities of the model.
[0132] 3. Dialogue Generation Module Based on Knowledge Text
[0133] The unstructured prior knowledge text is concatenated with the user input text and then fed into the autoregressive LLM to generate a reply dialogue. Figure 3 As shown in the figure, specifically, the user input text is "Are Lu Xun and Zhou Shuren the same person?", and the fifth output information, that is, the answer sentence is "Yes, they are the same person. Lu Xun's real name is Zhou Shuren. Lu Xun's pen name may be taken in memory of his mother Lu Rui". With the help of the excellent contextual learning ability of the autoregressive LLM, the text generated by the reply dialogue will be highly correlated with the prior knowledge text, ensuring the consistency of the model with the knowledge of the real world.
[0134] Through the three major modules of the embodiments of the present application, namely the information extraction module, the knowledge text generation module, and the dialogue generation module based on the knowledge text, not only can the knowledge of the dialogue large model be made consistent with the real world to avoid factual errors in the generated text, but also compared with the knowledge retrieval and ranking system based on the autoencoder large model, the dialogue communication ability is significantly improved.
[0135] Furthermore, in actual application, based on the dialogue large model, converting the triple information in the knowledge graph into unstructured text can be reused for the unsupervised pre-training of the dialogue large model, and through continuous iteration, the knowledge of the dialogue large model can be made more consistent with the real world; when the knowledge graph used is the domain knowledge graph of software engineering and software testing, this dialogue text generation technology can be immediately adapted to code automatic generation and test case automatic generation. The more relevant the knowledge graph is to the task to be adapted, such as the domain knowledge graph of the mobile cloud platform, the better the generated text content can be actually used in production tasks.
[0136] To implement the dialogue generation method of the embodiments of the present application, the embodiments of the present application also provide a dialogue generation device, as Figure 4 shown. The device includes:
[0137] A first processing unit 401, configured to convert first input information into a first language instruction based on a set first instruction template, and call a first autoencoder large language model LLM to process the first language instruction to obtain first output information; the first input information includes user input text; the first output information includes one or more first entities and corresponding entity types in the user input text.
[0138] A second processing unit 402, configured to convert second input information into a second language instruction based on a set second instruction template, and call a second autoencoder LLM to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information, and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained by performing a single-hop retrieval on the first output information in the knowledge graph; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities.
[0139] A third processing unit 403, configured to convert third input information into a third language instruction based on a set third instruction template, and call a third autoencoder LLM to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to the one or more second entities in the second output information.
[0140] The first generation unit 404 is configured to generate a response statement regarding the user input text based on the second output information and the third output information.
[0141] Wherein, in one embodiment,
[0142] the first input information further includes entity type description information; and / or,
[0143] the third input information further includes the entity type description information and relationship type description information; wherein,
[0144] the entity type description information and / or the relationship type description information is determined based on the prior knowledge of the knowledge graph.
[0145] In one embodiment, the entity type description information and / or the relationship type description information is updated as the knowledge graph is updated.
[0146] In one embodiment, the first generation unit 404 is configured to call the knowledge graph to process the second output information and the third output information to obtain triple information;
[0147] Construct a prior knowledge statement based on the triple information;
[0148] Generate a response statement regarding the user input text based on the prior knowledge statement and the user input text.
[0149] In one embodiment, the apparatus further includes:
[0150] A fourth processing unit, configured to input the triple information into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information before constructing the prior knowledge statement based on the triple information;
[0151] A first update unit, configured to update the triple information based on the relevant knowledge information.
[0152] In one embodiment, the first output information further includes a first confidence score for each of the first entities; the first confidence score is used to judge the credibility of the entity type corresponding to the first entity;
[0153] The second output information further includes a second confidence score for each of the second entities; the second confidence score is used to judge the relevance between each of the second entities in the candidate information and the corresponding first entity;
[0154] The third output information further includes a third confidence score; the third confidence score is used to judge the relevance between each of the second entities and different relationship types.
[0155] In one embodiment, the device further includes:
[0156] A fifth processing unit, configured to, after obtaining the triple information, calculate the average value of the second confidence score and the corresponding third confidence score corresponding to each triple in the triple information, and sort all the triples in the triple information according to the average value to obtain a sorting result;
[0157] A first judgment unit, configured to determine whether one or more triples in the triple information can be used to construct the prior knowledge statement based on the sorting result.
[0158] In one embodiment, a first generation unit 404 is configured to process the triple information based on a set fourth instruction template to obtain a fourth language instruction, and call a first autoregressive LLM to process the fourth language instruction to obtain the prior knowledge statement; the prior knowledge statement represents a natural language with smooth sentences based on the triple information.
[0159] In one embodiment, the first generation unit 404 is configured to call a second autoregressive LLM, and use the prior knowledge statement and the user input text as inputs to generate the answer statement.
[0160] In practical applications, the above units can be implemented by a processor in the dialogue generation device.
[0161] It should be noted that: for the dialogue generation device provided in the above embodiment, only the division of the above program modules is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the dialogue generation device provided in the above embodiment and the dialogue generation method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0162] Based on the hardware implementation of the above program modules, an embodiment of the present application further provides an electronic device, such as Figure 5 as shown, the electronic device 500 includes:
[0163] A communication interface 501, capable of interacting with other network nodes;
[0164] A processor 502, connected to the communication interface 501 to enable information interaction with other network nodes, is used to execute the methods provided by the above one or more technical solutions when running a computer program. The computer program is stored on a memory 503.
[0165] When the processor 502 runs a computer program to execute the above one or more dialogue generation methods, specifically:
[0166] The processor 502 is used for:
[0167] Based on a set first instruction template, convert first input information into a first language instruction, and call a first auto-encoding large language model (LLM) to process the first language instruction to obtain first output information; the first input information includes user input text; the first output information includes one or more first entities and corresponding entity types in the user input text;
[0168] Based on a set second instruction template, convert second input information into a second language instruction, and call a second auto-encoding LLM to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information, and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained by performing single-hop retrieval of the first output information in a knowledge graph; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities;
[0169] Based on a set third instruction template, convert third input information into a third language instruction, and call a third auto-encoding LLM to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to the one or more second entities in the second output information;
[0170] Generate a response statement regarding the user input text based on the second output information and the third output information.
[0171] Wherein, in one embodiment, the first input information further includes entity type description information; and / or,
[0172] The third input information further includes the entity type description information and relationship type description information; wherein,
[0173] The entity type description information and / or the relationship type description information are determined based on the prior knowledge of the knowledge graph.
[0174] In one embodiment, the entity type description information and / or the relationship type description information are updated as the knowledge graph is updated.
[0175] In one embodiment, the processor 502 is further configured to: process the second output information and the third output information by invoking the knowledge graph to obtain triple information;
[0176] Construct a priori knowledge statements based on the triple information;
[0177] Generate an answer statement regarding the user input text based on the a priori knowledge statement and the user input text.
[0178] In one embodiment, the processor 502 is further configured to:
[0179] Before constructing the a priori knowledge statement based on the triple information, input the triple information into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information;
[0180] Update the triple information based on the relevant knowledge information.
[0181] In one embodiment, the first output information further includes a first confidence score for each of the first entities; the first confidence score is used to judge the credibility of the entity type corresponding to the first entity;
[0182] The second output information further includes a second confidence score for each of the second entities; the second confidence score is used to judge the relevance between each of the second entities in the candidate information and the corresponding first entity;
[0183] The third output information further includes a third confidence score; the third confidence score is used to judge the relevance between each of the second entities and different relationship types.
[0184] In one embodiment, the processor 502 is further configured to:
[0185] After obtaining the triple information, calculate the average value of the second confidence score and the corresponding third confidence score for each triple in the triple information, and sort all the triples in the triple information according to the average value to obtain a sorting result;
[0186] Based on the sorting result, determine whether one or more triples in the triple information can be used to construct the a priori knowledge statement.
[0187] In one embodiment, the processor 502 is further configured to: process the triple information based on a set fourth instruction template to obtain a fourth language instruction, and call a first autoregressive LLM to process the fourth language instruction to obtain the prior knowledge statement; the prior knowledge statement represents a natural language with smooth sentences obtained based on the triple information.
[0188] In one embodiment, the processor 502 is further configured to: call a second autoregressive LLM, and use the prior knowledge statement and the user input text as inputs to generate the answer statement.
[0189] It should be noted that: The specific processing procedures of the processor 502 and the communication interface 501 can be understood with reference to the above method.
[0190] Of course, in actual applications, the various components in the electronic device 500 are coupled together through the bus system 504. It can be understood that the bus system 504 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 5 all kinds of buses are labeled as the bus system 504.
[0191] The memory 503 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device 500. Examples of these data include: any computer program for operating on the electronic device 500.
[0192] The method disclosed in the embodiments of the present application above can be applied to the processor 502 or implemented by the processor 502. The processor 502 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 502 or by instructions in software form. The above-mentioned processor 502 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 502 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 503. The processor 502 reads the information in the memory 503 and combines its hardware to complete the steps of the foregoing method.
[0193] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), general purpose processors, controllers, microcontroller units (MCUs), microprocessors, or other electronic components, and is used to execute the foregoing method.
[0194] It can be understood that the memory 503 in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (FlashMemory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), dynamic random access memory (DRAM, Dynamic Random Access Memory), synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), double data rate synchronous dynamic random access memory (DDRSDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random AccessMemory), sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random AccessMemory), direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but not limited to, these and any other suitable types of memories.
[0195] In an exemplary embodiment, the embodiments of the present application further provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium. For example, it includes a memory 503 storing a computer program. The above computer program can be executed by a processor 502 of an electronic device 500 to complete the steps described in the dialogue generation method. The computer-readable storage medium can be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0196] Exemplarily, the embodiments of the present application further provide a computer program product, including a computer program. The computer program can be executed by a processor 502 of an electronic device 500 to complete the steps described in the dialogue generation method.
[0197] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0198] As used herein, the term "one or more" means any one or any combination of at least two of a plurality. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C. As used herein, the term "one or more items" means any one or any combination of at least two of a plurality of items. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set composed of A, B, and C.
[0199] In addition, the technical solutions described in the embodiments of the present application can be arbitrarily combined without conflict.
[0200] The above is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A method for generating a dialogue, characterized in that: The method comprises: Based on a set first instruction template, convert the first input information into a first language instruction, and call a first autoencoder large language model LLM to process the first language instruction to obtain first output information; the first input information includes a user input text; the first output information includes one or more first entities and corresponding entity types in the user input text; Based on the set second instruction template, the second input information is converted into a second language instruction, and the second autoencoder LLM is called to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained after the first output information is input into the knowledge graph for single-hop retrieval; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities; Based on a set third instruction template, convert the third input information into a third language instruction, and call a third autoencoder LLM to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types that are most relevant to one or more second entities in the second output information; Based on the second output information and the third output information, a reply sentence regarding the user input text is generated.
2. The method according to claim 1, characterized in that The first input information also includes entity type description information; and / or, The third input information also includes the entity type description information and the relationship type description information; wherein, The entity type description information and / or the relationship type description information are determined based on prior knowledge of the knowledge graph.
3. The method according to claim 2, characterized in that The entity type description information and / or the relationship type description information is updated as the knowledge graph is updated.
4. The method according to claim 1, characterized in that: The step of generating a reply statement about the user input text based on the second output information and the third output information comprises: Calling the knowledge graph to process the second output information and the third output information to obtain triple information; Based on the triple information, construct a priori knowledge statement; Based on the prior knowledge sentence and the user input text, an answer sentence regarding the user input text is generated.
5. The method according to claim 4, characterized in that Before constructing a priori knowledge statement based on the triple information, the method further includes: The triple information is input into the knowledge graph for single-hop retrieval to obtain relevant knowledge information of the triple information; based on the relevant knowledge information, the triple information is updated.
6. The method according to claim 4, characterized in that The first output information also includes a first confidence score for each of the first entities; the first confidence score is used to determine the credibility of the entity type corresponding to the first entity; The second output information further includes a second confidence score for each of the second entities; The second confidence score is used to determine the relevance between each of the second entities in the candidate information and the corresponding first entity; The third output information also includes a third confidence score; The third confidence score is used to determine the relevance between each of the second entities and different relationship types.
7. The method according to claim 6, characterized in that After obtaining the triplet information, the method further includes: Calculating an average value of the second confidence score and the corresponding third confidence score corresponding to each triple in the triple information, and sorting all triples in the triple information according to the average value to obtain a sorting result; Based on the sorting result, it is determined whether one or more triples in the triple information can be used to construct the priori knowledge statement.
8. The method according to claim 4, characterized in that The constructing a priori knowledge statement based on the triple information includes: The triple information is processed based on a set fourth instruction template to obtain a fourth language instruction, and the first autoregressive LLM is called to process the fourth language instruction to obtain the prior knowledge sentence; the prior knowledge sentence represents a natural language with smooth sentences obtained based on the triple information.
9. The method according to claim 4 or 8, characterized in that: The step of generating the answer statement about the user input text based on the prior knowledge statement and the user input text comprises: The second autoregressive LLM is called to take the prior knowledge statement and the user input text as input to generate the answer statement.
10. A dialogue generation device, characterized in that: include: A first processing unit is used to convert first input information into a first language instruction based on a set first instruction template, and call a first autoencoder large language model LLM to process the first language instruction to obtain first output information; the first input information includes a user input text; the first output information includes one or more first entities and corresponding entity types in the user input text; a second processing unit, configured to convert the second input information into a second language instruction based on a set second instruction template, and call a second autoencoder LLM to process the second language instruction to obtain second output information; the second input information includes the user input text, the first output information and candidate information; the candidate information includes one or more second entities and corresponding relationship types obtained after the first output information is input into the knowledge graph for single-hop retrieval; the second output information includes one or more second entities in the candidate information that are most relevant to the one or more first entities; a third processing unit, configured to convert the third input information into a third language instruction based on a set third instruction template, and call a third autoencoder LLM to process the third language instruction to obtain third output information; the third input information includes the second output information and the candidate information; the third output information includes one or more relationship types most relevant to the one or more second entities in the second output information; The first generating unit is configured to generate a reply sentence regarding the user input text based on the second output information and the third output information.
11. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor is used to execute the method for dialogue generation as claimed in any one of claims 1 to 9.
12. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Method and system for performing medical question and answer by using large language model
CN116595131A
Data processing method and device based on knowledge graph, electronic equipment and medium
CN116821372A
Text structure recognition method based on large language model
CN117436441A
Method, medium and equipment for searching and outputting inquiry suggestion strategy based on knowledge graph
CN118245580A
Systems and methods for dynamic large language model prompt generation
US20240329942A1