Answer generation method and device, electronic equipment and storage medium

Through multiple rounds of recall and knowledge enhancement methods, the problem of low answer accuracy in dialogue system is solved, the accuracy of answers and question-and-answer efficiency is improved, and the knowledge ability and adaptability of dialogue robots are enhanced.

CN120256553APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410011745.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing dialogue system lacks factual knowledge, resulting in low accuracy in generating answers, especially in vertical fields, which is easy to generate unfounded answers, and users need to ask questions multiple times to obtain accurate answers, which reduces the efficiency of question and answers.

Method used

The answer is generated through multiple rounds of recall. First, the questions to be query are divided into multiple initial candidate questions, and each candidate question is recalled multiple rounds, and the initial answer is generated and knowledge enhancement is used to integrate. The domain knowledge base or third-party interface is used to generate knowledge, and the target answer is finally integrated to generate.

Benefits of technology

It improves the accuracy of the answers generated by the dialogue system, reduces the number of questions asked by users to the same question, improves the efficiency of question and answers, and enhances the knowledge ability and adaptability of the dialogue robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256553A_ABST
    Figure CN120256553A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to the technical field of artificial intelligence, and provides an answer generation method and device, electronic equipment and a storage medium, which are used for improving the accuracy of answers generated by a dialogue system and improving the question answering efficiency due to the fact that the number of questions asked by an object for the same question is reduced. The method comprises the steps of obtaining a to-be-queried question input based on an object, and generating at least one initial candidate question; for each initial candidate question, executing the following operations: generating at least one initial answer corresponding to the initial candidate question, and fusing the initial candidate question with the at least one initial answer to generate a fused candidate question; on the basis of the fused candidate questions, querying again to generate candidate answers corresponding to the fused candidate questions; the target answer corresponding to the to-be-queried question is generated based on the candidate answers corresponding to the at least one fusion candidate question, and the target answer generated based on the re-query is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the continuous development of machine learning technology and large language model technology, various types of large generative language models have developed very rapidly. Large language models contain powerful natural language capabilities and can play a very important role in the field of dialogue, such as being applied to dialogue systems.

[0003] A dialogue system is a computer system that simulates humans and aims to form a coherent and fluent dialogue with humans. When a large language model (LLM) is applied to a dialogue system, it can intelligently generate answers to questions raised by humans.

[0004] However, due to the lack of factual knowledge in current LLMs, the accuracy of the answers generated in some scenarios is not high. Specifically, although LLMs will memorize the facts and knowledge contained in the training corpus, they are unable to recall the facts. Therefore, LLMs are prone to generating answers with incorrect facts. For example, when asking the LLM: When did Einstein discover gravity? It may reply: Einstein discovered gravity in 1687. But in fact, the person who proposed the theory of gravity is Newton. Such problems will seriously damage the credibility of LLMs.

[0005] In addition, in vertical fields such as medicine, law, and e-commerce, since not much knowledge in these vertical fields is involved during model training, LLMs are prone to generating some unfounded answers.

[0006] In summary, how to improve the accuracy of the answers generated by the dialogue system is an urgent problem to be solved. Summary of the Invention

[0007] Embodiments of the present application provide an answer generation method, device, electronic device, and storage medium to improve the accuracy of the answers generated by the dialogue system, and at the same time improve the Q&A efficiency because the number of times an object asks the same question is reduced.

[0008] An answer generation method provided by an embodiment of the present application includes:

[0009] Obtain at least one initial candidate question generated based on the to-be-query question input by the object;

[0010] For each of the initial candidate questions, perform the following operations respectively:

[0011] Generate at least one initial answer corresponding to the initial candidate question, and fuse the initial candidate question with at least one of the initial answers to generate a fused candidate question;

[0012] Based on the fused candidate question, query again to generate a candidate answer corresponding to the fused candidate question;

[0013] Generate a target answer corresponding to the to-be-query question based on candidate answers respectively corresponding to at least one fusion candidate question.

[0014] An answer generation device provided by an embodiment of the present application includes:

[0015] A question generation unit, configured to obtain at least one initial candidate question generated based on an object input to-be-query question;

[0016] A query unit, configured to perform the following operations respectively for each of the initial candidate questions:

[0017] Generate at least one initial answer corresponding to the initial candidate question, and fuse the initial candidate question with at least one of the initial answers to generate a fusion candidate question;

[0018] Based on the fusion candidate question, query again to generate a candidate answer corresponding to the fusion candidate question;

[0019] An answer generation unit, configured to generate a target answer corresponding to the to-be-query question based on candidate answers respectively corresponding to at least one fusion candidate question.

[0020] Optionally, the query unit is specifically configured to:

[0021] Based on the fusion candidate question, query again in a domain knowledge base for knowledge enhancement to generate a candidate answer corresponding to the fusion candidate question; or,

[0022] Based on the fusion candidate question, query again by calling a third-party interface to generate a candidate answer corresponding to the fusion candidate question.

[0023] Optionally, the query unit is specifically configured to:

[0024] Based on a target language model, generate multiple different initial answers corresponding to the initial candidate question; the initial answers include pattern information of expected answers

[0025] Perform vector representation on multiple initial answers and the initial candidate question respectively to obtain respective corresponding embedding vectors;

[0026] Fuse the obtained multiple embedding vectors to obtain the fusion candidate question represented by a vector.

[0027] Optionally, the query unit is specifically configured to:

[0028] Average elements at corresponding positions in the obtained multiple embedding vectors to obtain a fusion vector, and use the fusion vector as the fusion candidate question represented by a vector.

[0029] Optionally, the query unit is specifically configured to:

[0030] Based on the target language model, generate an initial answer corresponding to the initial candidate question, where the initial answer includes a call tag for the third-party interface, and the call tag indicates the position in the initial candidate question where the third-party interface needs to be called and the category of the third-party interface;

[0031] Perform vector representation on the initial answer and the initial candidate question respectively to obtain respective embedding vectors;

[0032] Concatenate the obtained embedding vectors to obtain the fused candidate question represented by a vector.

[0033] Optionally, the domain knowledge base includes multiple text vectors; the fused candidate question is in the form of vector representation; then the query unit is specifically configured to:

[0034] Match the fused candidate question represented by a vector with multiple text vectors in the domain knowledge base respectively, and generate multiple candidate answers based on the successfully matched text vectors;

[0035] Wherein, the device further includes a pre-construction unit, which is used to generate the text vectors in the domain knowledge base in the following manner:

[0036] Segment the text included in the domain knowledge base into multiple text blocks;

[0037] Perform vector representation on multiple text blocks respectively to obtain multiple text vectors.

[0038] Optionally, the pre-construction unit is specifically configured to:

[0039] If the domain knowledge base includes knowledge documents, segment them into multiple text blocks according to the sentences and paragraphs in the knowledge documents;

[0040] If the domain knowledge base includes a knowledge graph, segment it into multiple text blocks according to the relationships between entities in the knowledge graph.

[0041] Optionally, the knowledge graph includes at least one of the following:

[0042] Encyclopedic knowledge-based knowledge graph, common sense-based knowledge graph, specific domain-based knowledge graph, multi-modal knowledge graph.

[0043] Optionally, the question generation unit is specifically configured to:

[0044] Perform text segmentation on the question to be queried to generate multiple initial candidate questions;

[0045] The answer generation unit is specifically configured to:

[0046] Integrate the candidate answers corresponding to the multiple fusion candidate questions to generate the target answer corresponding to the question to be queried.

[0047] Optionally, the device further includes:

[0048] A security post-processing unit, configured to, if the target answer contains content to be filtered, replace the content to be filtered with a fixed formula, or refuse to answer the question to be queried;

[0049] Wherein, the content to be filtered includes at least one of preset sensitive information and preset prohibited information.

[0050] An electronic device provided by an embodiment of the present application includes a processor and a memory. Wherein, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any one of the above answer generation methods.

[0051] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of the above answer generation methods.

[0052] An embodiment of the present application provides a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the above answer generation methods.

[0053] The beneficial effects of the present application are as follows:

[0054] Embodiments of the present application provide an answer generation method, apparatus, electronic device, and storage medium. Since the present application generates answers through multi-round recall, during this process, the query question to be processed is divided, and at least one initial candidate question can be generated. Then, multi-round recall is performed on each initial candidate question to achieve more fine-grained answer recall, so as to obtain candidate answers corresponding to each initial candidate question. Among them, during multi-round recall, the initial answer is recalled for the first time. Based on this, the fused candidate question obtained by fusing the initial answer and the initial candidate question contains both the information of the object question and the answer information. Therefore, the candidate answer obtained by querying again based on the fused candidate question has undergone a certain degree of knowledge enhancement compared to the initial answer obtained by the initial recall, making the candidate answer obtained by the second query contain richer and more accurate content. On this basis, the candidate answers corresponding to each initial candidate question are integrated, and the finally integrated generated answer is used as the target answer corresponding to the query question, which can, to a certain extent, avoid generating answers with incorrect facts or unfounded answers, improve the accuracy of the answers generated by the dialogue system, and at the same time, since the number of times the object asks the same question is reduced, the question-and-answer efficiency is improved.

[0055] In addition, when the answer generation method proposed in the present application is applied to a dialogue robot, it can effectively enhance the knowledge ability of the dialogue robot, expand the application fields of the dialogue robot, avoid some typical problems of current large language models, increase the adaptability of robot conversations and improve the dialogue experience related to the knowledge problems of the dialogue robot, and increase the stickiness of the object using the dialogue robot.

[0056] Other features and advantages of the present application will be described in the following specification, and some of them will become obvious from the specification or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0058] Figure 1 is an optional schematic diagram of an application scenario in an embodiment of the present application;

[0059] Figure 2 is a schematic diagram of a conversation scenario with a dialogue robot in an embodiment of the present application;

[0060] Figure 3 is another schematic diagram of a conversation scenario with a dialogue robot in an embodiment of the present application;

[0061] Figure 4 It is the implementation flowchart of a method for generating answers in an embodiment of this application;

[0062] Figure 5 It is the implementation flowchart of a method for generating candidate answers in an embodiment of this application;

[0063] Figure 6 It is a schematic diagram of a method and system structure for knowledge enhancement of a dialogue robot based on a large language model in an embodiment of this application;

[0064] Figure 7A It is a logical schematic diagram for generating the target answer corresponding to the question to be searched under Strategy 1 in an embodiment of this application;

[0065] Figure 7B It is another logical schematic diagram for generating the target answer corresponding to the question to be searched under Strategy 1 in an embodiment of this application;

[0066] Figure 8 It is the implementation flowchart of a method for generating candidate answers in an embodiment of this application;

[0067] Figure 9A It is a logical schematic diagram for generating the target answer corresponding to the question to be searched under Strategy 2 in an embodiment of this application;

[0068] Figure 9B It is another logical schematic diagram for generating the target answer corresponding to the question to be searched under Strategy 2 in an embodiment of this application;

[0069] Figure 10 It is a schematic diagram of the structure of a Transformer seq2seq model in an embodiment of this application;

[0070] Figure 11 It is a schematic diagram of the interaction process of a method and system for personal role - setting dialogue based on a large language model in an embodiment of this application;

[0071] Figure 12A It is a logical schematic diagram of the interaction between a terminal device and a server in an embodiment of this application;

[0072] Figure 12B It is another logical schematic diagram of the interaction between a terminal device and a server in an embodiment of this application;

[0073] Figure 13 It is a schematic diagram of the composition structure of an answer - generating device in an embodiment of this application;

[0074] Figure 14It is a schematic diagram of a hardware composition structure of an electronic device applying an embodiment of the present application;

[0075] Figure 15 It is a schematic diagram of a hardware composition structure of another electronic device applying an embodiment of the present application. Detailed implementation manners

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the technical solutions of the present application. Based on the embodiments described in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0077] Some concepts involved in the embodiments of the present application will be introduced below.

[0078] Dialogue system: Also known as a conversation agent, a computer system that simulates human conversations with people, that is, an intelligent agent, aiming to form a coherent and smooth conversation with humans. The communication methods mainly include voice, text, pictures, and of course, other methods such as gestures and touch, which are not specifically limited in this article.

[0079] Third-party interface: Refers to the application programming interface (API) of a third party relative to the dialogue system. In the case where the questions raised by the object are relatively complex, the dialogue system may not be able to generate accurate answers solely relying on the large language model. In this case, the third-party interface can be called to query more accurate content and generate the final result in combination with the query content.

[0080] Invocation mark: Represents the position in the initial candidate question where the third-party interface needs to be called and the category of the third-party interface to be called. In the embodiments of the present application, for a question, it may only be necessary to call a third-party interface at a certain position, or it may be necessary to call the third-party interface multiple times at different positions in the question, and the categories of the third-party interfaces required at different positions may be the same or different. That is, the number and category of the third-party interfaces required for a question are not limited in the present application.

[0081] Domain knowledge base: A new knowledge base proposed in this application for knowledge enhancement of dialogue models (such as large language models). This knowledge base includes a document library of various existing knowledge texts and a graph knowledge base constructed from stored historical knowledge. Among them, knowledge documents refer to relevant knowledge in the form of existing documents related to various fields such as e-commerce, law, and medicine; knowledge graphs refer to relevant knowledge in the form of existing graphs related to various fields such as e-commerce, law, and medicine.

[0082] Pattern information: Refers to a standard style of answers, which can be obtained by summarizing multiple answers.

[0083] The embodiments of this application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on natural language processing (NLP) and machine learning (ML) in artificial intelligence.

[0084] AI is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines, so that the machines have the functions of perception, reasoning, and decision-making.

[0085] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models and foundation models, can be widely applied to downstream tasks in various major directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0086] NLP is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, is developed from the LLM in the NLP field. After fine-tuning, large language models can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0087] LLM refers to a computer model that can process and generate natural language; it represents a major advancement in the field of artificial intelligence and is expected to change the field through the acquired knowledge. LLM can predict the next word or sentence by learning the statistical laws and semantic information of language data. As the input data set and parameter space continue to expand, the capabilities of LLM will also increase accordingly. It is used in a variety of application fields, such as robotics, machine learning, machine translation, speech recognition, image processing, etc., so it is called a Multimodal Large Language Model (MLLM).

[0088] The target language model in the embodiments of this application is trained using machine learning or deep learning techniques, and this model can be an LLM. Machine learning is a multi-disciplinary cross-discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0089] In addition, the target language model in the embodiments of this application can also be trained using Instruction Tuning and / or Reinforcement Learning with Human Feedback (RLHF) techniques.

[0090] Instruction fine-tuning means generating instructions separately for each task, fine-tuning on several full-shot tasks, and then evaluating the generalization ability on specific tasks. Among them, the model parameters are unfrozen, usually on a large number of publicly available NLP task data sets, to stimulate the understanding ability of the language model. By giving more obvious instructions, the model is allowed to understand and give correct feedback.

[0091] In the embodiments of the present application, based on instruction fine-tuning, large language models (LLMs) can be further trained on a data set including (instruction, output) pairs to enhance the capabilities and controllability of LLMs. Specifically, the special feature of instruction fine-tuning lies in the structure of its data set, that is, the pairing composed of human instructions and expected outputs. This structure enables instruction fine-tuning to focus on enabling the model to understand and follow human instructions, thereby enhancing the capabilities and controllability of large language models.

[0092] RLHF is an extension of reinforcement learning (RL). It incorporates human feedback into the training process, providing a natural and user-friendly interactive learning process for machines. In addition to the reward signal, the RLHF agent receives feedback from humans and learns from a broader perspective and with higher efficiency, similar to the way humans learn from another person's expertise. By building a bridge between the agent and humans, RLHF allows humans to directly guide the machine and enables the machine to master the decision-making elements significantly embedded in human experience. As an effective alignment technique, RLHF can help reduce harmful content generated by LLMs to a certain extent and improve information integrity.

[0093] After training the target language model based on the above technologies, the target language model can be applied to generate one or more answers corresponding to the question, realizing the basic dialogue function of the dialogue system.

[0094] In addition, it should be noted that the target language model in the embodiments of the present application can be trained online or offline, and no specific limitation is made here. In this article, offline training is taken as an example for illustration.

[0095] The design concept of the embodiments of the present application is briefly introduced below:

[0096] A social network, that is, Social Networking Services (SNS), refers to an Internet application that takes certain social relationships or common interests as a link and provides communication and interaction services for online aggregated objects in various forms. The social relationship network established in this way with the relationship between people as the core is mapped on the Internet to form an Internet application centered on objects and people-oriented. Social networks are typical applications of Web 2.0 and also typical manifestations of the Innovation 2.0 model in the Internet field, and their meanings include hardware, software, services, applications, etc. For example, instant messaging platforms are social networks with a huge scale and a large number of objects. There are various massive object relationships formed by these social networks, and there are also various object groups, such as various interest groups, etc. There are also various dialogue robots emerging on social networks to help objects establish more social relationships and meet various demands of objects.

[0097] A dialogue system is a computer system that simulates humans and aims to form a coherent conversation with humans. Due to the natural complexity of natural language, dialogue systems involve a very large number of NLP subtasks.

[0098] In current technical solutions, dialogue systems, especially those implemented based on large language models (LLMs), mainly face the following typical problems:

[0099] Due to the lack of factual knowledge, the answers generated by LLMs are not accurate enough in some scenarios. Specifically, LLMs will memorize the facts and knowledge contained in the training corpus, but LLMs cannot recall facts and often have hallucination problems, thus generating statements with incorrect facts.

[0100] In addition to the examples listed in the background technology, taking the very popular Chat Generative Pre-trained Transformer (ChatGPT) as an example, it is trained based on a large number of publicly available text corpora on the Internet. If you ask ChatGPT about general knowledge of the Internet, it can answer well. However, it has some obvious limitations. ChatGPT cannot answer questions whose answers are not in its training data. For example, if you ask ChatGPT "Who won the 2022 Football World Cup?", it will not be able to answer because it has not been trained with any information after September 2021.

[0101] In addition, real-world knowledge changes. Once the model is trained, the inherent limitations do not allow them to update the integrated knowledge unless the model is retrained, which requires a large cost of retraining and the collection, processing, and deployment of new corpora.

[0102] Without retraining the dialogue model, the current dialogue model cannot be well generalized to the knowledge of various unseen vertical fields and updated knowledge in practical applications. In many specific fields such as medicine, education, law, e-commerce, etc., there are a large number of very professional, specific and constantly updated knowledge bases, and it is very difficult for these knowledge to be conveniently used by the dialogue robot system. If fine-tuning / secondary pre-training is carried out in a specific field, it still cannot avoid its "nonsense", and if the training is not good, it is very likely to have "catastrophic forgetting" and lose some general capabilities.

[0103] Moreover, due to the low accuracy of the answers generated by the current dialogue system, the questioner may ask the dialogue system multiple times for the same question, reducing the Q&A efficiency; therefore, how to improve the accuracy of the answers generated by the dialogue system to improve the Q&A efficiency is a problem that needs to be solved.

[0104] In view of this, the embodiments of the present application propose an answer generation method, device, electronic device and storage medium. Since the present application generates answers through multiple rounds of recall, in this process, the question to be queried is divided, and at least one initial candidate question can be generated, and each initial candidate question is recalled multiple times to achieve a more fine-grained answer recall, so as to obtain the candidate answer corresponding to each initial candidate question; among them, during multiple rounds of recall, the initial answer is recalled for the first time. Based on this, the fusion candidate question obtained by fusing the initial answer and the initial candidate question contains both the information of the object question and the answer information. Therefore, the candidate answer obtained by querying again based on the fusion candidate question has a certain knowledge enhancement compared with the initial answer obtained by the initial recall, so that the candidate answer obtained by the second query contains richer and more accurate content. On this basis, the candidate answers corresponding to each initial candidate question are integrated, and the finally integrated generated answer is used as the target answer corresponding to the question to be queried, which can avoid generating answers with false facts or unfounded answers to a certain extent, improve the accuracy of the answers generated by the dialogue system, and at the same time reduce the number of times the object asks the same question, thereby improving the Q&A efficiency.

[0105] In addition, when the answer generation method proposed by the present application is applied to a dialogue robot, it can effectively enhance the knowledge ability of the dialogue robot, increase the application fields of the dialogue robot, avoid some typical problems of the current large language model, increase the adaptability of robot dialogue and improve the dialogue experience of knowledge-related problems of the dialogue robot, and increase the stickiness of the object using the dialogue robot.

[0106] The preferred embodiments of the present application will be described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present application, and are not intended to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0107] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiment of the present application. This application scenario diagram includes two terminal devices 110 and a server 120.

[0108] In the embodiment of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, intelligent home appliances, in-vehicle terminals, etc.; a client related to answer generation can be installed on the terminal device, and the client can be software (such as a browser, intelligent dialogue software, etc.), or a web page, a small program, etc. The server 120 is the background server corresponding to the software or the web page, the small program, etc., or a server dedicated to answer generation, and the present application does not make a specific limitation. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms.

[0109] It should be noted that the answer generation method in each embodiment of the present application can be executed by an electronic device, and the electronic device can be the terminal device 110 or the server 120, that is, the method can be executed independently by the terminal device 110 or the server 120, or can be executed jointly by the terminal device 110 and the server 120.

[0110] For example, when jointly executed by the terminal device 110 and the server 120, assume that an intelligent dialogue software is installed on the terminal device 110, and the server 120 is the background server corresponding to the intelligent dialogue software. Specifically, an object (such as a user) can input a query problem through the intelligent dialogue software installed on the terminal device 110. The server 120 obtains the query problem input by the object from the terminal device 110 and generates at least one initial candidate problem based on the query problem. Furthermore, the server 120 performs multiple rounds of recall on each initial candidate problem to generate candidate answers. Among them, for each initial candidate problem, the server 120 respectively performs the following operations: generating at least one initial answer corresponding to the initial candidate problem, and fusing the initial candidate problem with at least one initial answer to generate a fused candidate problem; based on the fused candidate problem, querying again to generate a candidate answer corresponding to the fused candidate problem. After that, the server 120 generates a target answer corresponding to the query problem based on the candidate answers corresponding to at least one fused candidate problem. Finally, the server 120 feeds back the target answer to the terminal device 110, and the terminal device 110 presents it to the object through the intelligent dialogue software.

[0111] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.

[0112] In an alternative embodiment, the communication network is a wired network or a wireless network.

[0113] It should be noted that Figure 1 The above is only an example. In fact, the number of terminal devices and servers is not limited and is not specifically defined in the embodiments of the present application.

[0114] In the embodiments of the present application, when the number of servers is multiple, the multiple servers can form a blockchain, and the server is a node on the blockchain; such as the answer generation method disclosed in the embodiments of the present application, data such as problems, answers, and feature vectors involved therein can be stored on the blockchain. For example, domain knowledge bases, query problems, initial candidate problems, fused candidate problems, initial answers, candidate answers, target answers, embedding vectors, fused vectors, text vectors, etc.

[0115] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0116] Optionally, the answer generation method in the embodiments of the present application can be applied to a dialogue system. Due to the natural complexity of natural language, the dialogue system involves a very large number of NLP subtasks.

[0117] Generally, a dialogue system is usually divided into three types: task-based, question-and-answer, and open-domain. These three types are not completely orthogonal. Specifically, during a dialogue, a single-round dialogue or a multi-round dialogue can be carried out. A single-round dialogue refers to one conversation, that is, one question and one answer, which is independent of the context. A multi-round dialogue refers to multiple conversations, centered around the intention, connecting with the context until the task is completed and the conversation turn ends.

[0118] With the continuous development of machine learning technology and large language model technology, especially since the emergence of ChatGPT, various types of large generative language models have developed very rapidly. Large language models contain rich background knowledge and powerful natural language capabilities, and can play a very important role in the field of dialogue, especially when combined with the business of various service vertical fields, such as digital humans, virtual assistants, digital robots, etc.

[0119] Taking a certain instant messaging platform as an example, there are a large number of objects at different levels on the platform, and various robots with different functions and capabilities can be developed; for example, digital idol robots can perform role settings and conversations, and different roles may require different capabilities. For example, an expert-type role requires a strong knowledge background.

[0120] These digital humans can interact and communicate with objects and fans in groups, channels, and live broadcasts. Through the answer generation method proposed in this application, core dialogue capabilities can be provided for her (or him), and at the same time, knowledge enhancement capabilities that conform to the characteristics of her (or his) character setting field can be provided, realizing the enhancement and update of various knowledge capabilities during the dialogue process.

[0121] As Figure 2 shown, it is a schematic diagram of a dialogue scenario with a dialogue robot in an embodiment of this application. Figure 2 The shown dialogue scenario is during a live broadcast. The dialogue robot can act as the host and interact and communicate with objects and fans during the live broadcast.

[0122] Another example is Figure 3 shown, which is another schematic diagram of a dialogue scenario with a dialogue robot in an embodiment of this application. Figure 3 The shown dialogue scenario is in a video call scenario of a private chat / group chat. During the video call, the dialogue robot can act as one party of the call and chat with a real object.

[0123] It should be noted that the above-listed several dialogue scenarios are only simple examples. In addition, there can be other dialogue scenarios, which will not be elaborated one by one in this article.

[0124] It should be noted that the above-mentioned dialogue robot belongs to a chat robot, which is a dialogue system in the open domain. Different from traditional task-oriented dialogue systems, open-domain dialogue systems pay more attention to free and fluent conversations, can cover a variety of topics and questions, and have a very wide range of domains, requiring the dialogue robot to have strong comprehensive knowledge capabilities.

[0125] The dialogue system in the embodiments of the present application corresponds to a mixed form of task type, question and answer type, and open domain, mainly dominated by the open domain.

[0126] In addition, it can be understood that in the specific implementation of the present application, when it comes to relevant data such as object information (such as object age, gender, behavior, etc.), when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0127] Next, in combination with the above-described application scenarios, the answer generation method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0128] Refer to Figure 4 As shown, it is a flowchart of the implementation of an answer generation method provided by an embodiment of the present application. Taking the server as the execution subject as an example, the specific implementation process of this method is as follows in S41 - S43:

[0129] S41: The server obtains at least one initial candidate question generated based on the query question input by the object.

[0130] In the embodiments of the present application, still taking the above-listed dialogue system as an example, the object can input its own question, that is, the query question, in the relevant client of the dialogue system.

[0131] Specifically, the object can input the query question in any way, such as through a virtual keyboard, a physical keyboard, voice input, video / voice call, live interaction, etc. The query question can include, but is not limited to, the following content: text, picture, voice, video, gesture, etc.

[0132] In the embodiments of the present application, after the object inputs the query question, the query question can be parsed to generate at least one initial candidate question.

[0133] Specifically, the query question in various content forms can be converted into text form, and then text segmentation can be performed to obtain one or more initial candidate questions.

[0134] For example, if the query problem input by the object is in the form of speech, the speech-to-text technology can be used to convert the query problem into text form. For another example, if the query problem input by the object is in the form of graphics and text, the features of the picture can be extracted and converted into text describing the main content of the image; and so on.

[0135] In the embodiments of the present application, for a query problem in text form, one or more initial candidate problems can be generated through text segmentation. Among them, if one initial candidate problem is generated, the initial candidate problem is the query problem itself. If multiple initial candidate problems are generated, it means that the query problem is divided into multiple sub-problems.

[0136] In the embodiments of the present application, the query problem input by the object can be divided into one or more initial candidate problems based on the above method or other methods. Furthermore, for each initial candidate problem, multi-round recall is performed in the following manner:

[0137] S42: For each initial candidate problem, the server respectively performs the following operations S421 to S422:

[0138] S421: Generate at least one initial answer corresponding to the initial candidate problem, and fuse the initial candidate problem with the at least one initial answer to generate a fused candidate problem.

[0139] S422: Based on the fused candidate problem, query again to generate a candidate answer corresponding to the fused candidate problem.

[0140] Among them, the first round of recall in the multi-round recall refers to generating an initial answer. In the embodiments of the present application, the generated initial answer can be fused with the initial candidate problem, and after fusion, at least one more recall can be performed to obtain a candidate answer corresponding to the fused candidate problem.

[0141] Among them, the process of querying again based on the fused candidate problem to generate a candidate answer corresponding to the fused candidate problem can be understood as a knowledge enhancement process, that is, based on the fused candidate problem that combines problem information and answer information, recall is performed again. In the embodiments of the present application, the recall based on the fused candidate problem belongs to knowledge-enhanced recall, which is used to generate a more accurate answer.

[0142] In the embodiments of the present application, the knowledge-enhanced recall strategies include but are not limited to the following two:

[0143] Strategy 1: Based on the fused candidate problem, query again in the domain knowledge base for knowledge enhancement to generate a candidate answer corresponding to the fused candidate problem.

[0144] Considering various existing dialogue robots and various business vertical fields, such as e-commerce, law, medicine, etc., there are a large number of knowledge documents and various knowledge graphs built in the existing stock world, storing a large amount of various strictly verified and quality-guaranteed knowledge, and the processing logic and historical services for generating this knowledge also exist. However, this knowledge and content have not been well utilized in the new dialogue robot system. Therefore, this application utilizes the above-mentioned knowledge and content and proposes a domain knowledge base for knowledge enhancement of large language models.

[0145] In the embodiments of this application, the composition of the domain knowledge base can be a knowledge text corpus or a knowledge graph precipitated historically. That is, the domain knowledge base includes at least one of the document libraries of various existing knowledge texts and the graph knowledge base built from stock historical knowledge. Here, the knowledge graph stores structured knowledge in the form of a set of (entity, relationship, entity) triples, which is very convenient for establishing relationships between different entities, expanding, and can also be very conveniently converted into specific knowledge entries.

[0146] Optionally, according to the different stored information, the knowledge graphs for knowledge enhancement include but are not limited to the following four categories:

[0147] Encyclopedic knowledge-based knowledge graphs, common-sense knowledge graphs, specific-domain knowledge graphs, and multi-modal knowledge graphs.

[0148] Among them, the multi-modal knowledge graph refers to the result obtained by processing visual and text information through a multi-modal model, and finally the graph actually stores text information.

[0149] In the embodiments of this application, Strategy 1 refers to fact retrieval and query of the knowledge fragment text (i.e., text blocks) in the domain knowledge base through multi-round recall, and then the relevant recall results are input into the LLM language model to obtain the final result.

[0150] Generally speaking, Strategy 1 adopts a method of multiple recalls and re-fusion rewriting to generate more accurate answers. The main solutions are as follows:

[0151] As Figure 5 shown, it is a flowchart of an embodiment for generating candidate answers in the embodiments of this application, including the following steps S51 to S54:

[0152] S51: Based on the target language model, generate multiple different initial answers corresponding to the initial candidate question; the initial answers contain the pattern information of the expected answer.

[0153] Among them, the target language model can be a conventional language model (such as an LLM). Based on this language model, N initial answers can be generated according to the questions requested by the object. The initial answers generated at this time are very likely to have knowledge errors, so the initial answers can be called "false answers" here. Among them, N is a positive integer greater than 1.

[0154] In the embodiment of the present application, when generating initial answers based on the LLM, the sample mode can be adopted to ensure that the N generated initial answers are different.

[0155] In the embodiment of the present application, the mode information refers to a standard style of answers, which can be obtained by summarizing multiple initial answers.

[0156] S52: Respectively perform vector representations on multiple initial answers and initial candidate questions to obtain their corresponding embedding vectors.

[0157] Specifically, a vectorization model can be used to respectively perform embedding processing on the N generated false answers and the initial candidate questions of the object to obtain N + 1 embedding vectors. The embedding vectors mentioned here refer to the embedding vectors corresponding to the initial answers and initial candidate questions in this article.

[0158] Among them, the vectorization model can be any model that can encode text into vectors, such as the pre-trained language representation model BERT (fully spelled Bidirectional Encoder Representations from Transformers), the basic model of the LLM language model, etc., which are not specifically limited in this article.

[0159] S53: Fuse the obtained multiple embedding vectors to obtain a fusion candidate question in vector representation.

[0160] In the embodiment of the present application, when multiple vectors are fused, a fused vector can finally be obtained. This fused vector is the fusion candidate question in vector representation in the embodiment of the present application.

[0161] Among them, there are many ways to fuse multiple vectors, such as summation, averaging, multiplication, etc., which are not specifically limited in this article.

[0162] Taking the averaging fusion method as an example, one implementation manner of S53 is:

[0163] Average the elements at the corresponding positions in the obtained multiple embedding vectors to obtain a fusion vector, and use the fusion vector as the fusion candidate question in vector representation.

[0164] Specifically, the following formula can be used to average N + 1 vectors:

[0165]

[0166] Among them, is the Nth generated initial answer, q ij is the initial candidate question of the object, and f is the vectorization operation. Therefore, is the embedding vector of the kth initial answer, f(q ij ) is the embedding vector of the initial candidate question, is the final fusion vector.

[0167] After generating the fusion vector, the fusion vector can be matched with the text vectors in the domain knowledge base, and candidate answers can be generated according to the matching results. The specific implementation method is as follows:

[0168] S54: Match the fusion candidate question represented by the vector with multiple text vectors in the domain knowledge base respectively, and generate multiple candidate answers based on the successfully matched text vectors.

[0169] Finally, use the fusion vector to recall the top N answers from the document library as candidate answers. Since the fusion vector contains both the information of the object question and the pattern information of the desired answer, the recall effect can be enhanced.

[0170] Specifically, when matching the fusion candidate question represented by the vector with multiple text vectors in the domain knowledge base, any vector matching method can be used.

[0171] For example, using Faiss (fully spelled as Facebook AI Similarity Search) or full-text retrieval (ElasticSearch, ES), etc., can effectively perform approximate nearest neighbor (Approximate Nearest Neighbor, ANN) search on vectors for matching.

[0172] Specifically, the domain knowledge base in the embodiments of the present application is not a traditional database, but a database for knowledge enhancement, including a document library containing various existing knowledge texts and a graph knowledge base constructed by stock historical knowledge.

[0173] Optionally, the domain knowledge base contains multiple text vectors; that is, the knowledge documents, knowledge graphs, etc. contained in the domain knowledge base can be pre-converted into the form of text vectors and stored in combination with the corresponding texts.

[0174] In an optional implementation manner in the embodiments of the present application, the domain knowledge is preprocessed in the following manner to generate the above text vectors:

[0175] First, the text included in the domain knowledge base is segmented into multiple text blocks (i.e., small text pieces). Then, the multiple text blocks are respectively vectorized to obtain multiple text vectors, and each text vector corresponds to a text block.

[0176] In the embodiments of the present application, for different types of texts, different text segmentation methods can be adopted.

[0177] For example, for the knowledge documents in the domain knowledge base, multiple text blocks can be segmented according to the sentences and paragraphs in the knowledge documents.

[0178] For example, each sentence can be used as a text block, several sentences can be used as a text block, each paragraph can be used as a text block, or several paragraphs can be used as a text block, etc.

[0179] In addition, it can also be divided according to a preset fixed length. Specifically, according to the order of sentences and paragraphs in the knowledge document, it is divided in turn, and every L characters are divided into a text block.

[0180] The text blocks obtained in this way can be of a preset fixed length, such as 60 characters, 120 characters, etc. Considering that it is not practical when the text is cut too fragmented, the 120 length is better than the 60 length. Therefore, this article takes the 120 length as an example.

[0181] For another example, for the knowledge graph in the domain knowledge base, it can be divided according to the relationship between a group of entities in the knowledge graph to obtain multiple text blocks.

[0182] That is to say, for the knowledge graph, since the knowledge graph stores structured knowledge in the form of a set of (entity, relationship, entity) triples, therefore, when performing text segmentation, it can also be divided according to the set of triples, that is, each set of triples can be used as a text block.

[0183] It should be noted that the above-listed several text segmentation methods are only simple examples. In addition, there can be other methods, which will not be elaborated here one by one.

[0184] When generating text vectors, specifically, each text block can be input into a trained language model to generate a vector representation. The generated text vector can also be called an embeeding vector, that is, the vector obtained through embedding processing.

[0185] Among them, the language model here can be any model that can encode text into a vector, such as the basic model of BERT, LLM language model, etc., which is not specifically limited in this article.

[0186] Based on the above method, the embedding vectors corresponding to each text block can be obtained. Finally, the text block and embedding vector pairs can be stored in a vector database or a <Key, Value> store, where the Key is the embedding vector and the Value is the text block.

[0187] Furthermore, Faiss or ES can be used to effectively perform ANN search on vectors for Key matching, rather than performing exact Key matching in a traditional database, so as to achieve the effect of knowledge enhancement and improve the accuracy of the generated answers.

[0188] The following briefly describes the application of Strategy 1 in the embodiments of the present application in combination with an actual dialogue robot scenario:

[0189] Refer to Figure 6 shown, which is a schematic diagram of a method and system structure for knowledge enhancement of a dialogue robot based on a large language model listed in the present application.

[0190] As Figure 6 shown, the basic corpus information of the dialogue robot mainly describes the setting of the role positioning of the robot by the object, including but not limited to: Question Answering (QA), persona QA, portrait setting, etc. Among them, automatic QA refers to the data collected by automatically setting persona prompts (Prompts) through a third-party dialogue model, such as the fourth-generation Generative Pre-trained Transformer (GPT4), usually in the form of one question and one answer, mainly to improve the data collection efficiency, and can be manually rechecked during actual use; persona QA mainly refers to questions based on the role positioning of the dialogue robot, such as "Who are you?" and "Where are you from?" The portrait setting refers to the setting and positioning of basic information such as age, gender, and hobbies in the persona of the dialogue robot.

[0191] Corpus data processing mainly converts various corpus information into single-round or multi-round dialogue forms for instruction fine-tuning of the language model when fine-tuning the large language model.

[0192] The dialogue corpus is used to store the dialogue corpus and persona description Q&A pairs collected from various channels in the early stage, so that the object can clearly understand and clarify the clear positioning of the communicating robot. In addition, the LLM large language model can also be trained and debugged based on these voice data in the dialogue corpus.

[0193] Among them, the data of Profile QA comes from the previous automatic QA and persona QA. However, here more emphasis is placed on the configuration of the role positioning of the dialogue machine. It can be considered that Profile QA is a subset of the basic corpus information of the dialogue robot mentioned above.

[0194] Among them, the basic corpus information of the dialogue robot, corpus data processing, and dialogue corpus are key parts in the existing dialogue system robot system, and will not be elaborated in detail in this application.

[0195] In Figure 6 the arrow between the dialogue corpus and the knowledge enhancement module indicates that the ability enhancement of the language model built based on the dialogue corpus depends on the next knowledge enhancement module. The core focus of this application is also to enhance the knowledge of the dialogue robot in the existing dialogue system, so that the dialogue robot can flexibly handle various fields, enabling the model to capture text semantics and also the latest real-world knowledge.

[0196] Specifically, this process is divided into two main stages. The first stage is the preprocessing step, which is used to perform text segmentation on the domain knowledge base, generate the embeeding vectors corresponding to the text blocks, and construct a vector index for approximate nearest neighbor search. Usually, Faiss or ES can be used to implement the specific index retrieval. For details, please refer to the above embodiments, and the repeated parts will not be elaborated. After generating the index, the next stage is to query and recall multiple knowledge contents, which will be described in detail below:

[0197] In the knowledge recall stage, the core is to use embedding and ANN search, and then fuse the results generated by prompt generation and the LLM large language model to obtain the final merged result. The core idea here is to adopt the vector recall method to recall the knowledge document fragments (i.e., text blocks) related to the object question from the domain knowledge base and input them into the LLM to enhance the model's answer quality. However, in many cases, especially during the dialogue process, the questions of the object are very colloquial and the descriptions are relatively vague, which will seriously affect the vector recall quality and thus the model's answer effect. Therefore, the details of the recall processing are very important.

[0198] Based on this, this application proposes the recall processing method shown in Strategy 1. Through multiple rounds of recall, the fact retrieval query is performed on the knowledge fragment text in the domain knowledge base, and then the relevant recall results are input into the LLM language model to obtain the final result. The specific process can be simply summarized as: loading the file -> reading the text -> text segmentation -> text vectorization -> question vectorization -> matching the top N most similar to the question vector in the text vectors -> adding the matched text as context and the question to the prompt -> submitting to the LLM to generate an answer.

[0199] Among them, loading text means loading the domain knowledge base file, reading the text therein, such as knowledge documents, knowledge graphs, etc., and then performing text segmentation on these texts to obtain text blocks, and encoding the text blocks to obtain text vectors. In addition, after the query question is also vectorized, the top N most similar text vectors to the question vector can be matched among the multiple text vectors included in the domain knowledge base as the initial answer; then, the matched text is added to the prompt together with the context and the question (corresponding to the fusion of the above initial answer and the query question to generate a fusion candidate question), and submitted to the LLM to generate the final answer.

[0200] It should be noted that in Figure 6 , all LLMs refer to models aligned through reinforcement learning with safe corpora. Refer to Figure 6 As shown, specifically, it refers to aligning the model data results with the ultimate human expectation through the RLHF method, which is a process of model fine-tuning to ensure that the model cannot output content that violates the law, is harmful, false, etc., ensuring that the model output results are harmless, meet expectations, and reduce hallucinations.

[0201] In addition, it should be noted that in the scenario of long text generation, if only one recall is performed, the effect is often not good, the generated text is too long, or it is easy to generate content with weak relevance to the object request (query, that is, the query question in this article). Therefore, for the query question of the object, the query text can be segmented and disassembled into multiple questions (i.e., initial candidate questions). For each initial candidate question, the text content "most relevant" to the question can be searched in the domain knowledge base. In the case of query text segmentation, the search results corresponding to the final segmentation results need to be further merged.

[0202] In the embodiment of the present application, the finally merged result also needs to go through a security post-processing module to filter out sensitive content or content not allowed to be generated by the current operation strategy, and replace it with fixed words or refuse to answer. This step is optional and not necessary.

[0203] In summary, in a dialogue robot based on the basic capabilities of a large language model, the content of the domain knowledge base is introduced, and then the method of retrieving the external domain knowledge base is used to enhance the dialogue ability of the overall dialogue machine. Moreover, the external domain knowledge base can be updated and supplemented regularly and dynamically, so that the large language model for online dialogue does not need to be retrained, and the catastrophic forgetting easily caused by training can also be avoided.

[0204] The following discusses the specific implementation logic of Strategy 1 in different cases:

[0205] Refer to Figure 7A and Figure 7BAs shown, it is a logical schematic diagram of two ways to generate the target answer corresponding to the problem to be searched under Strategy 1 in the embodiments of the present application.

[0206] Among them, Figure 7A is the logical schematic diagram corresponding to the case where the problem to be searched is not segmented.

[0207] That is to say, in S41, based on the query problem input by the object, only one initial candidate problem is generated. This initial candidate problem is the query problem itself. Furthermore, for the query problem, based on a conventional language model, N initial answers are generated, such as Figure 7A the initial answer 1, initial answer 2,..., initial answer N in Figure 7A Then, these N initial answers are fused with the query problem to generate a fused candidate problem. Finally, in a knowledge-enhanced manner, the top N text vectors matching the fused candidate problem are queried from the domain knowledge base, and the corresponding N candidate answers, that is, the target answers, are generated, such as

[0208] Figure 7B is the logical schematic diagram corresponding to the case where the problem to be searched is segmented. Taking the segmentation into 3 initial candidate problems as an example, similar to Figure 7A For each initial candidate problem, first, based on a conventional language model, N initial answers are generated respectively, such as Figure 7B the initial answers 1-1, 1-2,..., 1-N corresponding to the initial candidate problem 1 in Figure 7BFuse candidate answer 1-1, candidate answer 1-2, …, candidate answer 1-N corresponding to fusion candidate question 1; fuse candidate answer 2-1, candidate answer 2-2, …, candidate answer 2-N corresponding to fusion candidate question 2; fuse candidate answer 3-1, candidate answer 3-2, …, candidate answer 3-N corresponding to fusion candidate question 3. Finally, fuse these 3×N candidate answers to generate N target answers.

[0209] Specifically, that is, select one from the candidate answers corresponding to each fusion candidate question, and then fuse the three candidate answers to obtain one target answer. In fact, different orders of selecting candidate answers each time will also affect the fusion result, so there are many actual fusion methods.

[0210] Such as Figure 7B in, it takes candidate answer 1-1, candidate answer 2-1, candidate answer 3-1 to fuse and obtain target answer 1, candidate answer 1-2, candidate answer 2-2, candidate answer 3-2 to fuse and obtain target answer 2, …, candidate answer 1-N, candidate answer 2-N, candidate answer 3-N to fuse and obtain target answer N as an example.

[0211] Of course, candidate answer 1-1 can also be fused with candidate answer 2-2 and candidate answer 3-2, candidate answer 1-1 can also be fused with candidate answer 2-1 and candidate answer 3-2, candidate answer 1-1 can also be fused with candidate answer 2-2 and candidate answer 3-1, etc. The fusion results corresponding to different fusion methods are different, and will not be elaborated one by one here.

[0212] It should be noted that the above Figure 7A or Figure 7B The listed logic diagrams are just simple examples, and other related logics or deformations are equally applicable to the embodiments of the present application, and will not be elaborated one by one here.

[0213] In summary, the above strategy 1 effectively enhances the reasoning ability and interpretability of the basic LLM in the dialogue robot system by using an external knowledge base, enables the LLM to consider the latest knowledge when generating sentences and better understand the behavior of the LLM, and increases the actual application of the dialogue robot in more fields. To a certain extent, it can avoid the costs of retraining the large language model in the dialogue system and the collection, processing and deployment of new corpora, solve part of the hallucination problem of the generation results of the large language model and avoid the problem of catastrophic forgetting to a certain extent, and effectively experience the dialogue experience during the communication of the dialogue robot.

[0214] In addition, introducing a domain knowledge base can also enable a large number of knowledge documents in the existing stock world accumulated over time and construct various knowledge graphs. A large amount of strictly verified and quality-guaranteed knowledge content stored can be effectively applied and inherited, enabling the dialogue robot itself to serve more scenarios and businesses based on a higher starting point.

[0215] In summary, it can effectively enhance the knowledge ability of the dialogue robot, expand the application fields of the dialogue robot, give further play to the historical cumulative knowledge text content, avoid some typical problems of current large language models, increase the adaptability of robot conversations and improve the dialogue experience related to the knowledge problems of the dialogue robot, and increase the stickiness of users using the dialogue robot.

[0216] Strategy two: Based on the fused candidate questions, query again by calling a third-party interface to generate candidate answers corresponding to the fused candidate questions.

[0217] As Figure 8 shown, it is a flowchart of an embodiment for generating candidate answers in an embodiment of the present application, including the following steps S81 to S84:

[0218] S81: Based on the target language model, generate an initial answer corresponding to the initial candidate question. The initial answer contains a call mark for the third-party interface, and the call mark indicates the position in the initial candidate question where the third-party interface needs to be called and the category of the third-party interface.

[0219] Among them, the target language model can be a conventional language model (such as LLM), and the model can decide when to trigger the knowledge-enhanced recall operation by itself. Specifically, it is essentially up to the model to understand the intention of the object query. For example, when the object asks about the treatment methods of a very professional and rare disease, at this time, when the model recognizes the corresponding intention, it knows that it will be better to reply through a professional knowledge base or a third-party medical service API.

[0220] Specifically, the target language model is a pre-trained language model with the ability to mark APIs, that is, after automatic learning, the model can decide by itself when to call third-party capabilities and which third-party capabilities to call, etc.

[0221] In the embodiment of the present application, when training the target language model, the prompt can be designed and examples can be provided to let the model know that when encountering the need to query knowledge, ask questions and output according to the format. Among them, designing the prompt means giving a specific example and a specific mark, and the marked part is to tell the model to call the third-party interface. Specifically, one implementation method: it fine-tunes the language model in a self-supervised manner to let the model learn to automatically call the API.

[0222] Among them, the format of asking questions can be [Serch("questions automatically raised by the model")] (called active recall identification). Specifically, "questions automatically raised by the model" are essentially the understanding and judgment of the object query, which can be understood as query is the upstream input. After asking the question, the question generated by the model can be used to recall the answer, and the result can be directly obtained by calling a third party, and the result is unique. The result recalled here can be understood as the initial answer in the embodiment of this application.

[0223] After recalling the answer, put the answer in front of the object query, then remove the active recall mark and continue to generate answers. When the next active recall mark is generated (when the model needs other auxiliary third-party tools to determine the intent, that is, when a question needs to access different third-party tools, this is also learned), remove the content recalled last time from the prompt to prevent the content returned last time from interfering with the next recall, and keep the parameter content of the next call input correct.

[0224] Specifically, during model training, each API call can be represented as a tuple c = (ac, ic), where ac is the name of the API, ic is the corresponding input (in the embodiment of the present application, the input is the question, and the essence here is to let the model generate a sequence of descriptions of calling a third party), and the API output is r, including two cases of output API calls and excluding output API calls. Among them, the typical linearized sequences of API calls excluding and including output are respectively expressed as: e(c) = <api>ac(ic)< / api> and e(c, r) = <api>ac(ic)->r< / api> .in" <api> ”、"< / api> " and "→" are special tokens used to indicate the location of LLM instruction boundaries and output boundaries.

[0225] In the embodiment of the present application, the core of the above model training is to allow the program to automatically identify which places (words) need to be expanded to obtain third-party results through the API based on a piece of original normal corpus text (i.e., the ic in the previous text).

[0226] In an embodiment of the present application, after the model is trained, an initial answer with a call mark can be generated during the first recall based on the trained target language model. The call mark in the initial answer indicates which parts (words) in the initial candidate question need to be expanded to obtain third-party results through an API, and the category of the API that needs to be called.

[0227] In the embodiments of the present application, the API call execution is to complete each involved API. Some of these APIs may require the use of Python scripts, some may involve other neural networks, or some other search systems. However, it should be noted that for the response to each API call ci, a separate text sequence r is required.

[0228] Generally speaking, Strategy 2 refers to fine-tuning the model in the above manner so that when the model learns to answer some relatively complex knowledge-related questions, the corresponding API is called for knowledge enhancement to generate more accurate answers.

[0229] S82: Respectively perform vector representations on the initial answer and the initial candidate questions to obtain their respective corresponding embedding vectors.

[0230] Specifically, this process can also utilize a vectorization model to respectively perform embedding processing on the generated initial answer and the initial candidate questions of the object to obtain 2 embedding vectors. The embedding vectors mentioned here refer to the embedding vectors corresponding to the initial answer and the initial candidate questions in this article.

[0231] S83: Concatenate the obtained embedding vectors to obtain a fused candidate question in vector representation.

[0232] Specifically, directly concatenate the embedding vector corresponding to the initial answer and the embedding vector corresponding to the initial candidate question to obtain a concatenated vector, and this concatenated vector is the fused candidate question in vector representation.

[0233] Specifically, the concatenated vector can be expressed as

[0234] where is the embedding vector of the initial answer, and f(q ij ) is the embedding vector of the initial candidate question.

[0235] Assume has a dimension of d1, and has a dimension of d2, then has a dimension of d1 + d2.

[0236] S84: Based on the fused candidate question, query again by calling a third-party interface to generate candidate answers corresponding to the fused candidate question.

[0237] In the embodiments of the present application, when querying by calling a third-party interface, there is no restriction on the type and quantity of the APIs to be called, which is specifically determined according to the question to be queried itself. Specifically, there is no restriction that only one fixed API can be called. For different questions, different APIs can be called, and for the same question, different APIs can also be called.

[0238] For example, querying the weather and then outputting a graphical weather result may be completed based on two API services. Another example is that API 1 is required to query a certain disease, and API 2 is required to query relevant data cases of the disease, etc.

[0239] It should be noted that the above-listed API call situations are only simple examples, and the actual requirements are related to the problem to be queried itself, which will not be elaborated one by one here.

[0240] Refer to Figure 9A and Figure 9B shown in the figure, which are the logical schematic diagrams of two methods for generating the target answer corresponding to the problem to be searched under Strategy 2 in the embodiments of the present application.

[0241] Among them, Figure 9A is the logical schematic diagram corresponding to the situation where the problem to be searched is not segmented.

[0242] That is to say, in S41, based on the problem to be queried input by the object, only one initial candidate problem is generated. This initial candidate problem is the problem to be queried itself. Furthermore, for the problem to be queried, based on a conventional language model, an initial answer is generated, such as the initial answer in Figure 9A , and this initial answer contains an API call marker. Then, this initial answer is fused with the problem to be queried to generate a fused candidate problem. Finally, in a knowledge-enhanced manner, a third-party API is called to query the text vector matching the fused candidate problem and generate a corresponding candidate answer, that is, the target answer, as shown in Figure 9A shown.

[0243] Figure 9B is the logical schematic diagram corresponding to the situation where the problem to be searched is segmented. Taking the segmentation into three initial candidate problems as an example, similar to Figure 9A , for each initial candidate problem, first, based on a conventional language model, a corresponding initial answer containing an API call marker is generated, such as the initial answer 1 corresponding to the initial candidate problem 1 in Figure 9B ; the initial answer 2 corresponding to the initial candidate problem 2; the initial answer 3 corresponding to the initial candidate problem 3. Then, for each initial candidate problem, this initial answer is vector-concatenated with the initial candidate problem to generate a fused candidate problem. For example, the initial candidate problem 1 is fused with the initial answer 1 to generate the fused candidate problem 1; the initial candidate problem 2 is vector-concatenated with the initial answer 2 to generate the fused candidate problem 2; the initial candidate problem 3 is vector-concatenated with the initial answer 3 to generate the fused candidate problem 3. Then, in a knowledge-enhanced manner, a third-party API is called to query the text vector matching the fused candidate problem and generate a corresponding candidate answer, as shown in Figure 9BFuse the candidate answer 1 corresponding to the candidate question 1; fuse the candidate answer 2 corresponding to the candidate question 2; fuse the candidate answer 3 corresponding to the candidate question 3. Finally, fuse these 3 candidate answers to generate a target answer.

[0244] It should be noted that the Figure 9A or Figure 9B listed logic diagrams are just simple examples, and other related logics or deformations are equally applicable to the embodiments of the present application, which will not be elaborated here one by one.

[0245] In addition, it should be noted that Strategy 1 and Strategy 2 in the embodiments of the present application are two completely different ways in essence. The essence of Strategy 1 is still an improvement on the system content and the vector recall matching method. Strategy 2 is equivalent to function call (Function Call), introducing external APIs and service capabilities, with stronger scalability.

[0246] Of course, in addition to the knowledge enhancement methods listed in the above Strategy 1 or Strategy 2, other methods of generating target answers based on domain knowledge bases and / or API calls are equally applicable to the embodiments of the present application. In addition, the above Strategy 1 and Strategy 2 can also be used in combination, that is, for multiple initial candidate questions corresponding to the question to be queried, some initial candidate questions can generate candidate answers based on Strategy 1, and some initial candidate questions can generate candidate answers based on Strategy 2. Finally, the candidate answers are merged to generate the target answer, etc., which will not be elaborated here one by one.

[0247] The following is a brief introduction to the target language model in the embodiments of this application. The main architecture of the target language model can be based on the Transformer model. For example, the third-generation Generative Pre-trained Transformer (GPT3) and the 3.5th-generation Generative Pre-trained Transformer (GPT3.5). GPT3.5 has a series of model versions, as well as Alpaca, Camel... Most of the models are based on the decoder-only model structure and are not open source. The three models of OPT, BLOOM, and LLaMA are mainly for the open-source field. The Chinese open-source available one is the General Language Model (GLM). LLaMA is a collection of basic language models with four parameter scales of 7B, 13B, 33B, and 65B. GLM-130B is an open bilingual (English-Chinese) bidirectional dense pre-trained language model with 130 billion parameters, and is pre-trained using the algorithm of the General Language Model (GLM); there is also, for example, Yuan One in the enterprise world, which has a parameter scale of over a hundred billion and a pre-training corpus of over 2 trillion tokens, with strong Chinese understanding and creation capabilities, logical reasoning capabilities, and reliable task execution capabilities. These large language models can form a basic model, that is, the basic model (BaseModel) of the LLM mentioned in this application. Here, it is not limited to a fixed large language model. As long as it is a model using the generated Transformer architecture, it can be classified into this category. In actual implementation, models based on LLaMa and GLM are mainly selected as the basis, and on this basis, instruction fine-tuning can be carried out, that is, the task is described in the form of natural language. Almost all large models after Instruction fine-tuning are optimized through operations such as instruction fine-tuning, human feedback, and alignment on the basis of the basic language model.

[0248] For example, the underlying layer of the basic LLM language model adopts the Transformer structure. The Transformer structure completely uses the attention mechanism for sequence modeling and has achieved the best results in the machine translation task, breaking the traditional mode that the encoder-decoder model must combine with the Recurrent Neural Network (RNN). Without losing or even improving the effect, the model parallelism is greatly improved.

[0249] As shown Figure 10 in the figure, it is a schematic diagram of a Transformer seq2seq (sequence to sequence) model structure in an embodiment of the present application. The key parts of the Transformer structure network are as follows:

[0250] (1) Multi-Head Attention: Applying the self-attention mechanism to the sequence can simultaneously explore the mutual relationships between each item in the sequence and all other items.

[0251] In the embodiment of the present application, the input of the language model is a text sequence, and item refers to the token of each word in the sequence. Using Multi-Head Attention, information can be mined from different vector subspaces.

[0252] (2) Feed-Forward Network (FFN): Adding a feed-forward network layer after the attention mechanism endows the model with non-linear expression ability and can explore the interaction relationships between different dimensions.

[0253] (3) Transformer Layer: A Transformer layer is composed of a Multi-Head Attention Layer and a Feed-Forward Network. Among them, both the Attention Layer and the FFN use a residual network in the output part and perform layer normalization processing, corresponding to Figure 10 Add&Norm (summation & normalization) in it.

[0254] (4) Stacking Transformer Layers: Stacking multiple Transformer Layers together can learn more complex and higher-order interaction information.

[0255] It should be noted that the above-listed basic structure of the target language model and the key parts in this basic structure are only simple examples. In addition, other structures can also be adopted, which are not specifically limited in this article.

[0256] S43: The server generates a target answer corresponding to the question to be queried based on the candidate answers respectively corresponding to at least one fusion candidate question.

[0257] Specifically, if in S41, based on the query problem input by the object, only one initial candidate problem is generated (i.e., the query problem is not segmented), then there will only be one fused candidate problem correspondingly. Therefore, in S43, the candidate answer corresponding to the fused candidate problem can be directly used as the final target answer. As listed above Figure 7A or Figure 9A as listed.

[0258] If in S41, the query problem input by the object is segmented to generate multiple initial candidate problems; then in S43, it is necessary to integrate the candidate answers corresponding to the multiple fused candidate problems to generate the target answer corresponding to the query problem. As listed above Figure 7B or Figure 9B as listed.

[0259] In the embodiments of the present application, when integrating multiple candidate answers, simple splicing can be directly performed, and adjustments can be made to the discontinuous parts; or the candidate answers can be integrated according to the logical relationship between the initial candidate problems, etc. This is not specifically limited herein.

[0260] In the above implementation manner, by generating multiple initial candidate problems for the query problem according to the style, and performing multiple rounds of recall for each initial candidate problem separately, it is possible to reduce the situation that in the scenario of long text generation, if only one recall is performed, the effect is often not good, the generated text is too long, or it is easy to generate content with weak relevance to the object's request, thereby improving the accuracy of the generated answer.

[0261] Specifically, after generating the target answer corresponding to the query problem, the server can directly feedback the target answer to the terminal device, and present it to the object, notify the object, or describe it to the object, etc. through the terminal device.

[0262] Optionally, the server can further perform a security analysis on the target answer, that is, analyze whether the target answer contains content to be filtered. If the target answer contains content to be filtered, the content to be filtered is replaced with a fixed formula, or the query problem is refused to be answered.

[0263] Among them, the content to be filtered includes at least one of preset sensitive information and preset prohibited information.

[0264] The preset sensitive information is some sensitive words preset in advance according to experience or requirements, such as some sensitive words related to personal privacy (such as age, weight, etc.), or some sensitive words related to illegal and disciplinary violations, some uncivilized terms, some false facts, etc. This is not specifically limited herein.

[0265] The preset disabled information refers to some content that is not allowed to be generated according to the current operating strategy based on experience or requirements. For example, it is not allowed to generate remarks related to the views of some sensitive events, and it is not allowed to generate remarks that guide illegal and disciplinary behavior. This article does not make specific limitations.

[0266] After replacing the content to be filtered with a fixed statement, the server can feedback the updated target answer to the terminal device, and present it to the object, notify the object, or describe it to the object through the terminal device, etc.

[0267] If it is determined to refuse to answer, the server can feedback relevant information or instructions indicating refusal to answer to the terminal device, and the terminal device notifies the object to refuse to answer the question.

[0268] Alternatively, the target language model in the embodiments of the present application can be trained through security post-processing to ensure that the generated answers are secure and there is no content to be filtered.

[0269] In the above implementation manner, through the above security post-processing, it can be ensured that the target answer finally presented to the object is more credible and more in line with the expectations of the object, which is beneficial to improving the stickiness of the object.

[0270] In order to illustrate the relationship and role of the dialogue robot knowledge enhancement model and service among other modules in a typical instant messaging system, the present application also introduces an interaction process of a personal role setting dialogue method and system based on a large language model, such as Figure 11 shown in the figure, which is a schematic diagram of the interaction process of a personal role setting dialogue method and system based on a large language model in the embodiments of the present application. Among them, the part with black background and white characters is the key module involved in the present application, and the modules in the part with white background and black characters are the services where the social communication functions of a typical existing instant messaging system are located. The main purpose is to provide an example for understanding the role played by the model and service mentioned in the present application.

[0271] The following briefly describes the interaction process of the personal role setting dialogue method and system based on the large language model, such as Figure 11 shown in the figure:

[0272] Among them, the most important is the dialogue robot knowledge enhancement model and service. Through this part, the system completes the generation of role dialogues. The main functions of each service module of a dialogue robot knowledge enhancement method and system based on a large language model are described in detail as follows:

[0273] 1. Terminal (front end, such as the content production end).

[0274] (1) The [terminal] communicates with the message and content service access server through the [terminal] and messages to complete the up and down of the message function. In addition, content producers who specialize in producing Professional Generated Content (PGC), User Generated Content (UGC), Multi-Channel Network (MCN), or Professional User Generated Content (PUGC) provide local or shot videos through the mobile terminal or the back-end interface API system. These are the main content sources for distributing content and can also be considered a broad sense of [terminal].

[0275] In the embodiments of the present application, messages and content can be directly synchronized between [terminals], such as Sd1. For another example, during the communication between [terminals], the [terminal] and the message and content service access server can send messages or obtain content, such as Sa1; for another example, the [terminal] and the message and content service access server can synchronize messages or issue content distribution results, such as Sa6.

[0276] (2) The [terminal] is the carrier for carrying the functions of various scenarios in the content and social business ecosystem. For example, functions such as finding friends in mobile instant messaging software, adding strangers as friends, one-on-one chatting with friends, group chatting with friends, communicating with the anchor during live broadcasts, posting content posts on channels, and posting dynamic statuses in spaces (these are all sub-business scenarios). The communication objects can be real objects or various chatbots, and these chatbots can have virtual avatars and conduct immersive conversations with the objects.

[0277] In the embodiments of the present application, if the communication object is a chatbot, the knowledge ability of the chatbot can be enhanced based on the answer generation methods listed in the embodiments of the present application, the application fields of the chatbot can be increased, the adaptability of the chatbot conversation can be increased, the conversation experience related to the knowledge problems of the chatbot can be improved, and the stickiness of the object using the chatbot can be increased.

[0278] (3) The object can publish content through the [terminal]. When publishing content, usually the upload server interface address is obtained first, and then the local file is uploaded. During the shooting process, the local graphic and text content can be paired with music, filter templates, and filter beautification functions, etc.

[0279] (4) The [terminal] communicates with the statistical reporting and analysis interface server to collect detailed object behaviors and feedback data in each sub-business scenario in the social network scenario, and saves the collected data in the statistical analysis database as the basic data source for analyzing the object data attributes of the platform, which is used to guide what types of role-based chatbots need to be built and what knowledge in which fields they need to possess.

[0280] In the embodiments of the present application, between the terminal and the statistical reporting and analysis interface server, object social network behavior analysis and feedback reporting can be performed, that is, the above information of the object is reported to the statistical reporting and analysis interface server, such as Sc1.

[0281] II. Message and content service access server.

[0282] (1) The message and content service access server synchronizes with the terminal to complete the uplink and downlink communication and synchronization of messages.

[0283] For related steps such as Sa1 and Sa6 listed above, they will not be repeated here.

[0284] (2) The message and content service access server docks the message content with the message system and the message and content database to complete the message storage processing logic.

[0285] In the embodiments of the present application, such as Sa2, after the message and content service access server receives the message sent by the terminal, it can write the message into the message system, and then, through the message system, write and store it in the message and content database, such as Sa3.

[0286] (3) The message and content service access server communicates directly with the content production terminal. The content submitted from the front end, usually the title of the content, the publisher, the abstract, the cover image, the release time, or the directly shot video enters the server through this server and stores the file in the message and content database.

[0287] (4) The message and content service access server writes the meta information of the video content, such as file size, cover image link, bit rate, file format, title, release time, author, etc. into the message and content database.

[0288] III. Message and content database.

[0289] (1) The message and content database temporarily stores the messages of the object conversation, realizes the roaming of messages and the synchronization of multi-terminal messages, such as point-to-point messages and group messages. Add a chatbot to the friend list, and the communication with each other is also through the way of sending messages.

[0290] (2) As the core module of the message system, the message and content database optimizes the storage and indexing processing of messages with high efficiency and is the information source for multi-terminal message synchronization.

[0291] (3) The message and content database is the core database of content. All the meta-information of the content published by producers is stored in this business database. The key points are the meta-information of the content itself, such as file size, cover image link, bit rate, file format, title, release time, author, whether it is original or a first release, and also the classification of the content during the manual review process (including first-level, second-level, and third-level classifications and tag information. For example, for a video explaining Huawei mobile phones, the first-level classification is technology, the second-level classification is smart phones, the third-level classification is domestic mobile phones, and the tag information is x1 brand, x2 model, shooting artifact).

[0292] (4) When receiving a video file, the up and down content interface service performs standard transcoding operations on the content. After transcoding is completed, it asynchronously returns the meta-information, mainly including file size, bit rate, specifications, and the intercepted cover image. These information will be stored in the message and content database.

[0293] IV. Message System.

[0294] (1) The message system is responsible for the entire flow scheduling and distribution of social message synchronization and communication, including point-to-point personal-to-person (Consumer to Consumer, C2C) messages and group messages, etc.

[0295] Such as Sa4, the message is sent down to the message and content service access server, and then the message and content service access server forwards the message to the other end.

[0296] (2) The message system is responsible for communicating with the message and content database to complete the distribution and processing of messages.

[0297] For the relevant steps such as Sa3 above, the repeated parts will not be elaborated again.

[0298] V. Statistical Reporting and Analysis Interface Service.

[0299] (1) The statistical reporting and analysis interface service communicates with the terminal, receiving various feedbacks during the process of message consumption and distribution reported, such as reports and feedbacks on the quality of content distribution, satisfaction scores on the dialogue results of the chatbot, etc.

[0300] For the relevant steps such as Sc1 above, the repeated parts will not be elaborated again.

[0301] (2) The terminal reports the object behavior data in different sub-business scenarios. After real-time data cleaning, it is stored in different storage engines, and the corresponding data and feedback information are mined in combination with the content flow of different sub-business scenarios.

[0302] For the relevant steps such as Sc1 above, the repeated parts will not be elaborated again.

[0303] (3) The statistical reporting and analysis interface service collects various content quality problems actively feedback and reported by the consumer-side objects, and also includes the feedback and reporting of various interaction behaviors between the objects and the generated results of the dialogue system.

[0304] As in the relevant steps of Sc1 above, the repeated parts will not be elaborated.

[0305] (4) After the statistical reporting and analysis interface service cleans the reported results, it saves the cleaned results in the statistical analysis database respectively, which is used to evaluate and measure the performance of the dialogue robot itself and whether it achieves the expected effect, and to guide the subsequent improvement direction.

[0306] In the embodiment of the present application, as in Sc2, after the statistical reporting and analysis interface service receives the data reported by the receiving end, it cleans these data, and then statistical information and samples can be sorted out and written into the statistical analysis database.

[0307] VI. Statistical analysis database.

[0308] The statistical analysis database is used to communicate with the statistical reporting and analysis interface service, and saves the message content after desensitization processing and the preliminary processing results of cleaning and verifying the data of different sub-business scenarios.

[0309] As in the relevant steps of Sc2 above, the repeated parts will not be elaborated.

[0310] VII. Dialogue robot knowledge enhancement model.

[0311] According to the above detailed description, it is used to combine with the dialogue system by introducing a domain knowledge base, and improve the knowledge ability of the machine in the way of enhanced vector multi-modal recall retrieval, so as to have a good adaptability to dialogues in different fields.

[0312] In the embodiment of the present application, as in Sb2, by service-ifying the dialogue robot knowledge enhancement processing service, a dialogue robot knowledge enhancement model can be obtained. Then, as in Sb3, a knowledge-enhanced robot can be constructed using a large language model.

[0313] VIII. Dialogue robot knowledge enhancement processing service.

[0314] (1) The dialogue robot knowledge enhancement processing service is used to service-ify each sub-processing system of the above dialogue robot. The overall process is divided into two main stages. The first stage is the preprocessing step, which is used to generate embedding vectors and construct a vector index for approximate nearest neighbor search. After generating the index, the next stage is querying. Faiss or ES is used as the engineering implementation solution for vector retrieval query to implement the retrieval and enhancement of the specific external knowledge base. For details, please refer to the relevant part of Strategy 1 above. The repeated parts will not be elaborated.

[0315] (2) The dialogue robot knowledge enhancement processing service communicates with the platform system business service to complete the application of the dialogue robot in the social network.

[0316] In the embodiment of the present application, as shown in Sb1, the platform system business service can call the knowledge-enhanced dialogue robot service. Specifically, the object can communicate with the dialogue robot through the terminal in the social network, access different sub-services through the message and the content service access server (such as communicating with the anchor in the above live broadcast, etc.), and the platform system business service communicates through Sb2 and the platform system business service to complete the application of the dialogue robot in the social network.

[0317] IX. Platform System Business Service.

[0318] It generally refers to the operation system of the platform, which is used for the operation and maintenance related to services such as channel services, space services, content recommendation systems, group services, and emoji services, etc.

[0319] X. Knowledge and Corpus.

[0320] (1) According to the above description, the knowledge and corpus include a document library of existing knowledge texts and a graph knowledge base constructed from the stock historical knowledge.

[0321] (2) Keep the knowledge and corpus updated and upgraded regularly to ensure that the dialogue robot can perceive the latest knowledge in the field.

[0322] In addition, the knowledge and corpus also store fine-tuning sample data for role settings, which is used for the fine-tuning of large language models, such as Sb4.

[0323] XI. Large Language Model (a pre-trained model).

[0324] Here, it is not limited to a fixed large language model. As long as it is a model that uses a large amount of rich Internet basic corpus to construct a large-scale generative Transform architecture, it can be classified into this category. For example, when specifically implemented here, language models based on the LLaMa and GLM models can be used, and the ability of continued pre-training is utilized.

[0325] For the relevant steps such as Sb4 above, the repeated parts will not be elaborated again.

[0326] It should be noted that the above personal role setting dialogue method and system are only simple examples. In addition, other relevant dialogue systems are also applicable to the embodiments of the present application, and will not be elaborated one by one here.

[0327] As shown in Fig. 12A, it is a schematic diagram of the interaction logic between a terminal device and a server in the embodiment of the present application. Figure 12AIt is shown that the object inputs the question to be queried, "When did Einstein discover gravity?", on the chat page with the dialogue robot. The terminal device can then send the question to be queried input by the object to the server. After receiving the question to be queried, the server can adopt the Figure 4 method shown to generate the target answer, specifically including: the server generates at least one initial candidate question based on the question to be queried; furthermore, it performs multi-round recall on each initial candidate question to generate candidate answers; among them, for each initial candidate question, the server respectively performs the following operations: generates at least one initial answer corresponding to the initial candidate question, and fuses the initial candidate question with at least one initial answer to generate a fused candidate question; based on the fused candidate question, it queries again to generate the candidate answer corresponding to the fused candidate question. After that, the server generates the target answer corresponding to the question to be queried based on the candidate answers corresponding to at least one fused candidate question. Finally, the server feeds back the target answer to the terminal device, and the terminal device presents it to the object. As Figure 12A shown, the target answer is "Einstein did not discover gravity, but proposed the general theory of relativity to explain the nature of gravity. The general theory of relativity holds that gravity is caused by the curvature of space around an object. Einstein proposed this theory in 1915."

[0328] Refer to Figure 12B shown, which is another schematic diagram of the interaction logic between the terminal device and the server in the embodiments of the present application. Among them, Figure 12B taking the video call scenario between a real object (such as a real object) and the dialogue robot as an example, during the video call, what the real object says can be used as the question to be queried, and the dialogue robot can respond to what the real object says to realize the dialogue between the real object and the dialogue robot. Specifically, the terminal device can send what the object says during the video as the question to be queried to the server. After receiving the question to be queried, the server can adopt the Figure 4 method shown to generate the target answer, which will not be repeated here. Finally, the server feeds back the target answer to the terminal device, and the terminal device presents it to the object, as Figure 12B shown.

[0329] It should be noted that Figure 12A or Figure 12B the interaction logics between the terminal device and the server in the several scenarios listed are only simple examples. In addition, other scenarios are also applicable to the embodiments of the present application, and the interaction logics between the terminal device and the server in other scenarios are similar, which will not be elaborated here one by one.

[0330] Based on the same inventive concept, the embodiments of the present application also provide an answer generation device. As Figure 13As shown, it is a schematic structural diagram of the answer generation device 1300, which may include:

[0331] A question generation unit 1301, configured to obtain at least one initial candidate question generated based on the object input to be queried.

[0332] A query unit 1302, configured to perform the following operations respectively for each initial candidate question:

[0333] Generate at least one initial answer corresponding to the initial candidate question, and fuse the initial candidate question with the at least one initial answer to generate a fused candidate question.

[0334] Based on the fused candidate question, query again to generate a candidate answer corresponding to the fused candidate question.

[0335] An answer generation unit 1303, configured to generate a target answer corresponding to the to-be-query question based on the candidate answers corresponding to each of the at least one fused candidate question.

[0336] Optionally, the query unit 1302 is specifically configured to:

[0337] Based on the fused candidate question, query again in the domain knowledge base for knowledge enhancement to generate a candidate answer corresponding to the fused candidate question; or,

[0338] Based on the fused candidate question, query again by calling a third-party interface to generate a candidate answer corresponding to the fused candidate question.

[0339] Optionally, the query unit 1302 is specifically configured to:

[0340] Based on the target language model, generate multiple different initial answers corresponding to the initial candidate question; the initial answer contains the pattern information of the expected answer

[0341] Perform vector representation on the multiple initial answers and the initial candidate question respectively to obtain their corresponding embedding vectors.

[0342] Fuse the obtained multiple embedding vectors to obtain a fused candidate question in vector representation.

[0343] Optionally, the query unit 1302 is specifically configured to:

[0344] Average the elements at the corresponding positions in the obtained multiple embedding vectors to obtain a fused vector, and use the fused vector as the fused candidate question in vector representation.

[0345] Optionally, the query unit 1302 is specifically configured to:

[0346] Based on the target language model, an initial answer corresponding to the initial candidate question is generated. The initial answer contains a call tag for a third-party interface, and the call tag indicates the position in the initial candidate question where the third-party interface needs to be called and the category of the third-party interface;

[0347] The initial answer and the initial candidate question are respectively vectorized to obtain their corresponding embedding vectors;

[0348] The obtained embedding vectors are concatenated to obtain a fused candidate question in vector representation.

[0349] Optionally, the domain knowledge base contains multiple text vectors; the fused candidate question is in vector representation; then the query unit 1302 is specifically configured to:

[0350] Match the fused candidate question in vector representation with the multiple text vectors in the domain knowledge base respectively, and generate multiple candidate answers based on the successfully matched text vectors;

[0351] Wherein, the device further includes a pre-construction unit 1304, which is used to generate the text vectors in the domain knowledge base in the following manner:

[0352] Segment the text included in the domain knowledge base into multiple text blocks;

[0353] Respectively vectorize the multiple text blocks to obtain multiple text vectors.

[0354] Optionally, the pre-construction unit 1304 is specifically configured to:

[0355] If the domain knowledge base contains knowledge documents, segment them into multiple text blocks according to the sentences and paragraphs in the knowledge documents;

[0356] If the domain knowledge base contains a knowledge graph, segment it into multiple text blocks according to the relationships between entities in the knowledge graph.

[0357] Optionally, the knowledge graph includes at least one of the following:

[0358] Encyclopedic knowledge-based knowledge graph, common sense-based knowledge graph, specific domain-based knowledge graph, multi-modal knowledge graph.

[0359] Optionally, the question generation unit 1301 is specifically configured to:

[0360] Perform text segmentation on the question to be queried to generate multiple initial candidate questions;

[0361] The answer generation unit 1303 is specifically configured to:

[0362] Integrate the candidate answers corresponding to the multiple fused candidate questions to generate the target answer corresponding to the question to be queried.

[0363] Optionally, the device further includes:

[0364] A security post-processing unit 1305, configured to replace the content to be filtered with a fixed formula or refuse to answer the question to be queried if the target answer contains the content to be filtered;

[0365] The content to be filtered includes at least one of preset sensitive information and preset prohibited information.

[0366] Since the present application generates answers through multiple rounds of recall, in this process, the question to be queried is divided, at least one initial candidate question can be generated, and each initial candidate question is recalled multiple times to achieve a more fine-grained answer recall to obtain a candidate answer corresponding to each initial candidate question; among them, during multiple rounds of recall, the initial answer is recalled for the first time. Based on this, the fused candidate question obtained by fusing the initial answer and the initial candidate question contains both the information of the object question and the answer information. Therefore, the candidate answer obtained by querying again based on the fused candidate question has a certain knowledge enhancement compared to the initial answer obtained by the initial recall, making the candidate answer obtained by the second query contain richer and more accurate content. On this basis, the candidate answers corresponding to each initial candidate question are integrated, and the finally integrated generated answer is used as the target answer corresponding to the question to be queried, which can avoid generating answers with incorrect facts or unfounded answers to a certain extent, improve the accuracy of the answers generated by the dialogue system, and at the same time improve the question-and-answer efficiency because the number of times the object asks the same question is reduced.

[0367] In addition, when the answer generation method proposed in the present application is applied to a dialogue robot, it can effectively enhance the knowledge ability of the dialogue robot, increase the application fields of the dialogue robot, avoid some typical problems of current large language models, increase the adaptability of robot conversations and improve the dialogue experience of knowledge-related problems of the dialogue robot, and increase the stickiness of the object using the dialogue robot.

[0368] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0369] For the convenience of description, the above - mentioned parts are divided into various modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.

[0370] After introducing the answer - generation method and device of the exemplary embodiment of this application, next, an electronic device according to another exemplary embodiment of this application will be introduced.

[0371] Those skilled in the art can understand that various aspects of this application can be implemented as a system, a method, or a program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0372] Based on the same inventive concept as the above - mentioned method embodiment, an electronic device is also provided in an embodiment of this application. In one embodiment, the electronic device can be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device can be as Figure 14 shown, including a memory 1401, a communication module 1403, and one or more processors 1402.

[0373] The memory 1401 is used to store the computer program executed by the processor 1402. The memory 1401 mainly includes a program - storage area and a data - storage area. Among them, the program - storage area can store an operating system and programs required to run the instant - messaging function, etc.; the data - storage area can store various instant - messaging information and operation instruction sets, etc.

[0374] The memory 1401 can be a volatile memory, such as a random - access memory (RAM); the memory 1401 can also be a non - volatile memory, such as a read - only memory, a flash memory, a hard disk drive (HDD), or a solid - state drive (SSD); or the memory 1401 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1401 can be a combination of the above - mentioned memories.

[0375] The processor 1402 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1402 is used to implement the above answer generation method when calling the computer program stored in the memory 1401.

[0376] The communication module 1403 is used to communicate with the terminal device and other servers.

[0377] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1401, communication module 1403 and processor 1402 is not limited. In the embodiments of the present application Figure 14 it is described that the memory 1401 and the processor 1402 are connected through a bus 1404. The bus 1404 is depicted in Figure 14 in thick lines. The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1404 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 14 it is only depicted by a thick line in

[0378] The memory 1401 stores a computer storage medium, and the computer storage medium stores computer-executable instructions. The computer-executable instructions are used to implement the answer generation method of the embodiments of the present application. The processor 1402 is used to execute the above answer generation method, as Figure 4 shown.

[0379] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device may be as Figure 15 shown, including: a communication component 1510, a memory 1520, a display unit 1530, a camera 1540, a sensor 1550, an audio circuit 1560, a Bluetooth module 1570, a processor 1580 and other components.

[0380] The communication component 1510 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology. The electronic device can help users send and receive information through the WiFi module.

[0381] The memory 1520 can be used to store software programs and data. The processor 1580 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1520. The memory 1520 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 1520 stores an operating system that enables the terminal device 110 to operate. In this application, the memory 1520 can store the operating system and various application programs, and can also store a computer program for executing the answer generation method of the embodiments of this application.

[0382] The display unit 1530 can also be used to display information input by the user or information provided to the user, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1530 may include a display screen 1532 disposed on the front of the terminal device 110. Among them, the display screen 1532 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1530 can be used to display the object operation interface of the intelligent dialogue system in the embodiments of this application.

[0383] The display unit 1530 can also be used to receive input digital or character information, and generate a signal input related to the user settings and function control of the terminal device 110. Specifically, the display unit 1530 may include a touch screen 1531 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.

[0384] Among them, the touch screen 1531 can cover the display screen 1532, or the touch screen 1531 and the display screen 1532 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply called a touch display screen. In this application, the display unit 1530 can display application programs and corresponding operation steps.

[0385] The camera 1540 can be used to capture static images, and the user can publish the images captured by the camera 1540 through an application. The camera 1540 can be one or multiple. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1580 to convert it into a digital image signal.

[0386] The terminal device may further include at least one sensor 1550, such as an acceleration sensor 1551, a distance sensor 1552, a fingerprint sensor 1553, and a temperature sensor 1554. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, an optical sensor, and a motion sensor.

[0387] The audio circuit 1560, the speaker 1561, and the microphone 1562 may provide an audio interface between the user and the terminal device 110. The audio circuit 1560 may transmit the electrical signal converted from the received audio data to the speaker 1561, and the speaker 1561 converts it into a sound signal for output. The terminal device 110 may also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1562 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1560, converted into audio data, and then the audio data is output to the communication component 1510 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 1520 for further processing.

[0388] The Bluetooth module 1570 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1570 to perform data interaction.

[0389] The processor 1580 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 1520 and calling data stored in the memory 1520, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1580 may include one or more processing units; the processor 1580 may also integrate an application processor and a baseband processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1580. In this application, the processor 1580 can run the operating system, application programs, user interface display, and touch response, as well as the answer generation method of the embodiments of this application. In addition, the processor 1580 is coupled to the display unit 1530.

[0390] In some possible implementation manners, various aspects of the answer generation method provided in this application may also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the answer generation method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute the steps as shown in Figure 4 shown in.

[0391] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0392] The program product of the embodiments of the present application can adopt a portable compact disk read-only memory (CD-ROM) and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0393] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.

[0394] The computer program contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.

[0395] The computer program for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program can be executed entirely on the user's electronic device, partially on the user's electronic device, executed as a stand-alone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device can be connected to the user's electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or, it can be connected to an external electronic device (for example, by using an Internet service provider to connect through the Internet).

[0396] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of this application, the features and functions of the two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0397] In addition, although the operations of the method of this application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0398] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable computer programs.

[0399] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0400] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0401] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.

[0402] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0403] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for generating answers, characterized in that, The method includes: Obtaining at least one initial candidate question generated based on a to-be-query question input by an object; For each of the at least one initial candidate question, the following operations are respectively performed: Generating at least one initial answer corresponding to the initial candidate question, and fusing the initial candidate question with the at least one initial answer to generate a fused candidate question; Based on the fused candidate question, querying again to generate candidate answers corresponding to the fused candidate question; Generating a target answer corresponding to the to-be-query question based on the candidate answers corresponding to each of the at least one fused candidate question.

2. The method according to claim 1, wherein The querying again based on the fused candidate question to generate candidate answers corresponding to the fused candidate question includes: Based on the fused candidate question, querying again in a domain knowledge base for knowledge enhancement to generate candidate answers corresponding to the fused candidate question; or Based on the fused candidate question, querying again by calling a third-party interface to generate candidate answers corresponding to the fused candidate question.

3. The method according to claim 1, wherein The generating at least one initial answer corresponding to the initial candidate question includes: Based on a target language model, generating multiple different initial answers corresponding to the initial candidate question; the initial answers contain pattern information of expected answers; The fusing the initial candidate question with the at least one initial answer to generate a fused candidate question includes: Performing vector representations on the multiple initial answers and the initial candidate question respectively to obtain respective corresponding embedding vectors; Fusing the obtained multiple embedding vectors to obtain the fused candidate question in vector representation.

4. The method according to claim 3, wherein The fusing the obtained multiple embedding vectors to obtain the fused candidate question in vector representation includes: Taking the average of elements at corresponding positions in the obtained multiple embedding vectors to obtain a fused vector, and using the fused vector as the fused candidate question in vector representation.

5. The method according to claim 1, characterized in that, The generating at least one initial answer corresponding to the initial candidate question includes: Based on a target language model, generating an initial answer corresponding to the initial candidate question, where the initial answer contains a call mark of the third-party interface, and the call mark indicates the position in the initial candidate question where the third-party interface needs to be called and the category of the third-party interface; The fusing the initial candidate question with the at least one initial answer to generate a fused candidate question includes: Performing vector representations on the initial answer and the initial candidate question respectively to obtain respective corresponding embedding vectors; Concatenating the obtained respective embedding vectors to obtain the fused candidate question in vector representation.

6. The method according to claim 2, wherein The domain knowledge base contains multiple text vectors; the fused candidate question is in vector representation; then the querying again based on the fused candidate question in the domain knowledge base for knowledge enhancement to generate candidate answers corresponding to the fused candidate question includes: Matching the fused candidate question in vector representation with the multiple text vectors in the domain knowledge base respectively, and generating multiple candidate answers based on the text vectors with successful matches; Wherein, the text vectors in the domain knowledge base are generated by the following method: Split the text included in the domain knowledge base into multiple text blocks; Respectively perform vector representation on the multiple text blocks to obtain multiple text vectors.

7. The method according to claim 6, wherein The splitting the text included in the domain knowledge base into multiple text blocks includes: If the domain knowledge base includes knowledge documents, split to obtain multiple text blocks according to the sentences and paragraphs in the knowledge documents; If the domain knowledge base includes a knowledge graph, split to obtain multiple text blocks according to the relationships between entities in the knowledge graph.

8. The method according to claim 7, wherein The knowledge graph includes at least one of the following: Encyclopedic knowledge-based knowledge graph, common sense-based knowledge graph, specific domain-based knowledge graph, multi-modal knowledge graph.

9. The method according to any one of claims 1 to 8, characterized in that The obtaining at least one initial candidate question generated based on an object-input query question includes: Perform text segmentation on the query question to generate multiple initial candidate questions; The generating a target answer corresponding to the query question based on candidate answers corresponding to at least one of the fusion candidate questions includes: Integrate the candidate answers corresponding to the multiple fusion candidate questions to generate a target answer corresponding to the query question.

10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: If the target answer contains content to be filtered, replace the content to be filtered with fixed phrases or refuse to answer the query question; Wherein, the content to be filtered includes at least one of preset sensitive information and preset prohibited information.

11. An answer generation device, characterized in that, It includes: A question generation unit, configured to obtain at least one initial candidate question generated based on an object-input query question; A query unit, configured to perform the following operations respectively for each of the initial candidate questions: Generate at least one initial answer corresponding to the initial candidate question, and fuse the initial candidate question with the at least one initial answer to generate a fusion candidate question; Based on the fusion candidate question, query again to generate a candidate answer corresponding to the fusion candidate question; An answer generation unit, configured to generate a target answer corresponding to the query question based on candidate answers corresponding to at least one fusion candidate question.

12. The device according to claim 11, wherein The query unit is specifically configured to: Based on the fusion candidate question, query again in the domain knowledge base for knowledge enhancement to generate a candidate answer corresponding to the fusion candidate question; or, Based on the fusion candidate question, query again by calling a third-party interface to generate a candidate answer corresponding to the fusion candidate question.

13. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that, It includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 10.

15. A computer program product, characterized in that, The method comprises a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the methods described in claims 1 to 10.

Citation Information

Cited By

  • Homogeneous API recommendation method and device based on artificial intelligence, electronic equipment and storage medium

    CN120974203A