Method and device for generating answers based on text attribute graph
By introducing a text attribute graph based method in RAG technology, the problem that text information and relationship information in semi-structured data are difficult to be effectively utilized, and more accurate and relevant answer generation is achieved.
Patent Information
- Application Number
- CN202510207267.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing search-enhanced generation (RAG) technologies are difficult to effectively process semi-structured data that contains both textual information and relational information, making it difficult to fully utilize the information in these data when generating answers.
Through a method based on the text attribute graph, we obtain the nodes and relationship edges in the question text and text attribute graph to match, generate a planned node and a planned relationship sequence, and input this information into the agent to traverse to generate an answer.
This method can make more comprehensive use of text information and relationship information in semi-structured data, improve the accuracy and relevance of generated answers, and avoid the hallucination problems of large language models when generating answers.
Smart Images

Figure CN120045679A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of machine learning, and in particular, to a method and apparatus for generating answers based on a text attribute graph. Background Art
[0002] Large Language Models (LLMs) have demonstrated powerful capabilities in performing tasks in various fields. However, in tasks that require strong professionalism and up-to-date knowledge, they often suffer from the problem of hallucination, generating seemingly reasonable but actually incorrect or fictional information. To solve this problem, a common approach is to use Retrieval-Augmented Generation (RAG) technology, which assists the model's answers by retrieving relevant information from external sources.
[0003] However, existing RAG technologies typically focus on a single type of external data, such as a vectorized text database or a knowledge graph, and are difficult to effectively process semi-structured data that contains both text information and relationship information. Therefore, a method is needed to jointly utilize the text information and relationship information in semi-structured data to assist the model in generating answers. Summary of the Invention
[0004] One or more embodiments of this specification describe a method and apparatus for generating answers based on a text attribute graph, which jointly utilize the text information and relationship information in semi-structured data to assist the model in answering relevant questions to generate answers.
[0005] In a first aspect, a method for generating an answer based on a text attribute graph is provided, including:
[0006] Obtain a question text and a text attribute graph, where the text attribute graph includes multiple nodes and relationship edges;
[0007] Perform a first match between the question text and each node according to the node information of each node, and obtain a number of planned nodes; the node information includes at least the node name;
[0008] Perform a second match between the question text and each relationship sequence in the text attribute graph, and obtain a number of planned relationship sequences; any relationship sequence includes a number of relationship edges connected end to end;
[0009] Input the problem text and plan information into the agent, and let the agent traverse in the text attribute graph according to the problem text and plan information to obtain a number of result nodes, and generate an answer based on the result nodes; the agent is implemented based on the first large language model; the plan information includes the plan nodes and the plan relationship sequence.
[0010] In some possible implementation manners, performing a first matching between the problem text and each node includes:
[0011] Taking the node names of the respective nodes as keywords respectively, performing keyword matching in the problem text, and determining the nodes with successful matching as plan nodes.
[0012] In some possible implementation manners, the node information further includes text information; performing a first matching between the problem text and each node further includes:
[0013] Performing similarity matching between the embedding representations corresponding to the text information of the respective nodes and the embedding representation of the problem text respectively, and determining the nodes with the top similarity rankings as plan nodes.
[0014] In some possible implementation manners, performing a second matching between the problem text and each relationship sequence in the text attribute graph includes:
[0015] Construct a problem prompt word according to the problem text, including a first text segment for instructing the relationship sequence generation model to generate a plan relationship sequence for assisting in answering the problem text;
[0016] Input the problem prompt word into the relationship sequence generation model to obtain a number of plan relationship sequences.
[0017] In some possible implementation manners, the relationship sequence generation model is trained through the following steps:
[0018] Obtain a sample problem and its corresponding sample answer and sample relationship sequence;
[0019] Construct a sample prompt word according to the sample problem, including a third text segment for instructing the relationship sequence generation model to generate a plan relationship sequence for assisting in answering the sample problem;
[0020] Fine-tune the pre-trained second large language model according to the sample prompt word, sample answer and sample relationship sequence to obtain the relationship sequence generation model.
[0021] In some possible implementation manners, the sample answer has a corresponding sample answer node in the text attribute graph; obtaining the sample relationship sequence includes:
[0022] According to the node information of each node, perform a third matching between the sample problem and each node to obtain a number of sample plan nodes;
[0023] Search for the relationship sequences that can reach the sample answer nodes from each sample plan node in the text attribute graph, and determine the search results as sample relationship sequences.
[0024] In some possible implementation manners, fine-tuning the pre-trained second large language model includes:
[0025] Construct a first probability distribution for each alternative relationship sequence according to the sample prompt words, sample answers, text attribute graph, and the sample relationship sequences;
[0026] Input the sample prompt words into the second large language model to obtain a second probability distribution of the second large language model for each alternative relationship sequence;
[0027] Fine-tune the second large language model according to the training loss determined by the difference between the first probability distribution and the second probability distribution.
[0028] In some possible implementation manners, the probability value of the first probability distribution when the alternative relationship sequence takes the value of the sample relationship sequence is greater than its probability value when the alternative relationship sequence takes the value of a non-sample relationship sequence.
[0029] In some possible implementation manners, the difference between the first probability distribution and the second probability distribution is determined according to the KL divergence.
[0030] In some possible implementation manners, the problem prompt words further include a second text segment, and the second text segment includes examples composed of several groups of sample problems and their corresponding sample relationship sequences; the relationship sequence generation model is a pre-trained third large language model.
[0031] In some possible implementation manners, the agent is obtained through the following steps:
[0032] Construct a target prompt word, which includes a fourth text segment defining the inference steps of the agent and a fifth text segment for exemplarily explaining the inference steps of the agent;
[0033] Input the target prompt word into the first large language model and configure it as an agent.
[0034] In some possible implementation manners, the agent inference steps include: an observation step of observing the current state, a thinking step of thinking based on the observed state, and an execution action step of performing corresponding actions based on the thinking result.
[0035] In some possible embodiments, the actions include: a search action of performing neighbor search on each planning node in the current set of planning nodes based on a sequence of planning relationships, a query action of retrieving similar nodes as planning nodes in the text attribute graph based on a problem text segment, and a completion action of completing an inference step and outputting a result node.
[0036] In some possible embodiments, the node information further includes text information; retrieving similar nodes as planning nodes in the text attribute graph based on a problem text segment includes:
[0037] Performing similarity matching between the embedding representations corresponding to the text information of each node and the embedding representation of the problem text segment respectively, and determining the nodes with the top similarity rankings as planning nodes.
[0038] In some possible embodiments, generating an answer based on the result node includes:
[0039] Determining the node names of each result node as the answer.
[0040] In some possible embodiments, the node information further includes text information; generating an answer based on the result node includes:
[0041] Determining the node names and text information of each result node as the answer.
[0042] In a second aspect, there is provided an apparatus for generating an answer based on a text attribute graph, including:
[0043] An acquisition unit configured to acquire a problem text and a text attribute graph, where the text attribute graph includes a plurality of nodes and relationship edges;
[0044] A first matching unit configured to perform a first matching between the problem text and each node according to the node information of each node to obtain a number of planning nodes; the node information includes at least a node name;
[0045] A second matching unit configured to perform a second matching between the problem text and each relationship sequence in the text attribute graph to obtain a number of planning relationship sequences; any relationship sequence includes a number of relationship edges connected end to end;
[0046] An answer generation unit configured to input the problem text and planning information into an agent, and cause the agent to traverse in the text attribute graph according to the problem text and planning information to obtain a number of result nodes, and generate an answer according to the result nodes; the agent is implemented based on a first large language model; the planning information includes the planning nodes and the sequence of planning relationships.
[0047] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0048] In a fourth aspect, a computing device is provided, including a memory and a processor. Among them, executable code is stored in the memory, and when the processor executes the executable code, the method of the first aspect is implemented.
[0049] The method and device for generating an answer based on a text attribute graph proposed in the embodiments of this specification first match the question text to be answered with the nodes and relationship edges in the text attribute graph respectively. Among them, the first match between the question text and the nodes obtains several planned nodes, and the second match between the question text and the relationship sequence containing head-to-tail connected relationship edges obtains several planned relationship sequences. Then, the obtained planned nodes and planned relationship sequences are used as auxiliary information and input into the agent together with the question text, and the agent is made to traverse in the text attribute graph according to the question text, planned nodes and planned relationship sequences to obtain several result nodes. Finally, the answer to the question can be generated based on the result nodes.
[0050] By jointly using the text information and relationship information in the semi-structured data (text attribute graph) in the embodiments of this specification, the large language model can more comprehensively and perfectly utilize various types of information in the semi-structured data when generating an answer, thereby improving the accuracy and relevance of the generated answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] To more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only the multiple embodiments disclosed in this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 A schematic diagram showing a text attribute graph according to an embodiment;
[0053] Figure 2 A schematic diagram showing an implementation scenario of a method for generating an answer based on a text attribute graph according to an embodiment;
[0054] Figure 3 A flowchart showing a method for generating an answer based on a text attribute graph according to an embodiment;
[0055] Figure 4 A flowchart showing a method for training a relationship sequence generation model according to an embodiment;
[0056] Figure 5 Schematic block diagram showing an apparatus for generating answers based on a text attribute graph according to an embodiment. Detailed implementation
[0057] The solution provided in this specification will be described below with reference to the accompanying drawings.
[0058] As mentioned above, existing RAG technologies typically focus on a single type of external data, such as a vectorized text database or a knowledge graph, and it is difficult to effectively process semi-structured data that contains both text information and relationship information. Specifically, technologies based on vectorized text databases require all information to be stored in text format, and then the content to be retrieved is separately matched with each text chunk in the text database. This method is difficult to utilize the inherent relationship information between text chunks and cannot perform step-by-step reasoning. Although knowledge graph-based technologies can traverse and reason along the paths (relationship edges) between different nodes in the knowledge graph, due to the limitations of the knowledge graph's own data format, it is difficult to store rich information on nodes and relationship edges, which limits the reasoning and expression capabilities of such methods.
[0059] To overcome the above problems, the embodiments of this specification propose a method for generating answers based on a text attribute graph. The method performs retrieval and reasoning based on semi-structured data, enabling large language models to simultaneously utilize the rich semantic information in text information and the rich relationship information between entities in knowledge graph data for reasoning and answer generation.
[0060] Semi-structured data can be represented as a heterogeneous text-attributed graph. Based on a conventional knowledge graph, all or some of the nodes also have their corresponding text data. Specifically, the graph contains multiple nodes and relationship edges, and any relationship edge is connected to two nodes. Nodes and relationship edges can each have multiple types. At least some of the nodes in the text attribute graph, in addition to having a node name, also include text information, which can be document information related to the node.
[0061] Formally, semi-structured data can be represented as a text attribute graph G = {V, E, R, f N , f R}, where V represents the set of nodes, E represents the set of relationship edges, R represents the set of relationship edge types, f N : V → text represents the mapping relationship from nodes to text information, where text represents text information, f R:E→R represents the mapping relationship from relationship edges to relationship edge types. It should be noted that the text attribute graph G is a directed graph. It can be understood that when the text obtained by mapping a certain node through f N is empty, it means that the node has no corresponding text information.
[0062] For example, for a text attribute graph related to transactions, it includes user nodes and commodity nodes. In addition to having the commodity name as the node name, the commodity node also includes text information in the form of documents. The text information of the commodity node can be the text information in the web page of the detailed page of the commodity on the shopping website, such as commodity descriptions, user evaluations, user Q&A, etc. In addition to having the user name as the node name, the user node also includes text information in the form of documents. The text information of the user node can be the text information in the personal space web page of the user on the shopping website, such as personal signatures, user dynamics, etc.
[0063] A more specific example can be as Figure 1 shown. Figure 1 FIG. shows a schematic diagram of a text attribute graph according to an embodiment. In Figure 1 it, the text attribute graph includes user nodes and commodity nodes. The user node includes the user name as the node name and also includes text information, which includes personal signatures and user dynamics; the commodity node includes the commodity name as the node name and also includes text information, which includes commodity types, commodity prices, commodity descriptions, user evaluations, user Q&A. The type of the relationship edge between the user node and the user node can be "friend", which is a bidirectional relationship edge. The type of the relationship edge between the user node and the commodity node can be "purchase", which is a unidirectional relationship edge. The type of the relationship edge between the commodity node and the commodity node can be "same type", which is a bidirectional relationship edge, or "gift", which is a unidirectional relationship edge.
[0064] It should be noted that Figure 1 only shows an example of the text attribute graph in a specific example, which does not constitute a limitation on the text attribute graph in the embodiments of this specification.
[0065] Figure 2 FIG. shows a schematic diagram of an implementation scenario of a method for generating answers based on a text attribute graph according to an embodiment. As Figure 2As shown in the figure, first, obtain the question text and the text attribute graph for assisting in generating answers. The text attribute graph contains multiple nodes and relational edges. Then, perform node matching between the question text and the nodes in the text attribute graph to obtain several planned nodes. Node matching specifically includes keyword matching in the question text according to the node names of the nodes, and determining the nodes with successful matching as planned nodes. Optionally, node matching also includes similarity matching between the embedding representations of the text information of the nodes and the embedding representation of the question text, and determining the nodes with the top similarity rankings as planned nodes.
[0066] Next, perform relational sequence matching between the question text and the relational sequences in the text attribute graph to obtain several planned relational sequences. The relational sequence includes several relational edges connected end to end in the text attribute graph. Relational sequence matching specifically includes using a sequence generation model based on a large language model, making it generate a relational sequence for assisting in answering the question text based on the question text, and using the output result of the large language model as the planned relational sequence. The detailed steps for using and training (fine-tuning) the sequence generation model will be described in the subsequent part.
[0067] Then, summarize the planned nodes and the planned relational sequences into planned information, and input them together with the question text into the agent. The agent traverses in the text attribute graph according to the question text and the planned information to obtain several result nodes, and generates answers based on the result nodes. The agent is implemented based on the first large language model.
[0068] Specifically, the reasoning steps of the agent can include three specific steps: the observation step, the thinking step, and the action execution step. After receiving the question text and the planned information, the agent observes based on the received content and adds the observed results to the observation results. Then, it thinks based on the observation results and performs corresponding actions according to the thinking results. If the action is to search or query, then the agent will perform the corresponding action in the text attribute graph and obtain the corresponding results. Then, the agent observes the obtained results, adds the newly observed results to the existing observation results, then thinks again based on the updated observation results, and performs corresponding actions according to the thinking results. This cycle is repeated for multiple rounds until the agent believes that its current observation results can answer the question text or reaches the preset round limit. Then, the agent will perform the completion action and output the result nodes according to the current observation results.
[0069] The detailed steps for configuring the agent will be described in the subsequent part.
[0070] The following combines specific embodiments to describe the specific implementation steps of the above method for generating answers based on the text attribute graph.
[0071] Figure 3 A flowchart showing a method for generating an answer based on a text attribute graph according to an embodiment. The execution entity of the method can be any platform, server, or device cluster with computing and processing capabilities. As Figure 3 shown, the method at least includes: Step 302, obtaining a question text and a text attribute graph, where the text attribute graph includes multiple nodes and relationship edges; Step 304, according to the node information of each node, performing a first match between the question text and each node to obtain a number of planned nodes; the node information at least includes a node name; Step 306, performing a second match between the question text and each relationship sequence in the text attribute graph to obtain a number of planned relationship sequences; any relationship sequence includes a number of relationship edges connected end to end; Step 308, inputting the question text and planned information into an agent, and making the agent traverse in the text attribute graph according to the question text and planned information to obtain a number of result nodes, and generating an answer based on the result nodes; the agent is implemented based on a first large language model; the planned information includes the planned nodes and planned relationship sequences.
[0072] The specific execution processes of the above steps are described below.
[0073] First, in Step 302, obtain a question text and a text attribute graph, where the text attribute graph includes multiple nodes and relationship edges.
[0074] The nodes in the text attribute graph include node names, and at least one node has corresponding text information. The text information is plain text data for recording details of the node itself and / or related to the node.
[0075] The text attribute graph G can include data in any field, such as data related to transactions. The question text can be denoted as q, which is plain text data. The method for generating an answer based on the text attribute graph described in the embodiments of this specification is used to determine a number of result nodes in the text attribute graph G for answering the question text q.
[0076] Then, in Step 304, according to the node information of each node, perform a first match between the question text and each node to obtain a number of planned nodes; the node information at least includes a node name.
[0077] Step 304 is used to find nodes in the text attribute graph G that have potential relevance to the question text q, called planned nodes. The set of planned nodes containing the planned nodes related to the question text q can be denoted as V q 。
[0078] In one embodiment, the first match between the question text and each node in Step 304 includes:
[0079] Taking the node names of each of the said nodes as keywords respectively, perform keyword matching in the said problem text, and determine each node with successful matching as a planned node.
[0080] When the node name of a certain node appears in the said problem text q, determine this node as a planned node and add it to V q .
[0081] The above embodiments can have a good matching effect when the words and expressions in the problem text q itself are clear and the node names of each node in the text attribute graph G also use standardized naming.
[0082] In a more specific embodiment, in order to make up for the deficiencies in the matching results of the above embodiments, the first matching in step 304 may further include similarity-based matching for supplementing the results in V q . In this embodiment, the node information further includes text information. The first matching of the said problem text with each node in step 304 further includes:
[0083] Perform similarity matching between the embedding representations corresponding to the text information of each of the said nodes and the embedding representation of the said problem text respectively, and determine each node with a higher similarity ranking as a planned node.
[0084] The embedding representations (embeddings) corresponding to the text information of each of the foregoing nodes, as well as the embedding representation of the problem text q, can be encoded using any text encoder, such as a Word2Vec encoder, a BERT (Bidirectional Encoder Representations from Transformers) encoder, etc., which are not limited here.
[0085] The foregoing similarity matching can be based on any similarity measurement method, such as cosine similarity, Euclidean distance, etc., which are not limited here.
[0086] When the similarity between the embedding representation corresponding to the text information of a certain node and the embedding representation corresponding to the problem text q ranks among the top (for example, among the top K, where K is a preset value), determine this node as a planned node and add it to V q .
[0087] The above method of similarity-based matching can be used as a supplement to keyword matching and is executed when the number of planned nodes obtained by keyword matching is less than a preset threshold.
[0088] In other embodiments, other methods can also be used to implement the first matching. For example, the question text can be segmented into multiple words, and then the embedding representations corresponding to the node names of the nodes are matched with the embedding representations of each word in the question text, and the nodes with a similarity exceeding a preset threshold are determined as the planned nodes.
[0089] Next, in step 306, the question text is secondarily matched with each relationship sequence in the text attribute graph to obtain a number of planned relationship sequences; any relationship sequence includes a number of relationship edges connected end to end.
[0090] In the text attribute graph G, if there exists a series of triples connected end to end, where the head node of any triple is the tail node of the previous triple, and the tail node of this triple is the head node of the next triple, then the respective relationship edges of this series of triples are called relationship edges connected end to end. The set of relationship sequences formed by all the relationship sequences in the text attribute graph G can be denoted as Z. It should be noted that a relationship sequence can contain only one relationship edge.
[0091] In some cases, the answer corresponding to the question text q can be directly inferred from the question text q itself without using the relationship sequences in the text attribute graph G to answer the question. Therefore, in one embodiment, the question text q itself is also defined as a special relationship edge. Furthermore, the respective relationship sequences also include relationship sequences containing the question text as a relationship edge.
[0092] Step 306 is used to find a relationship sequence in the text attribute graph G that has a potential association with the question text q, which is called a planned relationship sequence. The set of planned relationship sequences containing the planned relationship sequences related to the question text q can be denoted as Z q . Due to the potential relationship between the relationship sequence and the question text, which is more complex than the potential relationship between the node and the question text, the embodiments of this specification use a model to learn and reason about the relationship between the question text and the relationship sequence.
[0093] In one embodiment, step 306 includes step 3062 and step 3064.
[0094] In step 3062, a question prompt word is constructed according to the question text, which includes a first text segment for instructing the relationship sequence generation model to generate a planned relationship sequence for assisting in answering the question text.
[0095] In step 3064, the question prompt word is input into the relationship sequence generation model to obtain a number of planned relationship sequences.
[0096] Multiple forms of the first text segment can be used as long as they can instruct the relationship sequence generation model to generate a planned relationship sequence for assisting in answering the question text. For example, in a more specific embodiment, the first text segment includes: "Please generate a valid relationship sequence to help answer the following question: [question text]". Here, [question text] is a placeholder for filling in the specific content of the question text.
[0097] The relationship sequence generation model can be fine-tuned using a sample training set based on a pre-trained large language model; or it can directly use a pre-trained large language model and then use in-context learning to enable the large language model to generate a planned relationship sequence according to the given examples.
[0098] In one embodiment, the relationship sequence generation model is obtained by fine-tuning a pre-trained large language model using a sample training set. In this embodiment, the relationship sequence generation model is trained through steps 402 to 406 as shown in Figure 4 shown. Figure 4 The flowchart shows a method for training a relationship sequence generation model according to an embodiment.
[0099] In step 402, a sample question, its corresponding sample answer, and a sample relationship sequence are obtained.
[0100] The sample question and its corresponding sample answer can be obtained from a sample training set in the relevant field, which will not be elaborated here. In one embodiment, the sample answer has a corresponding sample answer node in the text attribute graph. The obtaining of the sample relationship sequence in step 402 includes steps 4022 and 4024.
[0101] In step 4022, according to the node information of each node, the sample question is third-matched with each node to obtain a number of sample plan nodes.
[0102] The specific steps of the third matching in step 4022 can refer to the method in step 304 and its embodiments, which will not be elaborated here.
[0103] Then, in step 4024, a relationship sequence that can reach the sample answer node from each sample plan node is searched in the text attribute graph, and the search result is determined as the sample relationship sequence.
[0104] For any target sample plan node in each sample plan node, if there is a relationship sequence whose starting node is the target sample plan node and whose ending node is the sample answer node, it is said that the target sample plan node "can reach" the sample answer node, and at the same time, this relationship sequence is determined as the sample relationship sequence.
[0105] The sample relationship sequence is a relationship sequence that has a potential relationship with the sample question. By gradually reasoning along the sample relationship sequence, the sample answer node can be finally reached, and this sample answer node can be used to answer the sample question. Therefore, let the large language model learn the relationship between the sample question and the sample relationship sequence, so that when performing the second matching based on the question text, a relationship sequence that has a potential relationship with the question text, that is, the planned relationship sequence, can be found.
[0106] Return to Figure 4 , and then, in step 404, construct a sample prompt word according to the sample question, which includes a third text segment used to instruct the relationship sequence generation model to generate a planned relationship sequence for assisting in answering the sample question.
[0107] The specific steps of constructing the sample prompt word can refer to step 3062 and will not be elaborated here.
[0108] In step 406, fine-tune the pre-trained second large language model according to the sample prompt word, the sample answer, and the sample relationship sequence to obtain a relationship sequence generation model.
[0109] The second large language model can be any pre-trained large language model. When fine-tuning, the sample relationship sequence can be used as the ground-truth corresponding to the sample prompt word, and a training loss can be constructed to fine-tune the second large language model.
[0110] In one embodiment, step 406 includes steps 4062 to 4066.
[0111] In step 4062, construct a first probability distribution for each alternative relationship sequence according to the sample prompt word, the sample answer, the text attribute graph, and the sample relationship sequence.
[0112] The first probability distribution is the probability distribution of the relationship alternative relationship sequence, which can be used as the ground-truth or label to fine-tune the large language model. Therefore, this first probability distribution should be mainly distributed at the position where the alternative relationship sequence takes the value of the sample relationship sequence. Specifically, in one embodiment, the probability value of the first probability distribution when the alternative relationship sequence takes the value of the sample relationship sequence is greater than its probability value when the alternative relationship sequence takes the value of a non-sample relationship sequence.
[0113] In a more specific embodiment, it can be set that each point taking the value of the sample relationship sequence evenly divides the total probability, and at the same time, the probability value of each point taking the value of a non-sample relationship sequence is 0. At this time, the first probability distribution can be as shown in formula (1):
[0114]
[0115] Among them, Q(z|a′, q′, G) represents the first probability distribution, z represents the alternative relationship sequence, a′ is the sample answer, q′ is the sample question, and Z q′ is the set of sample relationship sequences, and |Z q′ | represents the number of elements in the set, that is, the number of sample relationship sequences.
[0116] Then, in step 4064, the sample prompt words are input into the second large language model to obtain the second probability distribution of the second large language model for each alternative relationship sequence.
[0117] The second probability distribution can be denoted as P θ (z|q′), where θ represents the set of trainable parameters in the large language model.
[0118] Next, in step 4066, the second large language model is fine-tuned according to the training loss determined by the difference between the first probability distribution and the second probability distribution.
[0119] In one embodiment, the difference between the first probability distribution and the second probability distribution is determined according to the KL divergence (Kullback-Leibler divergence), as shown in formula (2).
[0120]
[0121] The first term of formula (2) is a value independent of θ. Therefore, only the latter term can be taken as the training loss. At this time, the training loss L is as shown in formula (3):
[0122]
[0123] The second large language model can be fine-tuned based on the training loss L, and then a relationship sequence generation model can be obtained.
[0124] In other embodiments, other methods can also be used to measure the difference between the first probability distribution and the second probability distribution, and then determine the training loss, such as using Jensen-Shannon divergence, etc., which will not be elaborated here.
[0125] It can be understood that the above describes the steps of fine-tuning the large language model based on a set of sample questions, sample answers, and sample relationship sequences. In other embodiments, the large language model can also be fine-tuned one or more rounds using a method similar to the foregoing steps based on multiple sets of sample questions, sample answers, and sample relationship sequences to improve the accuracy of the finally obtained relationship sequence generation model, which will not be elaborated here.
[0126] The fine-tuned relationship sequence generation model can generate corresponding planned relationship sequences according to the received question text.
[0127] In another embodiment, step 3064 can be based on context learning to enable a large language model to generate a planned relationship sequence. In this embodiment, the relationship sequence generation model is a pre-trained third large language model. The question prompt also includes a second text segment, and the second text segment includes examples composed of several groups of sample questions and their corresponding sample relationship sequences.
[0128] The third large language model can be any pre-trained large language model. Various forms of the second text segment can be used as long as it can include examples composed of several groups of sample questions and their corresponding sample relationship sequences. For example, in a more specific embodiment, the second text segment includes: "The following are multiple groups of examples: [Sample Question 1]: [Sample Relationship Sequence 1]; [Sample Question 2]: [Sample Relationship Sequence 2];... [Sample Question n]: [Sample Relationship Sequence n]." Among them, [Sample Question 1], [Sample Relationship Sequence 1], [Sample Question 2], [Sample Relationship Sequence 2], [Sample Question n], [Sample Relationship Sequence n] are placeholders for filling in corresponding specific contents.
[0129] The second text segment activates and utilizes the existing capabilities of the large language model by adding examples to the question prompt, so as to enable the large language model to respond immediately to new tasks and application scenarios without parameter updates or additional data training.
[0130] In other embodiments, other methods can also be used to perform a second matching between the question text and each relationship sequence in the text attribute graph. For example, the aforementioned sample questions and sample relationship sequences can be used to train a neural network model, and then the trained neural network model receives the embedded representation of the question text and outputs a planned relationship sequence. Or, the question text can be tokenized to obtain multiple words, and then the embedded representations corresponding to the types of each relationship edge in the relationship sequence are matched with the embedded representations of each word in the question text, and the relationship sequence with a similarity exceeding a preset threshold is determined as the planned relationship sequence.
[0131] Steps 304 and 306 respectively describe the specific steps of finding relevant planned nodes and relevant planned relationship sequences in the text attribute graph G according to the question text q.
[0132] Back to Figure 3, finally, in step 308, the problem text and the plan information are input into the agent, and the agent traverses in the text attribute graph according to the problem text and the plan information to obtain a number of result nodes, and generates an answer based on the result nodes; the agent is implemented based on a first large language model; the plan information includes the plan nodes and the plan relationship sequence.
[0133] First, the functions and inference processes of the agent are described below, and then the steps of configuring the agent according to the pre-trained first large language model are described.
[0134] An agent can be configured using a pre-trained large language model via specific prompt words. The agent performs corresponding traversal and inference in the text attribute graph according to the received problem text, plan nodes, and plan relationship sequence, and finally obtains a number of result nodes.
[0135] Specifically, the inference steps of the agent include an observation step, a thinking step, and an action execution step. The agent will loop between these steps until the corresponding result nodes are found or the preset loop iteration upper limit is reached.
[0136] The agent will maintain a set of plan nodes V t , where t is the round of inference. Initially, this set of plan nodes V 1 is the set of plan nodes V corresponding to the problem text q q .
[0137] In the observation step, the agent observes the current context content and its own traversal and inference results to obtain corresponding observation results. In the thinking step, the agent thinks based on the observation results and decides which action to perform. In the action execution step, the agent executes the specific action it decides in the thinking step.
[0138] Actions include search actions, query actions, and completion actions.
[0139] When the agent chooses to execute a search action, the agent performs neighbor search based on each plan node in the current set of plan nodes V t and adds the searched nodes to V t to form the set of plan nodes V used in the next round t+1 . For the first plan node in each plan node, neighbor search is performed based on the first plan relationship sequence in each plan relationship sequence. Specifically, the first plan relationship sequence is split into several first relationship edges, and then neighbor nodes of the first plan node are found based on each first relationship edge, and the search result is determined as the neighbor search result of the first plan node based on the first plan relationship sequence.
[0140] In a more specific embodiment, when there are too many search results for neighbor search based on the planned relationship sequence for each planned node in the current set of planned nodes V t the agent can select the top M nodes with a high relevance to the reasoning from them, and then add these nodes to V t to form the set of planned nodes V used in the next round t+1 .
[0141] In some cases, the planned nodes matched by the first match in step 304 may not meet the reasoning requirements of the agent. For example, when the problem text q contains multiple conditions (or sub-problems, e.g., "What is the result of simultaneously satisfying condition 1 and condition 2?") and one of the conditions fails to match a planned node with a high relevance in the first match of step 304, the agent may choose to execute a query action to re-query in the text attribute graph based on a certain fragment (or the whole) of the problem text
[0142] When the agent chooses to execute a query action, the agent first selects the fragment of the problem text to be queried (this fragment can also be the problem text itself), and then performs a similarity match with the text information of each node in the text attribute graph, and adds several nodes with a high similarity ranking to V t to form the set of planned nodes V used in the next round t+1 .
[0143] When the agent believes that the current state information can already obtain the answer, the agent will choose to execute a completion action, jump out of the aforementioned step loop, and output the result node
[0144] The above describes the reasoning process for the agent to obtain the result node. The following describes the steps to configure the agent according to the pre-trained first large language model
[0145] In one embodiment, the agent is obtained through the following steps
[0146] Construct a target prompt word, which includes a fourth text fragment defining the reasoning steps of the agent and a fifth text fragment for exemplarily explaining the reasoning steps of the agent. Then, input the target prompt word into the first large language model to configure it as an agent
[0147] In a more specific embodiment, the agent reasoning steps include: an observation step of observing the current state, a thinking step of thinking based on the observed state, and an execution action step of performing corresponding actions based on the thinking result
[0148] In a more specific embodiment, the actions include: a search action of performing neighbor search on each planning node in the current set of planning nodes based on the sequence of planning relationships, a query action of retrieving similar nodes in the text attribute graph as planning nodes based on the problem text segment, and a completion action of completing the inference step and outputting the result nodes.
[0149] In a more specific embodiment, the node information further includes text information; retrieving similar nodes in the text attribute graph as planning nodes based on the problem text segment includes: respectively performing similarity matching between the embedding representations corresponding to the text information of each node and the embedding representation of the problem text segment, and determining the nodes with the top similarity rankings as planning nodes.
[0150] In summary, in one embodiment, the fourth text segment may include: "To perform the question - answering task, alternate between thinking steps, action steps, and observation steps. The thinking step is used to reason about the current situation. The action step includes the following three actions: (1) Search [set of planning nodes]:... (2) Query [text segment]:... (3) Complete [set of result nodes]:... Tabs should be used to separate nodes. You should avoid redundant words when generating each step." The ellipsis parts in the fourth text segment are used to list the specific descriptions of each action.
[0151] The fifth text segment may include: "Here are some examples: Question:... Set of planning nodes:... Thinking 1:... Action 1: Search / Query... Observation 1:... Thinking 2:... Action 2: Search / Query... Observation 2:... Thinking n:... Action n: Complete..." The ellipsis parts in the fifth text segment are used to list the specific content of each thinking step.
[0152] After obtaining the result nodes, an answer can be generated based on the result nodes. In different embodiments, answers with different levels of detail can be generated according to specific requirements.
[0153] In one embodiment, generating an answer based on the result nodes includes: determining the node names of each result node as the answer.
[0154] In another embodiment, the node information further includes text information; generating an answer based on the result nodes includes: determining the node names and text information of each result node as the answer.
[0155] The method for generating an answer based on a text attribute graph proposed in the embodiments of this specification first matches the text information and relationship information related to the question text in the text attribute graph according to the question text to be answered, and then configures an agent based on a large language model, and uses the agent to jointly use the text information and relationship information to reason about the question text, so that the large language model can more comprehensively and perfectly utilize various types of information in the semi-structured data when generating an answer, thereby improving the accuracy and relevance of the generated answer.
[0156] According to an embodiment of another aspect, there is also provided an apparatus for generating an answer based on a text attribute graph. Figure 5 A schematic block diagram showing an apparatus for generating an answer based on a text attribute graph according to an embodiment. The apparatus can be deployed in any device, platform or device cluster with computing and processing capabilities. As Figure 5 shown, the apparatus 500 includes:
[0157] An acquisition unit 502, configured to acquire a question text and a text attribute graph, where the text attribute graph includes a plurality of nodes and relationship edges;
[0158] A first matching unit 504, configured to perform a first match between the question text and each node according to the node information of each node, and obtain a number of planned nodes; the node information includes at least a node name;
[0159] A second matching unit 506, configured to perform a second match between the question text and each relationship sequence in the text attribute graph, and obtain a number of planned relationship sequences; any relationship sequence includes a plurality of relationship edges connected end to end;
[0160] An answer generation unit 508, configured to input the question text and planned information into an agent, and cause the agent to traverse in the text attribute graph according to the question text and planned information, obtain a number of result nodes, and generate an answer according to the result nodes; the agent is implemented based on a first large language model; the planned information includes the planned nodes and planned relationship sequences.
[0161] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is caused to execute the method described in any of the above embodiments.
[0162] According to an embodiment of still another aspect, there is also provided a computing device, including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method described in any of the above embodiments is implemented.
[0163] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the related content.
[0164] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0165] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or apparatus comprising the element.
[0166] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0167] The specific implementation manners described above further elaborate on the purpose, technical solution and beneficial effects of the present invention. It should be understood that the above are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for generating answers based on a text attribute graph, comprising: Obtaining a question text and a text attribute graph, wherein the text attribute graph includes a plurality of nodes and relationship edges; According to the node information of each node, the question text is first matched with each node to obtain a plurality of plan nodes; the node information at least includes the node name; Performing a second matching between the question text and each relation sequence in the text attribute graph to obtain a plurality of planned relation sequences; Any relationship sequence consists of several relationship edges connected end to end; The question text and plan information are input into the intelligent agent, and the intelligent agent is made to traverse the text attribute graph according to the question text and plan information to obtain a number of result nodes, and generate answers according to the result nodes; the intelligent agent is implemented based on the first language model; the plan information includes the plan nodes and plan relationship sequences.
2. The method according to claim 1, wherein: The question text is first matched with each node, including: The node names of the nodes are respectively used as keywords, keyword matching is performed in the question text, and each node that successfully matches is determined as a planned node.
3. The method according to claim 2, wherein: The node information also includes text information; first matching the question text with each node also includes: The embedded representations corresponding to the text information of each node are matched with the embedded representations of the question text respectively in similarity, and each node with a high similarity ranking is determined as a planned node.
4. The method according to claim 1, wherein: Performing a second matching between the question text and each relation sequence in the text attribute graph includes: Constructing question prompt words according to the question text, including a first text segment for instructing the relation sequence generation model to generate a planned relation sequence for assisting in answering the question text; The question prompt words are input into a relation sequence generation model to obtain a number of planned relation sequences.
5. The method according to claim 4, wherein: The relationship sequence generation model is trained by the following steps: Obtain sample questions and their corresponding sample answers and sample relationship sequences; Constructing sample prompt words according to the sample question, which includes a third text segment for instructing the relation sequence generation model to generate a planned relation sequence for assisting in answering the sample question; According to the sample prompt words, sample answers and sample relationship sequences, the pre-trained second largest language model is fine-tuned to obtain a relationship sequence generation model.
6. The method according to claim 5, wherein: The sample answer has a corresponding sample answer node in the text attribute graph; Get sample relationship sequences, including: According to the node information of each node, the sample problem is matched with each node to obtain a number of sample plan nodes; In the text attribute graph, a relationship sequence that can reach a sample answer node from each sample plan node is searched, and the search result is determined as a sample relationship sequence.
7. The method according to claim 5, wherein: Fine-tune the pre-trained second largest language model, including: Constructing a first probability distribution for each candidate relationship sequence according to the sample prompt word, the sample answer, the text attribute graph and the sample relationship sequence; Inputting the sample prompt word into the second largest language model to obtain a second probability distribution of each candidate relation sequence of the second largest language model; The second large language model is fine-tuned according to a training loss determined by a difference between the first probability distribution and the second probability distribution.
8. The method according to claim 7, wherein: The probability value of the first probability distribution when the candidate relationship sequence is a sample relationship sequence is greater than the probability value when the candidate relationship sequence is a non-sample relationship sequence.
9. The method according to claim 7, wherein: The difference between the first probability distribution and the second probability distribution is determined according to KL divergence.
10. The method according to claim 4, wherein: The question prompt word also includes a second text segment, which includes an example consisting of several groups of sample questions and their corresponding sample relationship sequences; the relationship sequence generation model is a pre-trained third language model.
11. The method according to claim 1, wherein: The agent is obtained by the following steps: constructing a target prompt word, which includes a fourth text segment defining the agent's reasoning steps and a fifth text segment for exemplifying the agent's reasoning steps; The target prompt word is input into the first large language model and configured as an intelligent agent.
12. The method according to claim 11, wherein: The agent reasoning steps include: an observation step of observing the current state, a thinking step of thinking based on the observed state, and an execution step of executing corresponding actions based on the thinking results.
13. The method according to claim 12, wherein: The actions include: a search action of performing a neighbor search based on a plan relationship sequence for each plan node in the current plan node set, a query action of retrieving similar nodes as plan nodes in a text attribute graph based on a question text fragment, and a completion action of completing the reasoning steps and outputting a result node.
14. The method according to claim 13, wherein: The node information also includes text information; based on the problem text fragment, retrieving similar nodes in the text attribute graph as plan nodes includes: The embedded representations corresponding to the text information of each node are matched with the embedded representations of the question text fragments in terms of similarity, and each node with a high similarity ranking is determined as a planned node.
15. The method according to claim 1, wherein: Generate an answer according to the result node, including: The node name of each result node is determined as the answer.
16. The method according to claim 1, wherein: The node information also includes text information; generating an answer according to the result node includes: The node name and text information of each result node are determined as the answer.
17. A device for generating an answer based on a text attribute graph, comprising: An acquisition unit is configured to acquire a question text and a text attribute graph, wherein the text attribute graph includes a plurality of nodes and relationship edges; A first matching unit is configured to perform a first matching between the question text and each node according to node information of each node to obtain a plurality of plan nodes; the node information at least includes a node name; A second matching unit is configured to perform a second matching between the question text and each relationship sequence in the text attribute graph to obtain a plurality of planned relationship sequences; Any relationship sequence consists of several relationship edges connected end to end; The answer generation unit is configured to input the question text and plan information into an intelligent agent, enable the intelligent agent to traverse the text attribute graph according to the question text and plan information, obtain a number of result nodes, and generate answers according to the result nodes; the intelligent agent is implemented based on the first language model; the plan information includes the plan nodes and plan relationship sequences.
18. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 16.
19. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 16 is implemented.