Intelligent question and answer method based on multi-hop similar subgraph and large language model

By constructing a multi-hop similar subgraph and a low-rank matrix split fine-tuning large language model, the problem of insufficient utilization of knowledge graph semantic information and illusion of large language models in the existing technology is solved, and the accuracy and reliability of question-and-answer are improved.

CN120387514APending Publication Date: 2025-07-29ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510440276.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the existing domain knowledge question and answer technology, problem retrieval fails to make full use of knowledge graph semantic information based on semantic similarity, and large language models are prone to hallucinations in professional knowledge tasks.

Method used

Using a combination of multi-hop similar subgraphs and large language models, features are extracted through pre-training models, multi-hop similar subgraphs are constructed, and a low-rank matrix is used to split and fine-tune the large language model to generate answers.

Benefits of technology

The accuracy of knowledge graph-related subgraph retrieval and the accuracy of generated answers of large language models in domain knowledge question-and-answer tasks are improved, and hallucination phenomena are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387514A_ABST
    Figure CN120387514A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent question and answer method based on a multi-hop similar subgraph and a large language model, and belongs to the technical field of natural language processing. The method comprises the following steps: S1, respectively extracting characteristics of a question, a knowledge graph entity and a knowledge graph relationship by using a pre-training model; s2, constructing a multi-hop similar sub-graph based on the features extracted in the step S1, and outputting an answer sub-graph; s3, splitting the low-rank matrix according to the step length, and finely adjusting the large language model; and S4, inputting the questions and the answer sub-graphs output in the step S2 into the large language model processed in the step S3 for processing, and generating answers. By the adoption of the technical scheme, semantic information of the knowledge graph can be effectively utilized, generation of the large language model is enhanced by retrieving the similarity subgraph, the illusion problem possibly occurring in the large language model generation process is reduced, and the generation capacity of the large language model on a specific task is improved by conducting fine adjustment on the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing in artificial intelligence, and relates to an intelligent knowledge answering method based on multi-hop similar subgraphs and large language models. Background Art

[0002] Intelligent knowledge answering systems have been widely used in fields such as intelligent manufacturing and intelligent customer service. Such systems focus on solving tasks in specific fields and help users achieve their goals through natural language interaction. Large language models have achieved remarkable results in multiple natural language processing fields. Models such as Seq2Seq, Transformer, and BERT can effectively process complex data and are widely used in knowledge answering due to their excellent natural semantic understanding and text generation capabilities. At the same time, the introduction of knowledge graphs provides necessary domain knowledge support for domain knowledge answering systems. A knowledge graph is one of the key technologies of semantic networks. It describes entities and their relationships in a structured manner through a graph data structure, thereby revealing the connections between entities. Essentially, a knowledge graph connects various entities through certain relationships to form a structured network, further revealing the relationships between entities. Therefore, combining domain knowledge graphs with large language models for knowledge answering is an effective way to utilize knowledge.

[0003] However, the existing domain knowledge graph answering technology still has the following deficiencies: First, when retrieving questions, the answering technology mostly uses semantic similarity calculation for matching. Such methods only match the similarity between the question and the knowledge graph entities and do not utilize the graph structure and semantic information of the knowledge graph, resulting in low accuracy when dealing with professional terms in domain knowledge. Second, although large language models can perform natural language understanding and generation through a large number of parameters, in tasks that require precise reasoning and professional knowledge, due to the lack of relevant domain knowledge, hallucination phenomena will still occur.

[0004] After retrieval, the Chinese patent application number is 2023101009939, the application date is: April 25, 2023, and the invention title is: A Knowledge Graph Question Answering Method and System Based on the Power Grid Hidden Danger Investigation Scenario. This method provides a knowledge graph question answering method and system based on the power grid hidden danger investigation scenario, including obtaining question corpora in the power grid hidden danger investigation scenario, sorting out intent templates and question templates; using a pre-constructed professional domain model to perform named entity recognition on the question corpora to obtain named entity recognition results; determining whether there are specific entity categories and specific intent words in the input question sentence, matching with the intent templates for intent judgment, and classifying them into the question templates; generating cypher statements for the Neo4j graph database according to the results returned by entity extraction and intent classification; connecting to the neo4j graph database according to the generated cypher statements for querying to obtain corresponding query results; formatting the query results and outputting them back to the user. However, this method generates cypher statements to query the neo4j graph database after performing named entity recognition on the question corpora. The existing problem is that only the entities in the question corpora are recognized, and the semantic information in the knowledge graph cannot be utilized. In addition, after the results are queried, they are directly returned to the user, and the user receives discrete triple information, which cannot provide help to the user.

[0005] For another example, the Chinese patent application number is: 2025100808280, the application date is: February 25, 2025, and the invention title is: A Chinese Medical Question Answering Method and Device Based on a Knowledge Graph and a Large Language Model. This method inputs the initial question data into the first large language model for parsing and processing to obtain symptom information, symptom core entities, and question information; matches the symptom core entities with the entities in a pre-constructed knowledge graph to obtain target entities, where the target entities are entities in the knowledge graph with a semantic matching degree greater than a threshold with the symptom core entities; performs relevant knowledge retrieval in the knowledge graph based on the target entities to obtain target triples; converts the target triples into input information in text form; inputs the input information, the symptom information, and the question information into the second large language model for question answering processing to obtain a target answer. However, the problem with this method is that when using a large language model to answer medical questions, due to the lack of domain knowledge in the task of professional knowledge, hallucination problems are likely to occur. Summary of the Invention

[0006] 1. Problems to be Solved

[0007] To solve the problems in the existing domain knowledge question answering technology that the retrieval of questions is only based on semantic similarity, the semantic information in the knowledge graph cannot be fully utilized, and the hallucination problems that are likely to occur when using a large language model to generate answers for domain knowledge, the present invention provides an intelligent question answering method based on multi-hop similar subgraphs and a large language model.

[0008] 2. Technical Solution

[0009] To solve the above problems, the technical solution adopted by the present invention is as follows:

[0010] The present invention provides an intelligent question-answering method based on multi-hop similar subgraphs and large language models, including the following steps:

[0011] S1. Use a pre-trained model to extract the features of questions, knowledge graph entities, and knowledge graph relationships respectively;

[0012] S2. Construct a multi-hop similar subgraph based on the features extracted in step S1, and output an answer subgraph;

[0013] S3. Split the low-rank matrix according to the step size and fine-tune the large language model;

[0014] S4. Input the question and the answer subgraph output in step S2 into the large language model processed in step S3 for processing to generate an answer.

[0015] Furthermore, in step S1, a pre-trained model is used to extract the feature vectors of questions, knowledge graph entities, and knowledge graph relationships respectively.

[0016] Furthermore, in step S1, by calculating the similarity between the question feature vector and the knowledge graph entity feature vector, and combining a set threshold, a starting point set is determined. The starting point set is used to store all starting point entities. For each starting point entity in the starting point set, the attention information between the starting point entity and its neighbor nodes, and the attention information between the question and the neighbor nodes of the starting point entity are aggregated using the attention mechanism.

[0017] Furthermore, the cosine similarity or Manhattan distance method is used to calculate the similarity between the question and the knowledge graph entity. When the similarity is greater than the set threshold, the starting point entity is included in the starting point set.

[0018] Furthermore, in step S2, first, starting from the starting point entity, the action score is calculated according to the prior score, action value, and access times of the current action, and the action with the maximum score is included in the subgraph; then, the entity value is evaluated through a rolling model and a key model, and the entity value is backpropagated to update the access times and action values on the path; finally, the subgraph with the highest action value is selected from all subgraphs as the answer subgraph.

[0019] Furthermore, in step S2, for each starting point entity in the starting point set, the entities and relationships connected to the starting point entity are recorded as a selection action. For each selection action, a set of statistics is stored as follows:

[0020] {P(e, a), Q(e, a), N(e, a)}

[0021] In the formula:

[0022] e represents the feature vector of the current starting entity;

[0023] a represents the action of selecting a neighbor entity, which includes the selected relationship and the feature vector of the entity;

[0024] P(e, a) is the prior score, which calculates the probability of matching the feature vector e of the starting entity with each action a using a large language model;

[0025] Q(e, a) is the action value for evaluating the entity's selection of action a, with an initial value of 0;

[0026] N(e, a) is the number of visits, representing the number of times the current action has been visited, with an initial value of 0;

[0027] The calculation formula for the action score is as follows:

[0028]

[0029] Score(a) is the action score for the current entity to select action a;

[0030] c puct is a constant for balancing exploration nodes and exploitation nodes;

[0031] N parent (a) is the number of visits of the previous action of the current action;

[0032] Start retrieving nodes from the starting entity, select the action with the highest action score value, and include it in the multi-hop similar subgraph. If the action score values are equal, select the node with the highest prior score to include in the multi-hop similar subgraph.

[0033] Furthermore, evaluate the value of the newly added entity. The evaluation method is as follows:

[0034] For the entity feature vector newly added to the similar subgraph, denote it as e new , and the query feature vector is denoted as d. represents the aggregated result vector between the query feature vector and the neighbor entity feature vectors of the newly added entity. Calculate the entity value according to the following formula:

[0035]

[0036] In the formula: is the value evaluation result value of the newly added entity node;

[0037] λ and μ are weight parameters used to assign weights to the rolling model result value, the key model result value, and the reward value;

[0038] V roll-out (e new ) is the result value of the rolling model based on the pre-trained LLM, which is used to predict the future reward of the entity;

[0039] is the result value of the key model based on BERT, which is used to calculate the aggregated result vector whether the answer semantically conforms to the question feature vector d;

[0040] r(e new ) is the reward value. If the entity corresponding to the feature vector e new is the answer, then r(e new ) = 1; otherwise, r(e new ) = -1.

[0041] Furthermore, the entity value of the newly added entity node is backpropagated through the graph structure to update the action value and the access count of all ancestor nodes on the update route. The update rules are as follows:

[0042] N(e,a) = N(e,a)’ + 1

[0043]

[0044] N(e,a)’ is the access count of the previous state;

[0045] N(e,a) is the current access count;

[0046] Q(e,a) is the action value for evaluating the entity's selection of action a, and e is the current entity feature vector;

[0047] is the indicator function, which takes 1 when the entity corresponding to the feature vector e transfers to the entity corresponding to the feature vector e new in the state of action a, and takes 0 otherwise;

[0048] represents the value evaluation result value of e new in the j-th simulation, and j is the simulation round.

[0049] Furthermore, after the backpropagation update is completed, the operation of step S2 is repeated for the next starting entity. The end conditions for constructing the multi-hop similarity subgraph are as follows:

[0050] (1) Find the answer node; and / or

[0051] (2) Reach the set number of entity retrievals; and / or

[0052] (3) The current entity has no action for extension.

[0053] Furthermore, step S3 includes the following steps:

[0054] S3.1. Perform low-rank decomposition on the weight matrix of the large language model to obtain two low-rank matrices;

[0055] S3.2. Further decompose the two low-rank matrices respectively according to the splitting step length, and then introduce the Hadamard product calculation to obtain the final weight matrix.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] An intelligent question-answering method based on multi-hop similar subgraphs and large language models provided by the present invention, after using semantic similarity to query the starting entity similar to the question, utilizes the semantic information of the knowledge graph, and solves the problem of being unable to utilize the semantic information of the knowledge graph by aggregating the neighbor node information of the starting entity;

[0058] At the same time, when constructing the subgraph, the value of each step is evaluated, improving the accuracy of retrieving relevant subgraphs of the knowledge graph. More optimally, the large language model is fine-tuned by introducing the Hadamard product calculation to split and calculate the weight matrix to obtain the final weight matrix, enabling it to perform better in the task of domain knowledge questions, and significantly improving the accuracy of the large language model in generating answers in the domain knowledge question-answering task. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is the overall flowchart of an intelligent question-answering method based on multi-hop similar subgraphs and large language models of the present invention;

[0060] Figure 2 It is the working flowchart between modules of an intelligent question-answering method based on multi-hop similar subgraphs and large language models of the present invention;

[0061] Figure 3 It is the process example diagram of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] Combined with Figure 1 andFigure 2 , the present invention proposes an intelligent question - answering method based on multi - hop similar sub - graphs and large language models, and the specific steps are as follows:

[0064] Step S1: Use a pre - trained model to extract the features of questions, knowledge graph entities, and knowledge graph relationships respectively;

[0065] In the present invention, a pre - trained model is used to extract the feature vectors of questions, knowledge graph entities, and knowledge graph relationships respectively, as follows:

[0066] d = Embedding(D)

[0067] r i = Embedding(R i ), i ∈ [1, 2, …, l]

[0068] e j = Embedding(E j ), j ∈ [1, 2, …, n]

[0069] Where:

[0070] d is the feature vector of the question D;

[0071] r j is the feature vector of the i - th relationship R i in the knowledge graph relationship set R, and l is the number of relationships in the knowledge graph relationship set R;

[0072] e j is the feature vector of the j - th entity E j in the knowledge graph entity set E, and n is the number of entities in the knowledge graph entity set E;

[0073] After extracting the feature vectors of questions, knowledge graph entities, and knowledge graph relationships respectively, by calculating the similarity between the feature vector d and e j When sim(d, e j ) > γ, then the feature vector e j is included in the starting point set E s , and E s is used to store all the starting point entity feature vectors. γ is a hyper - parameter for controlling the similarity threshold. The present invention can use the cosine similarity or Manhattan distance method to calculate the similarity between the question feature vector and the knowledge graph entity feature vector. The value of the hyper - parameter γ for controlling the similarity threshold can be adjusted according to actual needs. When using the cosine similarity calculation, the hyper - parameter γ of the similarity threshold can be default set to 0.5.

[0074] After determining the starting point set E s through the above steps, for the starting point set Es For each starting entity feature vector \(e\), the attention mechanism is used to aggregate the attention information of the starting entity feature vector \(e\) to the feature vectors of its neighbor nodes, as well as the attention information of the query feature vector \(d\) to the feature vectors of the neighbor nodes of the starting entity feature vector \(e\); respectively as follows:

[0075]

[0076] In the formula,

[0077] ⊙ is the Hadamard product;

[0078] \(\eta\) d is the propagation information value of the query feature vector \(d\) to the feature vectors of the neighbor nodes of the starting entity feature vector \(e\);

[0079] \(\eta\) e is the propagation information value of the entity feature vector \(e\) to the feature vectors of its neighbor nodes;

[0080] is the neighbor propagation vector;

[0081] is the feature vector of the \(i\)-th neighbor entity of entity \(E\);

[0082] \(Neb(E)\) is the set of feature vectors of the neighbor entities of entity \(E\);

[0083] is the feature vector of the relationship between entity \(E\) and its neighbor entity ;

[0084] \(\delta\) is a hyperparameter used to allocate the proportion of the influence of query \(D\) and entity \(E\) on the propagation information;

[0085] \(e\) d is the aggregation result vector, representing the aggregation result after the query feature vector \(d\) propagates the information of the neighbor of the entity feature vector \(e\);

[0086] \(W1\) is a trainable weighting matrix used to adjust the information weighting of entity \(E\) and the propagation information value ;

[0087] \(W2\) is a trainable weighting matrix used to adjust the degree of information interaction between entity \(E\) and its neighbor nodes;

[0088] is the entity product function used to strengthen the connection between entity \(E\) and its neighbor nodes;

[0089] \(b\) is a trainable parameter used to introduce translation during the aggregation process;

[0090] \(\sigma\) is the ReLU activation function.

[0091] Step S2: Construct a multi-hop similarity subgraph based on the features extracted in Step S1 and output the answer subgraph;

[0092] Specifically, for each starting entity in the starting point set, the entities and relationships connected to the starting entity are recorded as a selection action. For each selection action, a set of statistics is stored as follows:

[0093] {P(e,a), Q(e,a), N(e,a)}

[0094] In the formula:

[0095] e represents the feature vector of the current starting entity;

[0096] a represents the action of selecting a neighbor entity, which includes the selected relationship and the feature vector of the entity;

[0097] P(e,a) is the prior score, and the large language model is used to calculate the probability of the feature vector e of the starting entity matching each action a;

[0098] Q(e,a) is the action value for evaluating the entity's selection action a, with an initial value of 0;

[0099] N(e,a) is the number of visits, indicating the number of times the current action has been visited, with an initial value of 0;

[0100] Starting from the starting entity, calculate the action score based on the prior score, action value, and number of visits of the current action, and include the action with the highest score in the subgraph.

[0101] Furthermore, the calculation formula for the action score is:

[0102]

[0103] Score(a) is the action score for the current entity to select action a;

[0104] c puct is a constant for balancing exploration nodes and exploitation nodes;

[0105] N parent (a) is the number of visits of the previous action of the current action;

[0106] Start retrieving nodes from the starting entity, select the action with the highest action score value and include it in the multi-hop similarity subgraph. If the action score values are equal, select the node with the highest prior score and include it in the multi-hop similarity subgraph.

[0107] Then, the entity value is evaluated through the rollout model and the critic model, and the entity value is backpropagated to update the visit count and action value on the path. Specifically, for the newly added entity value, the evaluation method is as follows:

[0108] The newly added entity feature vector is denoted as e new , and the query feature vector is denoted as d, represents the aggregation result vector between the query feature vector and the neighbor entity feature vectors of the newly added entity. The entity value is evaluated through the rollout model and the critic model, and the entity value is calculated according to the following formula:

[0109]

[0110] In the formula: is the evaluation result value of the newly added entity node;

[0111] λ, μ are weight parameters used to allocate the weights of the rollout model result value, the critic model result value, and the reward value;

[0112] V roll-out (e new ) is the result value of the rollout model based on the pre-trained LLM, which is used to predict the future reward of the entity;

[0113] is the result value of the critic model based on BERT, which is used to calculate whether the aggregation result vector semantically conforms to the answer of the query feature vector d;

[0114] r(e new ) is the reward value. If the entity corresponding to the feature vector e new is the answer, then r(e new ) = 1; otherwise, r(e new ) = -1.

[0115] After the evaluation is completed, the entity value of the newly added entity is backpropagated through the graph structure to update the action value and visit count of all ancestor nodes on the route. The update rule is as follows:

[0116] N(e,a) = N(e,a)’ + 1

[0117]

[0118] N(e,a)’ is the visit count of the previous state;

[0119] N(e,a) is the current visit count;

[0120] Q(e,a) is the action value for evaluating the entity to select action a, and e is the current entity feature vector;

[0121] is an indicator function that takes 1 when the entity corresponding to the feature vector e transfers to the entity corresponding to the feature vector e in the state of selecting action a, and takes 0 if the transfer fails; new corresponding entity, and 0 if the transfer fails;

[0122] represents the value evaluation result of e in the j-th simulation, where j is the simulation round. new corresponding entity, and 0 if the transfer fails;

[0123] After the backpropagation is completed and updated, repeat the operation of step S2 for the next starting entity, and calculate the updated action value Q and access count N to construct a multi-hop similar subgraph. The end conditions for the construction are:

[0124] (1) Find the answer node;

[0125] (2) Reach the set number of entity retrievals, which needs to be determined based on the experimental results, and the initial value can be set to 10;

[0126] (3) There are no actions for the current entity to expand.

[0127] When any one or more of the above (1) to (3) occur during the construction process, the construction is completed.

[0128] Finally, select the subgraph with the highest value (i.e., the highest calculated action value) from all subgraphs as the answer subgraph.

[0129] Step S3: Split the low-rank matrix according to the step size and fine-tune the large language model, which specifically includes the following steps:

[0130] Step S3.1: Perform low-rank decomposition on the weight matrix of the existing large language model to reduce the number of training parameters:

[0131] ΔW = BA

[0132] In the above formula,

[0133] ΔW is a full-rank weight matrix of d×k;

[0134] Matrix B is d×r, and matrix A is a low-rank matrix of r×k, where r << min(d,k);

[0135] Step S3.2: Further decompose matrix B and matrix A, and introduce the Hadamard product to calculate the final weight matrix:

[0136]

[0137] where step is the splitting step size, and matrix B iThe \(i\)-th column to the \((i + step - 1)\)-th column of matrix \(B\) are taken as a sub-matrix, matrix \(A\) j The \(j\)-th row to the \((j + step - 1)\)-th row of matrix \(A\) are taken as a sub-matrix.

[0138] Step S3.3: Use the segmented matrix \(B\) x and matrix \(A\) x for matrix multiplication to obtain matrix \(C\) x , and perform Hadamard product on the result to obtain the final weight matrix. The formula is as follows:

[0139] \(C\) x = \(B\) x · \(A\) x

[0140]

[0141] In the above formula: step is the segmentation step length, \(C\) x is a \(d×k\) matrix, and \(h\) is the forward propagation formula.

[0142] Step S4: Input the question and the answer sub-graph output in Step S2 into the large language model processed in Step S3 for processing to generate an answer.

[0143] Example 1

[0144] Combined with Figure 3 , assuming that there is an existing question \(D =\) "How to handle the situation where the fixing screw is not installed?", and there is a domain knowledge graph \(G\), where the entity set is \(E\) and the relationship set is \(R\). The intelligent question-answering method based on multi-hop similar sub-graphs and large language models provided in this example includes the following steps:

[0145] Step 1: Use a pre-trained model to extract features from the question, knowledge graph entities, and knowledge graph relationships. Assume there are entities \(E1 =\) "fixing bolt", \(E2 =\) "flange bolt", and relationship \(R1 =\) "fault";

[0146] \(d = Embedding(D)=[0.35, 0.67, -0.28, 0.45, \cdots]\)

[0147] \(e1 = Embedding(E1)=[0.12, 0.98, -0.45, 0.33, \cdots]\)

[0148] \(e2 = Embedding(E2)=[0.55, 0.44, 0.68, -0.23, \cdots]\)

[0149] \(r1 = Embedding(R1)=[0.27, -0.13, 0.44, 0.55, \cdots]\)

[0150] Calculate the similarity between the query and the entities in the knowledge graph, and select the entities with similarity greater than the threshold as the starting entities. Here, cosine similarity is used as the similarity calculation method, and the similarity threshold γ = 0.80 is set.

[0151]

[0152] Therefore, include e1 in the starting set E. s , and exclude e2.

[0153] Furthermore, for the feature vector e1 of the starting entity, aggregate the attention information of the starting entity and its neighbor nodes, as well as the attention information of the query and the neighbor nodes of the starting entity. Let δ = 0.5 and b = 0, and calculate using the formula of the present invention to obtain e d = [0.72, 0.1, 0.65, 0.24, …].

[0154] Step 2: Construction of multi-hop similar subgraphs;

[0155] First, starting from the starting entity, expand outward, and combine the historical statistics and prior probabilities of each step to select the relationships and entities to be included in the subgraph. Starting from the starting element e1, retrieve nodes and select entities outward to be included in the multi-hop similar subgraph. For the starting element e1, there are three neighbor nodes, so the initial statistics for the three actions are:

[0156] {P(e1, a1) = 0.7, Q(e1, a1) = 0, N(e1, a1) = 0}

[0157] {P(e1, a2) = 0.2, Q(e1, a2) = 0, N(e1, a2) = 0}

[0158] {P(e1, a3) = 0.1, Q(e1, a3) = 0, N(e1, a3) = 0}

[0159] Calculate the action scores of the three actions of e1 in the first round. Let c puct = 1.5, and the calculation result is:

[0160]

[0161] Furthermore, add the selected action a1 to the similar subgraph, and conduct value evaluation through the rolling model and the key model. Let λ = 0.4 and μ = 0.2;

[0162]

[0163] Finally, after the evaluation is completed, for the entities newly added with similar subgraphs, the entity values are backpropagated through the graph structure to update the action values and visit counts of all ancestor nodes on the route, so that the information obtained by the newly added subgraph entities during the evaluation affects the remaining nodes in the graph;

[0164] N(e1,a1) = N(e1,a1)' + 1 = 1

[0165]

[0166] Step 3: Split the low-rank matrix according to the step size and fine-tune the large language model;

[0167] First, disassemble the weight matrix of the large language model into a low-rank matrix and fine-tune the large language model, with step = 10;

[0168] ΔW = BA

[0169]

[0170] Furthermore, combine the question and the answer subgraph to form a prompt and input it into the large language model to generate an answer.

[0171] Suppose a similar subgraph g is obtained, which contains triples: [{Fixing bolts of the pump housing, Fault, Not installed}, {Fixing bolts of the pump housing, Located in, Hydrogen peroxide tank area}, {Measures for unfixed bolts, Install fixing bolts}...]. Then combine the question D and the similar subgraph g to form a prompt prompt, and input prompt into the fine-tuned large language model to generate an answer, as shown in the following table:

[0172]

[0173] Answer = LLM(prompt)

[0174] Answer is the generated answer. The output result is "The fixing bolts of the pump housing should be installed immediately, especially in the hydrogen peroxide tank area, to ensure that the bolts are firmly installed to prevent failures", and the Q&A is completed.

[0175] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent question-answering method based on multi-hop similar subgraphs and large language models, characterized in that: It includes the following steps: S1. Use a pre-trained model to extract the features of the question, knowledge graph entities, and knowledge graph relationships respectively; S2. Construct a multi-hop similarity subgraph based on the features extracted in step S1, and output an answer subgraph; S3. Split the low-rank matrix according to the step size and fine-tune the large language model; S4. Input the question and the answer subgraph output in step S2 into the large language model processed in step S3 for processing to generate an answer.

2. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 1, wherein In step S1, use a pre-trained model to extract the feature vectors of the question, knowledge graph entities, and knowledge graph relationships respectively.

3. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 2, wherein In step S1, by calculating the similarity between the question feature vector and the knowledge graph entity feature vector, and combining the set threshold, a starting point set is determined. The starting point set is used to store all starting entities. For each starting entity in the starting point set, the attention mechanism is used to aggregate the attention information between the starting entity and its neighbor nodes, and the attention information between the question and the neighbor nodes of the starting entity.

4. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 3, wherein The cosine similarity or Manhattan distance method is used to calculate the similarity between the question and the knowledge graph entity. When the similarity is greater than the set threshold, the starting entity is included in the starting point set.

5. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 3 or 4, characterized in that In step S2, first, starting from the starting entity, calculate the action score according to the prior score, action value, and access times of the current action, and include the action with the largest score in the subgraph; then, evaluate the entity value through the rollout model and the critic model, and backpropagate the entity value to update the access times and action values on the path; finally, select the subgraph with the highest action value from all subgraphs as the answer subgraph.

6. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 5, wherein In step S2, for each starting entity in the starting point set, the entities and relationships connected to the starting entity are recorded as a selection action. For each selection action, a set of statistics is stored as follows: {P(e,a),Q(e,a),N(e,a)} In the formula: e represents the feature vector of the current starting entity; a represents the action of selecting a neighbor entity, which includes the selected relationship and the feature vector of the entity; P(e,a) is the prior score, and the large language model is used to calculate the probability of the feature vector e of the starting entity matching each action a; Q(e,a) is the action value for evaluating the entity's selection action a, and the initial value is 0; N(e,a) is the access times, indicating the number of times the current action is accessed, and the initial value is 0; The calculation formula for the action score is as follows: Score(a) is the action score of the current entity's selection action a; c puct A constant for balancing exploration nodes and exploitation nodes; N parent (a) The number of accesses to the previous action of the current action; Retrieve nodes starting from the starting entity, select the action with the largest action score value and include it in the multi-hop similarity subgraph. If the scores of the action scores are equal, select the node with the largest prior score and include it in the multi-hop similarity subgraph.

7. The intelligent question answering method based on multi-hop similar subgraphs and large language models according to claim 6, wherein, For the evaluation of the entity value of the newly added entity, the evaluation method is as follows: For the entity feature vector newly added to the similar subgraph, it is denoted as e new , the question feature vector is denoted as d, represents the aggregation result vector between the question feature vector and the neighbor entity feature vectors of the newly added entity. Calculate the entity value according to the following formula: Wherein: is the value evaluation result value of the newly added entity node; λ and μ are weight parameters used to allocate the weights of the rollout model result value, critic model result value, and reward value; V roll-out (e new ) is the result value of the rolling model based on the pre-trained LLM and is used to predict the future reward of the entity; The result value of the key model based on BERT, which is used to calculate the aggregated result vector The answer that semantically conforms to the question feature vector d; r(e new ) is the reward value. If the entity corresponding to the feature vector e new is the answer, then r(e new ) = 1; otherwise, r(e new ) = -1.

8. The intelligent question answering method based on multi-hop similar subgraphs and large language models according to claim 6, wherein, The entity value of the newly added entity node is backpropagated through the graph structure to update the action values and access times of all ancestor nodes on the route. The update rule is as follows: N(e,a)=N(e,a)’+1 N(e,a)’ is the access times of the previous state; N(e,a) is the current access count; Q(e,a) is the action value for evaluating entity e to select action a, where e is the current entity feature vector; is an indicator function that takes the value 1 when the entity corresponding to the feature vector e transfers to the entity corresponding to the feature vector e new in the state of action a, and takes the value 0 otherwise; Indicates the value evaluation result value of e in the j-th simulation, where j is the simulation round. new ​ 9. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 6, wherein After the backpropagation update is completed, repeat the operation of step S2 for the next starting entity. The end conditions for constructing the multi-hop similarity subgraph are as follows: (1) Find the answer node; and / or (2) Reach the set number of entity retrieval times; and / or (3) There is no action for the current entity to expand.

10. The intelligent question-answering method based on multi-hop similar subgraphs and large language models according to claim 1, wherein Step S3 includes the following steps: S3.

1. Perform low-rank decomposition on the weight matrix of the large language model to obtain two low-rank matrices; S3.

2. Further decompose the two low-rank matrices respectively according to the slice step length, and then introduce the Hadamard product to calculate the final weight matrix.

Citation Information

Cited By

  • Data processing method and system based on large table model

    CN120723900A