Retrieval enhancement generation method and system based on precise subgraph matching and medium

By using precise subgraph matching technology, the problem of inaccurate answers in existing search enhancement generation is solved, and the topological logic consistency between search results and user queries is achieved, ensuring the accuracy and traceability of generated content.

CN120892575APending Publication Date: 2025-11-04CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510883002.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-28
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing search enhancement generation techniques rely on fuzzy matching during the matching process, which cannot guarantee the accuracy of the answer and ignores the topological dependencies of the query graph, resulting in inaccurate generated answers.

Method used

We employ precise subgraph matching technology, generating node and path embedding vectors for the data graph through offline preprocessing, constructing a hierarchical R-Tree index, and performing path-level dominance embedding and hierarchical index pruning during the online retrieval stage to ensure the topological isomorphism between the subgraph and the query graph, thereby achieving precise matching.

Benefits of technology

It ensures that the search results are completely consistent with the topological logic of the user query, eradicating structural illusions, supporting end-to-end logical chain verification across multiple paths, and ensuring that the generated content can be traced back to specific paths in the knowledge graph, eliminating fictitious relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892575A_ABST
    Figure CN120892575A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement generation method and system based on accurate sub-graph matching and a medium, and can solve the problem of inaccurate answer caused by the situation that complete dependence of a graph structure asked by a user cannot be obtained, a retrieval method and the like in related technologies. The method comprises the following steps: in an off-line preprocessing stage of a data graph G, constructing a layered R-Tree index for all paths with the length of 1-d in the data graph G; in the online retrieval stage, an accurate matching sub-graph set S and an approximate matching sub-graph set S'are obtained through an accurate sub-graph matching algorithm; wherein the precise matching sub-graph set S is a set of precise matching sub-graphs g of the data graph G, and the precise matching sub-graphs g and the normalized query graph q are isomorphic; the approximate matching sub-graph set S'is a set of approximate matching sub-graphs g 'of the data graph G, and the approximate matching sub-graphs g' are approximately matched with the normalized query graph q; and generating a credible answer based on the accurate matching sub-graph set S and the approximate matching sub-graph set S '.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of large language models, and particularly relates to a retrieval enhancement generation method and system based on accurate subgraph matching and a medium. BACKGROUND

[0002] With the breakthrough development of large language models (LLM) such as GPT-4 (Generative Pre-trained Transformer 4) and LLaMA (Large Language Model MetaAI), the paradigm of natural language processing (NLP) has been transformed in today's society. LLMs can model long-range dependencies through multi-head attention, breaking through the sequence processing bottleneck of traditional RNN (Recurrent Neural Network) and CNN (Convolutional Neural Network), and have strong context awareness, which can dynamically allocate weights based on positional encoding to accurately capture word order and semantic associations. Despite the impressive potential of LLMs, their underlying technical mechanisms still have inherent flaws, such as uncontrollable hallucination, which refers to the generation of irrelevant or factually incorrect content. This is because errors in the training data are amplified by the probabilistic model, or because of the local optimal trap caused by greedy decoding or nucleus sampling. To address the inherent flaws of hallucination generation in large language models (LLMs), retrieval-augmented generation (RAG) technology has emerged. The essence of RAG is to dynamically retrieve external knowledge and constrain the generation process, integrating the general language ability of LLMs with domain-specific knowledge, and establishing a technology paradigm of "real-time knowledge update + controllable generation logic". RAG mainly has three stages to enhance the ability of LLMs: in the retrieval stage, relevant information fragments are retrieved from external knowledge bases (such as databases and knowledge graphs) in real time; in the fusion stage, the retrieval results are converted into structured context (such as text fragments, subgraphs, and tables) and input into the prompt (Prompt); in the generation stage, LLM generates the final answer based on the retrieval content and its own parameter knowledge, ensuring the consistency of the output with the external knowledge.Although RAG technology significantly improves the generation reliability, its dependent retrieval mechanism still has deficiencies. The existing RAG retrieval mechanism is based on vector similarity and fuzzy matching of top-k results. The text is converted into a high-dimensional vector by an embedding model to capture semantic features, then the closeness between vectors is measured using cosine similarity, Euclidean distance, etc., and finally the top K results with the highest similarity are returned as the context of the generated answer. However, this method is obviously a fuzzy matching, which cannot be proven to be an exact matching in theory, and cannot guarantee the accuracy of the answer.

[0003] The existing representative technologies include two types of implementation schemes:

[0004] GraphRAG (Graph-based Retrieval-Augmented Generation) is a graph-enhanced retrieval paradigm, which is the representative of the second generation of RAG technology. Its core improvement is to achieve global semantic understanding through the hierarchical structure of the knowledge graph. The process design of GraphRAG is divided into two stages: indexing time and query time. In the Indexing Time stage, there are ① text extraction and blocking, ② element extraction and instance abstraction, and ③ knowledge graph construction. In the Query Time stage, there are ① community summary generation, ② query-focused summarization, and ③ global answer generation. GraphRAG converts the source text into a knowledge graph, avoiding the destruction of entity association like traditional RAG cutting text blocks. Through community summary, it covers the macro theme of the document set, solving the local perspective limitation problem of traditional RAG.

[0005] LightRAG (LIGHTRAG: SIMPLE AND FAST RETRIEVAL-AUGMENTED GENERATION) adopts a lightweight index structure and upgrades traditional RAG in multiple dimensions through the introduction of graph structure technology, significantly improving the complex query processing capability and system efficiency. LightRAG generates key-value pairs for each entity and relationship, and also adopts a two-level retrieval mechanism. The low-level retrieval focuses on specific entities and their attributes, and the high-level retrieval captures cross-entity theme associations, thereby achieving efficient retrieval. At the same time, it also avoids the destruction of entity association like traditional RAG cutting text blocks, and is much stronger than traditional RAG in multi-hop problem reasoning.

[0006] Compared with naiveRAG, GraphRAG and LightRAG have upgraded the form of cutting the dataset into text blocks into the form of generating a knowledge graph (graph), and can output the neighbor nodes of the retrieved entity nodes, so as to optimize the fragmented answers caused by the flat data representation and insufficient context awareness of the naiveRAG system, thereby greatly improving the retrieval speed.

[0007] The existing LightRAG and GraphRAG model the data graph (Data Graph) in the form of nodes or edges as the basic unit, and do not model the topological semantics of the complete path, and they only consider the data graph structure and ignore the topological dependency relationship of the query graph (query Graph), which leads to the fact that they not only rely on fuzzy matching in the matching principle, but also do not consider the complete dependency relationship in the query, and further leads to the fact that the answer generated by the RAG system in combination with the query and the context through the large language model is inaccurate. SUMMARY

[0008] The embodiment of the present application provides a retrieval enhancement generation method and system based on accurate subgraph matching, which can break through the matching accuracy bottleneck caused by the asymmetry of the "query-data" topological structure, realize accurate subgraph isomorphism matching, and avoid the problem of inaccurate answers caused by the fact that the complete dependency of the graph structure of the user's question cannot be obtained and the retrieval method.

[0009] To achieve the above object, the technical scheme is as follows:

[0010] In a first aspect, a retrieval enhancement generation method based on accurate subgraph matching is provided, comprising:

[0011] In the offline preprocessing data graph G stage, the node label embedding vector in the data graph G is generated by LLM; for the path in the data graph G with a length less than or equal to d, the path label embedding vector is generated by splicing the node label embedding vector; the node dominant embedding vector in the data graph G is generated according to the trained GNN model M; the path dominant embedding vector with a length of 1-d is generated by splicing according to the node dominant embedding vector; a hierarchical R-Tree index is constructed for all paths with a length of 1-d in the data graph G, wherein the leaf node stores the path label embedding vector and the path dominant embedding vector, and the non-leaf node stores the minimum circumscribed rectangle range; d is a preset value;

[0012] In the online retrieval stage, based on the LLM and entity normalization algorithm, the query information input by the user is processed to generate a normalized query graph q; the normalized query graph q is divided into query path set Q; based on the LLM and hierarchical R-Tree index, the candidate label embedding vector set of the unknown node in the query path set Q is obtained by using the candidate label algorithm of the unknown node; according to the query path set Q and the candidate label embedding vector set, the candidate path set P of all nodes in the query path set Q is obtained by using the Cartesian product method; based on the candidate path set P, the trained GNN model M and the hierarchical R-Tree index, the exact matching subgraph set S and the approximate matching subgraph set S' are obtained by using the exact subgraph matching algorithm; wherein the exact matching subgraph set S is a set of exact matching subgraphs g of the data graph G, and the exact matching g is isomorphic to the normalized query graph q; the approximate matching subgraph set S' is a set of approximate matching subgraphs g' of the data graph G, and the approximate matching subgraph g' is approximately matched with the normalized query graph q;

[0013] Based on the exact matching subgraph set S and the approximate matching subgraph set S', a trusted answer is generated.

[0014] In a second aspect, a retrieval enhancement generation system based on exact subgraph matching is provided, comprising:

[0015] An offline preprocessing module is configured to, in the offline preprocessing data graph G stage, generate node label embedding vectors in the data graph G by using the LLM; for the paths in the data graph G with a length less than or equal to d, the node label embedding vectors are spliced to generate path label embedding vectors; the node dominant embedding vectors in the data graph G are generated by using the trained GNN model M; based on the node dominant embedding vectors, the path dominant embedding vectors of the paths with a length of 1-d are spliced in the order of the paths; the hierarchical R-Tree index is constructed for all the paths with a length of 1-d in the data graph G, wherein the leaf nodes store the path label embedding vectors and the path dominant embedding vectors, and the non-leaf nodes store the minimum bounding rectangle range; d is a preset value;

[0016] The online retrieval module is used for an online retrieval stage, processes query information input by a user based on an LLM and an entity normalization algorithm to generate a normalized query graph q, performs path division on the normalized query graph q to obtain a query path set Q, obtains a candidate label embedding vector set of unknown nodes in the query path set Q based on an LLM and a hierarchical R-Tree index and by using a candidate label algorithm for unknown nodes, obtains a candidate path set P of all nodes in the query path set Q by means of a Cartesian product according to the query path set Q and the candidate label embedding vector set, and obtains an accurate matching subgraph set S and an approximate matching subgraph set S' based on the candidate path set P, a trained GNN model M and the hierarchical R-Tree index by means of an accurate subgraph matching algorithm; the accurate matching subgraph set S is a set of accurate matching subgraphs g of the data graph G, and the accurate matching subgraph g is isomorphic to the normalized query graph q; the approximate matching subgraph set S' is a set of approximate matching subgraphs g' of the data graph G, and the approximate matching subgraph g' is approximately matched with the normalized query graph q.

[0017] The answer generation module is used for generating a trusted answer based on the accurate matching subgraph set S and the approximate matching subgraph set S'.

[0018] In a third aspect, an electronic device is provided, and the electronic device includes:

[0019] one or more processors;

[0020] a memory device configured to store one or more programs;

[0021] When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to the first aspect.

[0022] In a fourth aspect, a computer program product is provided, and the computer program product includes a computer program or instructions, which, when executed on a computer, cause the computer to perform the method according to the first aspect.

[0023] The present application has the following beneficial effects:

[0024] The ESM-RAG (Exact Subgraph Matching Retrieval-Augmented Generation) of the application realizes a significant technical breakthrough and application value by deeply integrating the exact subgraph matching technology into the RAG framework, ensures that the retrieval result is completely consistent with the topological logic of the user query through subgraph isomorphism verification (dual constraints of label consistency and topological dominance), thereby radically curing structural hallucinations, supporting end-to-end logical chain verification across multiple paths, solving the fragmentation problem of traditional RAG, supporting natural language queries containing unknown node / edge labels through zero embedding zeroing and Cartesian product generation, and forcing the LLM to output based on the exact matching subgraph in the answer generation stage, eliminating fabricated relationships, and generating content that can be traced back to specific paths in the knowledge graph. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 A step diagram of a retrieval augmented generation method based on exact subgraph matching is schematically shown.

[0026] Figure 2 A block diagram of a retrieval augmented generation system based on exact subgraph matching is schematically shown.

[0027] Figure 3 A block diagram of an electronic device is schematically shown.

[0028] Figure 4 A block diagram of a computer readable medium is schematically shown. DETAILED DESCRIPTION

[0029] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example embodiments to those skilled in the art. Like reference numerals refer to like elements throughout the figures, and descriptions of the same or similar elements can be omitted or simplified in some instances.

[0030] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the application can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the application.

[0031] The block diagrams shown in the drawings are merely functional entities and do not necessarily have to correspond to physically independent entities. That is, these functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the drawings are merely exemplary illustrations and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be further decomposed, and some operations / steps can be combined or partially combined, so the actual execution order can be changed according to actual conditions.

[0033] It should be understood that although the terms first, second, third, etc. can be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another component. Therefore, the first component discussed below can be referred to as the second component without departing from the teachings of the present application concepts. As used herein, the term "and / or" includes all combinations of any one and one or more of the associated listed items.

[0034] Those skilled in the art can understand that the modules or flows in the drawings are not necessarily required for implementing the present application, and therefore cannot be used to limit the protection scope of the present application.

[0035] The primary purpose of the present application is to break through the matching accuracy bottleneck caused by the asymmetry of the "query-data" topology structure in the prior art, to realize accurate subgraph isomorphism matching, and to avoid the problem of inaccurate answers caused by the inability to obtain the complete dependency of the graph structure of the user's query.

[0036] To solve the above problems, the present application proposes an exact subgraph matching enhancement generation (ESM-RAG) framework, which ensures the topological isomorphism of the returned subgraph and the query graph in the retrieval stage. It considers the complete dependency relationship of the query graph, obtains the accurately matched subgraph through path-level domination embedding and hierarchical index pruning, realizes the accurate matching of the graph structure, and combines it with RAG (Retrieval-augmented Generation) to guide the content generation of the LLM (Large Language Model) with the most relevant information to query, effectively alleviating the illusion problem of the LLM (Large Language Model). The technical scheme is divided into two stages of offline preprocessing and online retrieval.

[0037] The symbols and descriptions involved in the embodiments of the present application are shown in Table 1:

[0038]

[0039]

[0040] Table 1: Symbols and descriptions

[0041] The formal definition in the present application is as follows:

[0042] Definition 1 (graph, G), a data graph G is represented by a four-tuple (V(G), E(G), φ(G), L(G)), wherein: · V(G) is a node set, containing vertices v i ;

[0043] · E(G) is an edge set, containing edges e ij = (v i ,v j ) represents the connection between nodes v i and v j ;

[0044] · φ(G) is a mapping function φ: V(G) × V(G) → E(G), used to define the mapping of node pairs to edges; · L(G) is a label function, associated with a label (v i ) for each node v i ∈ V(G).

[0045] Definition 2 (node, v), a node v in a data graph G is represented by a three-tuple (id v , l v , d v ), wherein: · id v is the unique identifier of the node;

[0046] · l v is the label of the node;

[0047] · d v is the description information of the node.

[0048] Definition 3 (edge, e), an edge e in a data graph G is represented by a four-tuple (id e , v i , v j , d e ), wherein:

[0049] · id e is the unique identifier of the edge;

[0050] · v i , v j ∈ V(G) are the source node and target node connected by the edge;

[0051] ·l e It is the label of the edge;

[0052] ·d e It is the descriptive information of the edge.

[0053] Definition 4 (Graph Isomorphism): Given two graphs G A =(V A E A ,φ A ,L A )and G B =(V B E B ,φ B ,L B If there exists a bijective function f:V that preserves the edge structure, A →V B , so that:

[0054] For all v i ∈V A , satisfying L A (v i ) = L B (f(v i ));

[0055] For all v i ,v j ∈V A If (v i ,v j )∈E A Then (f(v) i ),f(v j ))∈E B ;

[0056] Then it is called G A With G B Isomorphism (denoted as G) A ≡G B ).

[0057] Furthermore, if G A isomorphic to G B An derived subgraph g B Then G is called A It is G B Subgraph isomorphism.

[0058] Definition 5 (Subgraph Matching Query): Given a data graph G and a query graph q, the goal of a subgraph matching query is to retrieve all subgraphs g in G that are isomorphic to q (i.e., g ≡ q).

[0059] Definition 6 (Graph Edit Distance), given two graphs G1 = (V1, E1, φ1, L1) and G2 = (V2, E2, φ2, L2), the graph edit distance ged(G1, G2) is defined as the total cost of the minimum sequence of allowed edit operations required to transform G1 into G2. The allowed edit operations include:

[0060] • Insertion or deletion of a node;

[0061] • Insertion or deletion of an edge;

[0062] • Vertex replacement (label replacement)

[0063] According to the first specific embodiment of the present application, as shown in Figure 1 , the present application provides a retrieval enhancement generation method based on accurate subgraph matching, comprising the following steps:

[0064] (1) Offline pre-processing data graph G stage:

[0065] The offline stage is the preliminary preparation of the original data, which performs structured processing on the original data, i.e. extracts entities and relationships in the original data through LLM, and embeds entity labels to obtain node label embeddings, constructs an efficient pre-computed data structure for the online query stage, and obtains the dominant embedding of the node by training the GNN (Graph Neural Network) model M, and the label embedding and the dominant embedding of the path can be obtained by combining them. Label embedding only depends on the label (Label) of the node in the path and does not involve the structural topology information of the graph, and the dominant embedding involves the topology of the path and reflects the dominance of the path. The introduction of label embedding and dominant embedding speeds up the subgraph matching process and lays the performance foundation for subgraph matching query. This stage mainly includes the following steps:

[0066] S11, generating node label embedding vectors in data graph G through LLM.

[0067] In step S11, LLM is used to generate label embedding vectors o0(v i ) for all nodes v i in data graph G.

[0068] S12, for the paths in data graph G with length less than or equal to d, the node label embedding vectors are spliced to generate path label embedding vectors.

[0069] For the paths in data graph G with length less than or equal to d, the path label embedding vectors o0(p z ) are obtained by splicing the node label embeddings.

[0070] S13, Generate node-dominant embedding vectors in the data graph G based on the trained GNN model M.

[0071] First, train the GNN model M. Then, obtain the embedding vector for each node in graph G from M. For training the GNN model M: first obtain the dataset D, which is initially an empty set, i.e. For a vertex set V(G), for each node v in it i Perform the operation to extract v i Unit star diagram A unit star subgraph refers to a subgraph containing nodes v i The subgraph consisting of the unit star subgraph and all its 1-hop neighbors; then extract the unit star subgraph. unit star substructure in Next, the subgraph pair Add dataset D. Randomly shuffle the pairs in dataset D to avoid order bias during training. Input dataset D and learning rate η, divide dataset D into multiple batches B. For each batch B, compute the embedding vector of the subgraph pair using the current GNN (Graph Neural Network) model M; then apply the loss function formula... Calculate the loss function, where It is the central node v i The unit star subgraph, containing v i The local structures of its direct neighbors; Is with Related substructures; and These are their embedding vectors, and the cumulative loss L during the training process of the GNN model M to be trained is... e When the value is 0, the trained GNN model M is output. After obtaining M, M is used to obtain the graph G obtained in Module 1, thereby obtaining the label embedding and dominant embedding of each node in graph G. This is the basis for supporting efficient pruning. The GNN model training algorithm is shown in Algorithm 1.

[0072]

[0073] Algorithm 1: GNN Model Training Algorithm

[0074] S14: Based on the node-dominant embedding vector, concatenate the paths in order to generate a path-dominant embedding vector of length 1-d.

[0075] Applying the trained GNN model M to graph G yields the dominant embedding o(p) of the path of length 1-d originating from each node in G. zThe dominant embedding of a path is obtained by concatenating the dominant embeddings of each node in the path in the order they appear in the path.

[0076] S15 constructs a hierarchical R-Tree index for all paths of length 1-d in the data graph G, where leaf nodes store path label embedding vectors and path dominant embedding vectors, and non-leaf nodes store the minimum bounding rectangle range.

[0077] The path set of length 1-d obtained in step S14 is used to construct a hierarchical R-Tree index I. L The leaf nodes store the dominant embedding and label embedding of the path, while the non-leaf nodes store the range of the minimum bounding rectangle. During retrieval, efficient pruning is achieved by determining whether to continue the search by judging the numerical range of the non-leaf nodes in the hierarchical R-Tree, thus eliminating irrelevant branches and reducing time complexity.

[0078] (II) Online Search Stage:

[0079] The Online phase is the core process of the ESM-RAG (Exact Subgraph Matching Retrieval-Augmented Generation) framework in responding to user queries. Through dynamic query parsing, path-level exact matching, and verifiable answer generation, it achieves end-to-end processing from natural language input to reliable output, mainly including the following steps:

[0080] S21, based on LLM and entity normalization algorithm, processes the query information input by the user to generate a normalized query graph q.

[0081] The query information input by the user is a natural language question posed by the user. The Large Language Model (LLM) extracts entities from the user's query information, identifies key entities and relational constraints, and combines the resulting entities and relations into a query graph q. * And generate its node label embedding o0(q) i Then, entity normalization is performed, the goal of which is to align the entity names in the user query with the standard representations in the knowledge graph (G), eliminating representational differences (such as aliases, abbreviations, and spelling errors). This is achieved using an entity normalization algorithm, whose input is the query graph q containing the original entities. * First, copy q * We obtain the node set V(q), and then traverse each vertex in the node set. Generate label embeddings to obtain vertices Label embedding vector Then use FAISS_Search to search for matches in the knowledge base graph G. Most similar get entity q i Finally, replace q in V(q) with q i The normalized query graph q is obtained. Wherein, FAISS_Search can achieve millisecond response in the order of billions of data through Approximate Nearest Neighbor Search (ANN) and vector compression technology, reducing the time complexity. This step is the basis for ESM-RAG to achieve accurate retrieval, avoiding path missed detection or false detection caused by name difference, laying a core data guarantee for subsequent accurate subgraph matching. The entity normalization algorithm is shown in Algorithm 2.

[0082]

[0083] Algorithm 2: Entity normalization algorithm

[0084] S22, divide the normalized query graph q into a query path set Q.

[0085] The path division algorithm aims to generate a path set that covers all edges of the query graph and has the lowest total cost. The cost model is used to quantify the computational overhead of different division schemes, guiding the algorithm to select the optimal solution. The process of the path division algorithm is as follows: input the query graph q and the path length L, first initialize the optimal query path set, that is Initialize the current optimal cost, that is, CostQ(φ) = +∞, indicating that there is no feasible solution in the initial state. Then select the node q i with the highest degree from the query graph q as the starting point, because a high-degree node will cover more edges, reducing the computational overhead of subsequent path expansion. Then generate an initial path set P containing the starting node q i and the length L. The algorithm details are as follows: first, initialize the current path set localQ = {p q} and the current cost local_cost = 0, and expand the path set through the greedy strategy, select a path p connected with the current localQ, which satisfies two conditions: ① minimum overlap, that is, the new path p has the least overlap with the existing path. ② minimum weight, that is, the computational cost of the path p is the lowest, and the budget cost is determined by the pre-defined cost model. Then add p to localQ, and calculate the cumulative path cost local_cost ← local_cost + w(p). If the cost local_cost is better than the global optimal cost CostQ(φ) at this time, update Q and CostQ

[0086] ​(φ), continuing until all edges of the query graph q are covered, finally returning the set of paths Q that covers all edges of the query graph and has the lowest total cost. This algorithm uses a greedy strategy to select the path with the minimum overlap and the minimum weight, covering all associated edges and avoiding redundant calculations. The path partitioning algorithm can obtain the set of query paths Q obtained from the query graph q. At this time, the query paths Q contain unknown nodes. The completion of unknown nodes will be explained in detail in the following steps. The path partitioning algorithm based on the cost model is shown in Algorithm 3.

[0087]

[0088] Algorithm 3: Cost-based path partitioning algorithm

[0089] S23, based on LLM and hierarchical R-Tree index, uses the candidate label embedding vector set of unknown nodes in the query path set Q to obtain the candidate label embedding vector set of unknown nodes.

[0090] First, the label embedding of each query path in Q is obtained through LLM (the label embedding of a query path is obtained by concatenating the label embeddings of the nodes in the query path in order), and the label embeddings of unknown nodes are all set to zero. Then, the candidate label embedding algorithm for querying unknown nodes is used to obtain the candidate label embedding set of unknown nodes. The algorithm principle is that if the label embeddings of all nodes except the unknown node are consistent, then the label embedding of the node at the corresponding position of that path is taken as the possible value of the label embedding of the unknown node in the query graph. The candidate label embedding algorithm for querying unknown nodes is as follows: The algorithm takes the query path set Q and the R*-Tree index I of the data graph G as input. L To output the candidate label mapping table UnknownVertexLabel for unknown vertices, the main steps are as follows: First, initialize UnknownVertexLabel to an empty mapping table to store candidate labels for unknown vertices, such as {V1:[T1,T2],V2:[E3]}. Initialize UnknownVertexPathset to an empty set to store query paths containing unknown vertices. Then, iterate through the query path set Q, filtering all paths containing unknown vertices and adding them to UnknownVertexPathset, prioritizing paths covering different unknown vertices to avoid repeatedly processing the same unknown vertex. For each query path p containing an unknown vertex... q ∈Q: The path label embedding o0(p) is generated using the GNN (Graph Neural Network) model M generated by the algorithm in the offline process. q), and then the root node of the hierarchical R-Tree index is associated to the current path and inserted into the maximum heap H, aiming to use the heap to realize the top-down traversal of the hierarchical R-Tree. Next, the nodes in the heap H are processed in a loop: the top node N and its key value key(N) are popped out, and if key(N) is less than the minimum label embedding norm of all query paths, the retrieval of the current branch is terminated, indicating that the query path is not on the branch, thereby accelerating the retrieval speed; then the node is judged, if it is a non-leaf node, the child nodes N i of the node are traversed, and whether the MBR (Minimum Bounding Rectangle) of the child nodes intersects with the label embedding range of o0(p q ), if it intersects, the child nodes are inserted into the heap H; if it is a leaf node, the paths p z in the leaf node are traversed, and whether the label embedding o0(p z ) of the paths matches the label embedding o0(p q ) of the query path p q is verified, if it matches, the label of the corresponding unknown vertex in the path p z is added to UnknownVertexLabel, and finally the candidate label embedding set UnknownVertexLabel of the unknown node is obtained, and the candidate label embedding algorithm for querying the unknown node is shown in Algorithm 4.

[0091]

[0092]

[0093] Algorithm 4: Candidate label embedding algorithm for querying unknown node

[0094] S24, according to the query path set Q and the candidate label embedding vector set, the candidate path set P of all nodes in the query path set Q is obtained by means of Cartesian product.

[0095] The UnknownVertexLabel set obtained in step S23 is corresponded to the unknown nodes, and all possible label embedding combinations of the unknown nodes are constructed by calculating the Cartesian product, so as to avoid omission, expand the original query path set containing unknown nodes into multiple query path sets with all node labels known, and finally obtain the candidate path set P. The process mainly relies on the unknown node filling algorithm, and the algorithm is described in detail as follows: the input of the algorithm is the query path set Q and the UnknownVertexLabel generated in module 3, wherein Q is a query path set containing part unknown vertex labels, and the UnknownVertexLabel is a candidate label mapping set of unknown vertices, such as {V1: [T1, T2], V2: [E3]}. First, the IDs of all unknown vertices are extracted from the UnknownVertexLabel to obtain UnknownVtertexIDs, such as [V1, V2], then the candidate label list of each unknown vertex is extracted to obtain LabelOptions, such as the candidate labels of V1 are [T1, T2] and the candidate labels of V2 are [E3], and the Cartesian product of all candidate labels is calculated to generate all possible label combinations LabelCombinations, such as (T1, E3) and (T2, E3). Next, the label mapping and path filling are created: first, P is initialized as an empty set, then each label combination combo in LabelCombinations is traversed: a LabelMap is created, the unknown vertex ID is associated with the corresponding candidate label, such as [V1→T1, V2→E3], the original query path set Q is copied to generate Q_copy, so as to avoid directly modifying the original data; then each path and each node in Q_copy is traversed: if the node label is empty and its ID is in the LabelMap, the candidate label is filled, finally the filled Q_copy is added to P. The algorithm returns the set P of all possible completely labeled query path sets, which is very important in the subsequent main process, and is mainly used for generating all possible candidate query path sets. The output of the algorithm is also used as an input module in the subsequent precise matching, which is a core submodule in the precise matching. By dynamically generating all possible candidate path label combinations, the problem that the traditional method does not consider the query containing unknown label nodes is solved, and the reliability of complex graph data query is significantly improved. The unknown node filling algorithm is shown in algorithm 5.

[0096]

[0097]

[0098] Algorithm 5: Unknown node filling algorithm

[0099] S25. Based on the candidate path set P, the trained GNN model M, and the hierarchical R-Tree index, the exact matching subgraph set S and the approximate matching subgraph set S' are obtained through the exact subgraph matching algorithm.

[0100] Wherein, the exact matching subgraph set S is the set of exact matching subgraphs g of the data graph G, and the exact matching g is isomorphic to the normalized query graph q; the approximate matching subgraph set S' is the set of approximate matching subgraphs g' of the data graph G, and the approximate matching subgraph g' is approximately matched to the normalized query graph q.

[0101] Specifically, the dominant embeddings of paths in the query path set P are first obtained using the GNN model M trained in the offline phase, for subsequent precise subgraph matching. Then, S and S' are obtained using the precise subgraph matching algorithm. The detailed steps of the algorithm are as follows: The input to the algorithm is the output P of Algorithm 5, i.e., the fully labeled query path set, the GNN (Graph Neural Network) model M trained in the offline phase, and the hierarchical R-Tree index I of the data graph G. L First, initialize S and S' as empty sets; then, iterate through each query path set Q in the query path set P. ' , for Q ' Each query path p in q Perform the following operations: Initialize p q The candidate list `cand_list` and approx_list are used to generate two types of path embeddings through the GNN model M: ① label embedding based solely on vertex labels ② dominant embedding based on topological structure. Then comes the index traversal and heap processing stage. This stage first associates the root node of the index with the current query path set Q, inserts it into a max-heap H, and iteratively processes the nodes in heap H: popping the top node N and its key value `key(N)`. If `key(N)` is less than the minimum dominant embedding norm of all query paths, the current branch is terminated, indicating that the path does not exist with this branch, and pruning is performed directly. For nodes popped from the heap, there are two cases: if it is a non-leaf node, traverse its child nodes N. i Check if its MBR (Minimum Bounding Rectangle) matches the dominating region DR (o(p) of the query path. q )) and DR(o'(p q If the nodes intersect and the condition is met, insert the child node into heap H; if it is a leaf node, traverse the path p in the leaf node. z First, verify the tag embedding match (o0(p)). q )==o0(p z Then verify the dominance relationship (o(p)). q)≤o(p z And o'(p) q )≤o'(p z If satisfied, then p z Add the candidate to either `cand_list` (exact candidate) or `approx_list` (approximate candidate). Candidates satisfying both label embedding and dominance relation are added to `cand_list`, while those satisfying only label embedding are added to `approx_list`. Finally, for p... q ∈Q, connect p q For p, find all candidate paths in cand_list and obtain the exact matching subgraph g in S; q ∈Q, connect p q Get S from all candidate paths in approx_list. ' Approximate matching subgraph g in ' Then calculate the graph edit distance for all subgraphs, and collect the subgraphs whose edit distance is less than or equal to x into set S. ' Algorithm 6 is the core implementation of the candidate path retrieval and verification stage, supporting the final output of the overall process. It strictly filters candidate paths through the partial order relation of the dominant embedding to ensure the mathematical constraint of subgraph isomorphism. The precise subgraph matching algorithm is shown in Algorithm 6.

[0102]

[0103]

[0104] Algorithm 6: Precise Subgraph Matching Algorithm

[0105] Algorithm 6 concludes with a subgraph stitching operation, primarily relying on a subgraph assembly algorithm that uses a two-stage assembly process. The algorithm's inputs are the query graph q, the edit distance threshold x, and the query path set Q. ' The first assembly stage of the algorithm is for each path p. q The associated exact matching path set generates a Cartesian product C, where for each path combination p in the Cartesian product... c First, i will perform initialization operations: set graph g to an empty graph, set the node mapping table query_node_map to empty, and set conflict to False; then, for each path p in the path... c The operation involves two steps. The first step is to query each node v in the path. c The operation first places the node in p c The data nodes mapped in the middle are given to v d Then make a judgment, if v c If it already exists in the mapping table, then check the original mapping v. dIf it is equal to the new mapping, set conflict to True and exit the inner loop; if v is not equal, then... c If it does not exist in the mapping table, add the mapping query_node_map[v c ]←v d If v d If v is not in graph g, then v d Add; the second step is to add p c All edges in S are added to g. Finally, it is checked whether there is a conflict. If not, g is merged into S. The second assembly stage is for each path p. q The associated approximate matching path set generates a Cartesian product C. ' For each path combination p in the Cartesian product ' c First, i will perform initialization operations: [image g] ' Set the graph to empty, empty the node mapping table query_node_map, and set conflict to False; then for each path p in the path... ' c The operation involves two steps. The first step is to query each node v in the path. c The operation first places the node in p ' c The data nodes mapped in the middle are given to v d Then make a judgment, if v c If it already exists in the mapping table, then check the original mapping v. d If it is equal to the new mapping, set conflict to True and exit the inner loop; if v is not equal, then... c If it does not exist in the mapping table, add the mapping query_node_map[v c ]←v d If v d Not in picture g ' The middle will be v d Add; the second step is to add p ' c Add g to all edges in the middle ' Finally, it checks if there is a conflict. If not, it also checks the graph edit distance GraphEditDistance(g ' If q) is less than or equal to x, then g is set to x. ′ Merge into S ′ In the algorithm for assembling subgraphs, see Algorithm 7.

[0106]

[0107]

[0108] Algorithm 7: Assemble subgraph algorithm

[0109] The last step of Algorithm 7 uses graph edit distance for the assembly of approximate subgraph matching, which mainly relies on the graph edit distance calculation algorithm, which is used to calculate the graph edit distance, which is a measure of the difference between two graph structures, defined as the minimum number of editing operations required to convert graph g1 to graph g2. The input of the algorithm is two graphs g1 and g2, and the output is their minimum edit distance GED(g1, g2). The algorithm starts with initialization, setting the minimum edit distance to infinity, indicating that the valid solution has not been found, and then exhaustively enumerates all possible mappings f: V1→ V2∪{NULL}, and sequentially calculates the node operation cost and edge operation cost. Node cost calculation will have two-step judgment, if the node v of g1 is mapped to NULL, then the cost c v +1, if the node v is mapped to f(v) of g2 but the label is different, then the cost c v +1, for the nodes in g2 that are not mapped, each node cost is added by 1; the edge operation cost calculation will also have two-step judgment, if the endpoints of the edge (u i ,v j ) of g1 are mapped to NULL, then the cost c e +1, if the edge (u i ,u j ) does not exist in g2, then the cost c e +1, for the edges in g2 that are not covered by mapping, each edge cost is added by 1. Finally, update the minimum edit distance, the total cost c total =c v +c e , if the current total cost is the historical minimum value, then update the minimum edit distance, after the algorithm ends, return the global minimum edit distance, the specific implementation is shown in Algorithm 8:

[0110]

[0111]

[0112] Algorithm 8: Graph edit distance calculation

[0113] (Three) Answer generation stage:

[0114] S31, based on the accurate matching subgraph set S and the approximate matching subgraph set S', generate the trusted answer.

[0115] In the answer generation stage, the information of the matched subgraph is input into the large model as prompt and query to generate the answer. If the accurate matching subgraph set S is not empty, the accurate matching subgraph g is directly used as the prompt, and the large language model is used to generate the answer based on the semantic label lv and the description information d of the edge v and the relationship label l of the edge e and the description information d of the edge e If the exact match subgraph set S is empty but the approximate match set S' is not empty, an approximate match subgraph g is used as the prompt, and a large language model is used to generate an answer based on the semantic label l of the node of g ′ and the description information d of the edge ′ and the relationship label l of the edge v and the description information d of the edge v and the relationship label l of the edge e and the description information d of the edge e and the description information d of the edge v and the description information d of the edge v and the description information d of the edge e If none of the above conditions are met, a subgraph g" is constructed by searching for the 1-hop neighbors of the known entities in the query graph q, and g" is used as the prompt, and a large language model is used to generate an answer based on the description information d and the semantic label l of the relevant nodes in g"

[0116] The subgraph isomorphism detection is embedded in the RAG framework, breaking the limitation of traditional RAG relying only on vector similarity for fuzzy matching. In addition, the possible values of unknown nodes in the query are dynamically completed, facilitating the subsequent matching process. Through the use of label embedding and dominant embedding for double filtering, the exact match and approximate match are separated, achieving significant technical breakthroughs and application value. The search results are completely consistent with the topological logic of the user query, thereby completely eliminating structural hallucinations, supporting end-to-end logical chain verification across multiple paths, solving the fragmentation problem of traditional RAG, supporting natural language queries with unknown node / edge labels through zero embedding and Cartesian product generation, and forcing LLM to output based on exact match subgraphs in the answer generation stage, eliminating fabricated relationships, and generating content that can be traced back to specific paths in the knowledge graph.

[0117] According to the second specific embodiment of the present application, a retrieval enhancement generation system based on exact subgraph matching is provided, which adopts the method of the first specific embodiment, as shown in Figure 2 The retrieval enhancement generation system based on exact subgraph matching 400 includes:

[0118] The offline preprocessing module 410 is configured to generate, in an offline preprocessing data graph G stage, a node label embedding vector in the data graph G by using the LLM; generate, for a path with a length less than or equal to d in the data graph G, a path label embedding vector by splicing the node label embedding vectors; generate a node dominant embedding vector in the data graph G according to the trained GNN model M; splice the node dominant embedding vectors in a path order to generate a path dominant embedding vector with a length of 1-d based on the node dominant embedding vectors; and construct a hierarchical R-Tree index for all the paths with the length of 1-d in the data graph G, wherein the leaf nodes store the path label embedding vectors and the path dominant embedding vectors, and the non-leaf nodes store minimum bounding rectangles; and d is a preset value.

[0119] The online retrieval module 420 is configured to process, in an online retrieval stage, query information input by a user to generate a normalized query graph q based on the LLM and an entity normalization algorithm; divide the normalized query graph q to obtain a query path set Q; obtain a candidate label embedding vector set of an unknown node in the query path set Q by using a candidate label algorithm for an unknown node based on the LLM and the hierarchical R-Tree index; obtain a candidate path set P of all the nodes in the query path set Q by using a Cartesian product method according to the query path set Q and the candidate label embedding vector set; and obtain an exact match subgraph set S and an approximate match subgraph set S' by using an exact subgraph matching algorithm based on the candidate path set P, the trained GNN model M, and the hierarchical R-Tree index; wherein the exact match subgraph set S is a set of exact match subgraphs g of the data graph G, and the exact match g is isomorphic to the normalized query graph q; and the approximate match subgraph set S' is a set of approximate match subgraphs g' of the data graph G, and the approximate match subgraph g' is approximately matched to the normalized query graph q.

[0120] The answer generation module 430 is configured to generate a trusted answer based on the exact match subgraph set S and the approximate match subgraph set S'.

[0121] According to a third specific embodiment of the present application, an electronic device is provided, such as Figure 3 as shown in FIG. 8, Figure 3 is a block diagram of an electronic device according to an example embodiment.

[0122] The electronic device 200 according to this embodiment of the present application will be described below with reference to Figure 3 FIG. 8. Figure 3 The electronic device 200 shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present application.

[0123] As shown in FIG. 8, Figure 3As shown, the electronic device 200 is in the form of a general computing device. Components of the electronic device 200 can include, but are not limited to, at least one processing unit 210, at least one storage unit 220, a bus 230 that connects the various system components, including the storage unit 220 and the processing unit 210, a display unit 240, and the like.

[0124] The storage unit stores programming code that can be executed by the processing unit 210 so that the processing unit 210 performs the steps described in this specification in accordance with the various example embodiments of the present application. For example, the processing unit 210 can perform the steps shown in FIG. 2. Figure 1

[0125] The storage unit 220 can include a readable medium in the form of volatile storage such as a random access memory (RAM) 2201 and / or cache memory 2202, and can further include a read-only memory (ROM) 2203.

[0126] The storage unit 220 can also include a program / utility 2204 having a set of programs / modules 2205, each of which performs one or more of the steps described in this specification, and / or some combination of these examples. The programs 2205 can include an operating system. Application programs 2206, other program modules 2207, and program data 2208, examples of which are described above. Programs 2205 can also include a graphical user interface utility.

[0127] The bus 230 can represent one or more of several types of bus structures, including a storage bus or bus controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus architectures.

[0128] The electronic device 200 can also communicate with one or more external devices 200' such as a keyboard or pointing device, a Bluetooth device, etc. so that a user can interact with the electronic device 200 and / or one or more other computing devices (e.g., a router, a modem, etc.) so that the electronic device 200 can communicate with one or more other computing devices. Such interaction can occur through an input / output (I / O) interface 250. Additionally, the electronic device 200 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or the public network such as the Internet) through a network adapter 260. The network adapter 260 can communicate with the other modules of the electronic device 200 through the bus 230. It should be appreciated that although not shown, other hardware and / or software modules can be used in conjunction with the electronic device 200, including but not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc. ​

[0129] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware.

[0130] Therefore, according to the fourth specific embodiment of the present application, the present application provides a computer readable medium. As shown in the above, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or on a network, and includes a plurality of instructions to make a computing device (which can be a personal computer, a server, or a network device, etc.) execute the above-mentioned method according to the embodiments of the present application. Figure 4

[0131] The software product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0132] The computer readable storage medium can include a data signal carried in the baseband or as a part of a carrier wave propagating through the program code readable. Such a propagating data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate or transmit programs for use by or in conjunction with an instruction execution system, device or component. The program code contained on the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0133] ​The program code can be executed by one or more programmable processors, digital signal processors, ASICs, FPGAs, GPUs, microprocessors, etc. The program code can execute entirely on one or more user computers, partly on the user computers, as a stand-alone software package, partly on the user computers and partly on one or more remote computers or servers, or entirely on one or more remote computers or servers. In the latter scenario, the remote computers can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). These networks and / or computer systems can use many well-known protocols to communicate, such as TCP / IP, UDP, HTTP, HTTPs, SMTP, etc. The user computers can be one of many types of user computers, including, for example, a desktop computer, a notebook computer, a tablet computer, a handheld computer, a PDA, a cell phone, an embedded system, etc.

[0134] The computer readable medium described above can carry one or more programs which, when executed by one of the devices, cause the computer readable medium to perform the functions of the first embodiment.

[0135] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiment, and can also be changed in one or more devices different from the embodiment. The modules of the above-mentioned embodiment can be combined into one module, or can be further split into a plurality of sub-modules.

[0136] From the above description of the embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.) or a network, and includes a plurality of instructions to make a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) execute the method according to the embodiments of the present application.

[0137] The above merely describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for retrieval augmentation generation based on exact subgraph matching, characterized in that, Comprise: In the offline pre-processing data graph G stage, the node label embedding vector in the data graph G is generated by LLM; for the path with length less than or equal to d in the data graph G, the path label embedding vector is generated by splicing the node label embedding vector; the node dominant embedding vector in the data graph G is generated according to the trained GNN model M; the path dominant embedding vector with length 1-d is generated by splicing according to the node dominant embedding vector; the hierarchical R-Tree index is constructed for all paths with length 1-d in the data graph G, wherein the leaf node stores the path label embedding vector and the path dominant embedding vector, and the non-leaf node stores the minimum circumscribed rectangle range; d is a preset value; In the online retrieval stage, the query information input by the user is processed to generate the normalized query graph q based on LLM and entity normalization algorithm; the query path set Q is obtained by path division of the normalized query graph q; the candidate label embedding vector set of unknown nodes in the query path set Q is obtained by using the candidate label algorithm of unknown nodes based on LLM and hierarchical R-Tree index; the candidate path set P of all nodes in the query path set Q is obtained by calculating the Cartesian product according to the query path set Q and the candidate label embedding vector set; the accurate matching subgraph set S and the approximate matching subgraph set S' are obtained by using the accurate subgraph matching algorithm based on the candidate path set P, the trained GNN model M and the hierarchical R-Tree index; wherein the accurate matching subgraph set S is the set of accurate matching subgraphs g of the data graph G, and the accurate matching g is isomorphic with the normalized query graph q; the approximate matching subgraph set S' is the set of approximate matching subgraphs g' of the data graph G, and the approximate matching subgraph g' is approximately matched with the normalized query graph q; The trusted answer is generated based on the accurate matching subgraph set S and the approximate matching subgraph set S'.

2. The method of claim 1, wherein, Also include: For a vertex set V(G), extract the unit star subgraph of each node v i where the unit star subgraph refers to a subgraph containing node v i and all its 1-hop neighbors;​ Extracting unit star subgraphs Unit star substructures wherein pairing of subgraphs adding dataset D and shuffling the pairs of dataset D randomly; Input the data set D and the learning rate η into the GNN model M to be trained; dividing the dataset D into a plurality of batches B, for each batch B, computing, by the GNN model M to be trained, an embedding vector for the subgraph G . Calculate the loss function L(B); When the accumulated loss L of the GNN model M to be trained in the training process is 0 e the trained GNN model M is obtained.

3. The method of claim 2, wherein, The loss function L(B) is calculated by the following loss function formula: wherein, and corresponding to and embedding vectors.

4. The method of claim 3, wherein, The query information input by the user is processed to generate the normalized query graph q based on LLM and entity normalization algorithm, which comprises: extracting, by the LLM, entity from the query information input by the user to obtain a query graph q composed of entities and corresponding relationship constraints * and generating a node label embedding o0(q i ) thereof; copy query graph q * to get node set V(q); traversing each vertex in the set of nodes V(q) generating a label embedding vector for the vertex label embedding vector The most similar entity q in the data graph G is found by an approximate nearest neighbor search and vector compression techniques i ;​ In the set of nodes V(q) replace q i with to obtain the normalized query graph q.

5. The method of claim 4, wherein, The query path set Q is obtained by path division of the normalized query graph q, which further comprises: Select the node with the highest degree in the normalized query graph q as the starting point; According to the greedy strategy, the path expansion with the least overlap with the existing path and the lowest calculation cost is selected to generate the query path set Q covering all edges of the query graph with the lowest total cost.

6. The method of claim 5, wherein, The candidate label embedding vector set of unknown nodes in the query path set Q is obtained by using the candidate label algorithm of unknown nodes based on LLM and hierarchical R-Tree index, which comprises: Set the label embedding vector of unknown nodes in the query path set Q to zero; The candidate label set is generated by matching the path label embedding vector through the hierarchical R-Tree index.

7. The method of claim 6, wherein, The trusted answer is generated based on the accurate matching subgraph set S and the approximate matching subgraph set S', which comprises: Determine whether the accurate matching subgraph set S is empty or not; If the exact matching subgraph set S is a non-empty set, an exact matching subgraph g is taken as an input of the LLM, so that the LLM generates an answer according to the semantic label of the node of the exact matching subgraph g, the description information, the relation label of the edge, the description information of the edge, and the node information associated with the edge; If the exact matching subgraph set S is an empty set, it is judged whether the approximate matching subgraph set S' is an empty set; If the set of approximate matching subgraphs S' is a non-empty set, the approximate matching subgraph g ′ as input of the LLM, so that the LLM generates an answer according to the semantic label of the node, the description information of the node, the relation label of the edge, the description information of the edge, and the node information associated with the edge of the approximate matching subgraph g ′ .

8. The method of claim 7, wherein, Further comprising: If both the exact matching subgraph set S and the approximate matching subgraph set S' are empty sets, retrieve 1-hop neighbors of known entities in the normalized query graph q, construct a subgraph g ” , and input the subgraph g ” into the LLM to enable the LLM to generate an answer based on the description information of related nodes, the semantic label, the relationship label of edges, and the description information of edges in g ” .

9. A search enhancement generation system based on exact subgraph matching, characterized in that, Further comprising: The offline preprocessing module is configured to generate a node label embedding vector in the data graph G by the LLM in the offline preprocessing data graph G stage; for a path with a length less than or equal to d in the data graph G, the node label embedding vector is spliced to generate a path label embedding vector; a node dominant embedding vector in the data graph G is generated according to the trained GNN model M; the node dominant embedding vector is spliced in the path order to generate a path dominant embedding vector with a length of 1-d; a hierarchical R-Tree index is constructed for all paths with a length of 1-d in the data graph G, wherein the leaf node stores the path label embedding vector and the path dominant embedding vector, and the non-leaf node stores the minimum circumscribed rectangle range; d is a preset value; The online retrieval module is configured to process the query information input by the user to generate a normalized query graph q based on the LLM and the entity normalization algorithm in the online retrieval stage; the normalized query graph q is divided into a query path set Q; the candidate label embedding vector set of the unknown node in the query path set Q is obtained by using the candidate label algorithm of the unknown node based on the LLM and the hierarchical R-Tree index; the candidate path set P of all nodes in the query path set Q is obtained by taking the Cartesian product according to the query path set Q and the candidate label embedding vector set; the exact matching subgraph set S and the approximate matching subgraph set S' are obtained by the exact subgraph matching algorithm based on the candidate path set P, the trained GNN model M, and the hierarchical R-Tree index; wherein the exact matching subgraph set S is a set of exact matching subgraphs g of the data graph G, and the exact matching g is isomorphic to the normalized query graph q; the approximate matching subgraph set S' is a set of approximate matching subgraphs g' of the data graph G, and the approximate matching subgraph g' is approximately matched with the normalized query graph q; The answer generation module is configured to generate a trusted answer based on the exact matching subgraph set S and the approximate matching subgraph set S'.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium includes a computer program or instructions, when the computer program or instructions run on the computer, so that the computer executes the method as claimed in any one of claims 1-8. The computer readable storage medium includes a computer program or instructions, when the computer program or instructions run on the computer, so that the computer executes the method as claimed in any one of claims 1-8.