Data query method based on graph convolutional network
Through multi-layer semantic networks and hierarchical topological sequences based on graph convolution networks, the problems of information loss and insufficient semantic processing capabilities of traditional methods when processing heterogeneous data are solved, and efficient and accurate data query and real-time processing are achieved.
Patent Information
- Application Number
- CN202510353681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional methods have problems such as information loss, semantic distortion, weak semantic processing capabilities and difficulty in dealing with real-time queries and dynamic updates when processing heterogeneous data, resulting in data queries being inefficient and accurate enough.
The data query method based on graph convolution network is adopted, and efficient query of heterogeneous data is achieved by building multi-layer semantic networks, generating hierarchical topological sequences, performing semantic matching and structural verification, and combining GCN path embedding and pruning strategies.
It significantly improves the efficiency of semantic information integration and matching of heterogeneous data, improves the accuracy of semantic matching, reduces the computational complexity, and supports real-time processing of large-scale graph data.
Smart Images

Figure CN120296177A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data query, and specifically, to a data query method based on a graph convolutional network. Background Art
[0002] In the current data-driven era, with the explosive growth of multi-source heterogeneous data and the continuous increase in complexity, how to effectively manage and utilize these data has become a key problem that needs to be solved urgently. Multi-source heterogeneous data comes from different data sources and covers structured, semi-structured, and unstructured data types. These data often have different formats, structures, and semantic features, and this diversity and complexity pose great challenges to unified query and fusion. In the big data environment, in the face of the rapidly growing data volume and real-time requirements, traditional processing methods are difficult to meet the high efficiency and accuracy requirements of modern data analysis.
[0003] Traditional methods usually rely on a unified data model or a fixed fusion strategy to forcibly convert data from different sources into a unified format. However, this method has many limitations in dealing with heterogeneous data. First, there may be deep semantic associations between data from different sources, and simple format conversion often leads to information loss and semantic distortion, resulting in limited data expression ability and difficulty in supporting queries of complex semantic relationships. Second, the semantic processing ability of traditional methods is weak, and most only rely on static semantic matching rules and cannot capture the dynamic semantic connections or context associations between data, resulting in inaccurate matching results. Finally, with the expansion of the data scale and the increase in the update frequency, static fusion methods are difficult to meet the requirements of real-time query and dynamic update, and the problem of insufficient system flexibility becomes more prominent. Summary of the Invention
[0004] The purpose of the present invention is to provide a data query method based on a graph convolutional network to solve the problem that traditional methods usually rely on a unified data model or a fixed fusion strategy to forcibly convert data from different sources into a unified format as proposed in the above background art.
[0005] To achieve the above purpose, the present invention provides a data query method based on a graph convolutional network, including the following steps:
[0006] S1. Construct an extraction and multi-layer semantic network for heterogeneous data;
[0007] S2. Generate a multi-layer topology network conversion and a hierarchical topology sequence;
[0008] S3. Semantic matching;
[0009] S4. Subgraph matching based on GCN path embedding;
[0010] S5. Structure matching and isomorphism verification.
[0011] As a further improvement of this technical solution, the construction of the extraction of heterogeneous data and the multi-layer semantic network in the step S1 specifically includes the following steps:
[0012] S11. Extraction and preprocessing of heterogeneous data;
[0013] S12. Construction of a multi-layer semantic network;
[0014] S13. Node and relationship mapping of the multi-layer semantic network;
[0015] S14. Vectorized representation;
[0016] S15. Strengthening of graph representation based on GCN.
[0017] As a further improvement of this technical solution, the strengthening of graph representation based on GCN in the step S15 specifically includes the following steps:
[0018] S151. Overview of the GCN model;
[0019] S152. Data preparation and feature initialization;
[0020] S153. Convolutional layer and model training;
[0021] S154. GCN representation of the temporal graph.
[0022] As a further improvement of this technical solution, the conversion of the multi-layer topology network and the generation of the hierarchical topology sequence in the step S2 specifically include the following steps:
[0023] S21. Conversion of the multi-layer topology network;
[0024] S22. Generation of the hierarchical topology sequence.
[0025] As a further improvement of this technical solution, the generation of the hierarchical topology sequence in the step S22 specifically includes the following steps:
[0026] S221. Initialization: Take all nodes with in-degree 0 as the output of the first layer, mark their parallel relationships, and at the same time delete these nodes and their associated edges, and update the remaining network;
[0027] S222. Layer-by-layer recursion: Repeat the above operations, each time select a new set of nodes with in-degree 0 as the next layer until all nodes are output. The output sequence is represented in the following format: S = C1 / C2 / … / C n where C i is the set of nodes in the i-th layer, and " / " represents the inter-layer separator;
[0028] S223. Symbol representation rule: Each layer of nodes is wrapped by "{}", and parallel nodes within a layer are separated by ",", for example, {c1, c2}; the out-degree relationship of nodes is represented by "[]".
[0029] As a further improvement of this technical solution, the GCN representation of node semantic information in step S3 specifically includes the following steps:
[0030] S31. GCN representation of node semantic information;
[0031] S32. Edge-level semantic matching;
[0032] S33. Local consistency verification;
[0033] S34. Semantic matching output.
[0034] As a further improvement of this technical solution, the subgraph matching based on GCN path embedding in step S4 specifically includes the following steps:
[0035] S41. Pruning strategy;
[0036] S42. Index mechanism;
[0037] S43. Index-level pruning;
[0038] S44. Subgraph matching algorithm based on GCN.
[0039] As a further improvement of this technical solution, the pruning strategy in step S41 specifically includes the following steps:
[0040] S411. Path label pruning: For a path p of length l z , use GCN to obtain its node embedding, and then form a path label embedding vector o0(p z ) by concatenating or aggregating node labels. During the pruning process, if the label embedding o0(p z ) of the data graph path p z does not match the label embedding o0(p z ) of the query path o0(p q ), that is, o0(p z ) ≠ o0(p q ), then prune this path from the matching search space;
[0041] S412. Path dominance pruning: Through specific aggregation operations on the node vectors output by GCN, maximizing, minimizing, or dominating vector concatenation, obtain the path embedding o(p z ). If the embedding o(p q ) corresponding to the query path p q ) does not dominate (or is not dominated by) the data path pz embedding of o0(p z ), that is, this path is determined to be invalid in subgraph matching and is quickly excluded. If the level of the node to which the path belongs conflicts with the level of the corresponding node in the query path, level pruning is directly performed.
[0042] As a further improvement of this technical solution, the structure matching and isomorphism verification in step S5 specifically include the following steps:
[0043] S51. Global isomorphism of nodes and edges;
[0044] S52. Backtracking of the hierarchical topological sequence.
[0045] Compared with the prior art, the beneficial effects of the present invention are:
[0046] 1. In the present invention, through the construction of a multi-layer semantic network covering the entity relationship layer, attribute relationship layer, and time sequence relationship layer, the semantic information of heterogeneous data is comprehensively unified. Using graph embedding technology, high-dimensional vector representations are made for nodes and edges, ensuring the integrity and consistency of data semantics, and significantly improving the efficiency of semantic information integration and matching.
[0047] 2. In the present invention, based on the GCN model, semantic embedding and encoding are performed on nodes and edges, realizing multi-level semantic information capture. Combining hierarchical topological sequence constraints, only nodes in the same layer or adjacent layers are matched, reducing the search space and improving the accuracy of semantic matching at the same time.
[0048] 3. In the present invention, by constructing a hierarchical topological network, a structure matching mechanism based on the topological sequence is proposed. Combining node features, a standardized feature vector is generated. Further, through topological constraints and global isomorphism verification, the consistency of the local and global topological structures between the query graph and the data graph is ensured, greatly reducing the computational complexity of traditional graph isomorphism verification.
[0049] 4. In the present invention, a hierarchical topological sequence generation and path matching method based on GCN embedding are adopted to achieve efficient query of heterogeneous data. Through the indexing mechanism and pruning strategy, the real-time processing ability of large-scale graph data is significantly improved, providing flexible and efficient support for complex query scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a schematic diagram of the overall step flow of the data query method based on a graph convolutional network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] In a specific embodiment, as Figure 1 shown, the present invention provides a data query method based on a graph convolutional network, including the following steps:
[0053] Step 1: Extraction of heterogeneous data and construction of a multi-layer semantic network.
[0054] Step 1.1: Extraction and preprocessing of heterogeneous data.
[0055] First, heterogeneous data is extracted from multiple different data sources, including structured data (such as relational databases) and unstructured data (such as text files, social media data, etc.). During the extraction process, data cleaning and standardization techniques are applied to ensure the integrity and consistency of the data. The following formula is used for filling missing values:
[0056]
[0057] where X i represents the i-th data point, is the average value of all data points, and ∈ is a random error.
[0058] Step 1.2: Construction of a multi-layer semantic network.
[0059] After the data preprocessing is completed, the heterogeneous data is constructed into a multi-layer semantic network. Each network layer represents different types of semantic relationships: the entity relationship layer is composed of entity nodes and entity relationships, and its triple form is <entity 1, relationship, entity 2>, which describes the direct relationship between entities. There are internal associations among these three layers, forming a multi-dimensional semantic graph structure from entities to attributes and then to temporal changes.
[0060] The entity relationship layer G E =(V E , R E , T E ): where V E is the set of entity nodes, R E is the set of entity relationships, T E is the set of triples in the form of that describe the direct relationship between entities.
[0061] The attribute relationship layer G A =(VA , R A , T A ): where V A is a set of attribute nodes, R A is a set of attribute relationships, T A is a set of triples in the form of that describes the relationship between an entity and its attributes.
[0062] Temporal relationship layer G T = (V T , R T , T T ): where V T is a set of time nodes, R T is a set of temporal relationships, T T is a set of triples in the form of that describes the relationship changing over time.
[0063] When constructing these semantic layers, the relevance between layers needs to be ensured. For example, the entity nodes in the entity relationship layer can be used as the starting points of attributes in the attribute relationship layer, and the attribute nodes in the attribute relationship layer can be used as the associated objects of time nodes in the temporal relationship layer.
[0064] Step 1.3: Node and relationship mapping of the multi-layer semantic network.
[0065] To achieve a consistent representation between different data sources, the entities and relationships of each data source are mapped to the nodes and edges of the multi-layer semantic network. (The main purpose of mapping nodes and relationships to the multi-layer semantic network is to unify the semantic representations of heterogeneous data sources and lay a foundation for subsequent embedding and query operations.) Set the mapping function f: V source → V network . A mapping that satisfies the following conditions is considered consistent:
[0066] Semantic equivalence condition (The semantic equivalence condition ensures that synonymous entities in different data sources can be correctly aligned): For each pair of entities v s ∈ V source and v n ∈ V network , if f(v s ) = v n , then they must be semantically equivalent, denoted as v s ≡ v n , satisfying:
[0067] sim(v s , v n ) ≥ τ
[0068] where sim is the semantic similarity function and τ is a predefined similarity threshold.
[0069] Relationship preservation condition (ensuring the integrity of relationships in the semantic network): For each triple In the source data, if holds in the network, then r s and r are semantically equivalent, denoted as r s ≡ r. Through this multi-level network representation, various semantic relationships from different data sources can be effectively captured and integrated.
[0070] Step 1.4: Vector representation.
[0071] After completing the mapping of nodes and relationships, in order to process and analyze the data in the semantic network more efficiently, the nodes and edges of each layer of the network are converted into vector representations. By using graph embedding techniques such as Node2Vec or GraphSAGE, a high-dimensional vector representation z v can be generated for each node, where:
[0072]
[0073] z v is the vector representation of node v, N(v) is the set of neighbor nodes directly connected to node v, is the vector representation of neighbor node u. This embedding method ensures that the node representation not only retains semantic information but also supports efficient semantic matching and topological verification.
[0074] Step 1.5: Graph representation enhancement based on GCN.
[0075] Based on the vectorization of data in Step 1.4, the method of graph convolutional network (GCN, Graph Convolutional Network) is further introduced to perform deep representation learning on the nodes and their local subgraphs in the multi-layer semantic network, improving the representation ability for heterogeneous graphs and temporal graphs.
[0076] Step 1.5.1: Overview of the GCN model.
[0077] Compared with traditional graph embedding techniques such as Node2Vec and GraphSAGE based on random walks or simple neighbor aggregation, GCN aggregates the features of node neighbors through spectral domain convolution or spatial domain convolution, and can more flexibly fuse the features of neighbor nodes, graph structure information, and the context labels of the nodes themselves, thereby forming an embedding vector with stronger semantic and structural representation capabilities. Its core idea can be briefly summarized as follows:
[0078] Neighbor aggregation: The convolution operation of each layer aggregates the features of the node itself and its neighbor nodes together;
[0079] Weight sharing: The convolutional kernels in GCN are shared across the entire graph, different from the mechanism of traditional CNNs which target local regions (convolutional kernels) of images;
[0080] Hierarchical: After stacking multiple layers of GCN, nodes will gradually perceive neighbors at greater distances, thus capturing cross-local graph structures and semantic information.
[0081] Step 1.5.2: Data preparation and feature initialization.
[0082] In a heterogeneous multi-layer semantic network, each node may contain multiple features (such as text descriptions, attribute values, timestamps, etc.). When using GCN, these original features need to be appropriately encoded:
[0083] Label / attribute encoding: Methods such as one-hot vectors and Embedding layers (such as word vectorization for text attributes) can be used to convert discrete or text-based features into vectors that can be input into GCN.
[0084] Structural feature supplementation: The degree, level information, time slice information, etc. of the nodes can be concatenated into the initial feature vectors of the nodes.
[0085] Step 1.5.3: Convolutional layer and model training.
[0086] GCN convolutional layer: In the spectral domain method, GCN transforms node features through an approximate Laplacian matrix; in the spatial domain method, GCN aggregates and updates node and neighbor features through convolutional operations, with the formula as follows: where, (adding self-connections to the adjacency matrix), is the corresponding degree matrix, H (l) represents the node representation matrix of the l-th layer, W (l) are learnable parameters, and σ is the activation function.
[0087] Training strategy: For a multi-layer semantic network, some known node labels (such as entity types) or relationship labels can be used as supervision signals for supervised or semi-supervised training; if temporal information needs to be captured, the graphs of different time slices can also be made into sequence snapshots, and the node representations can be learned layer by layer or in segments in GCN.
[0088] Output vector: After training, each node v can obtain a high-dimensional vector o(v). In the subsequent "Step 3: Semantic matching" and "Step 5: Structural matching", similarity comparisons, candidate generation, and isomorphism verification will be directly based on these node vectors and edge vectors, thus making full use of the deep representations of GCN to capture semantics and structure.
[0089] Step 1.5.4: GCN Representation of the Timing Diagram
[0090] For heterogeneous networks with a time dimension (temporal relationship layer), the following are adopted:
[0091] Multi-snapshot GCN: The network is split into multiple "snapshots" in chronological order, and the GCN is trained / updated for each time slice separately;
[0092] Time-decaying weights: When constructing the adjacency matrix or degree matrix, smaller weights are assigned to edges that are too far apart in time, so that the GCN pays more attention to recent temporal correlations;
[0093] Edge attribute encoding: Edge attributes such as timestamps are added to the GCN convolution, so that nodes can perceive the temporal differences on the edges.
[0094] Step 2: Multi-layer Topological Network Transformation and Hierarchical Topological Sequence Generation
[0095] Step 2.1: Transformation of the Multi-layer Topological Network
[0096] Each node and edge in the multi-layer semantic network is further structured into a multi-layer topological network through topological transformation. The construction of the topological network mainly includes the following steps:
[0097] Each node v in the semantic network is represented as a subgraph that describes all the adjacency relationships of the node.
[0098] Define the levels and connection rules of the nodes: If there is a direct semantic relationship between two nodes, their subgraphs are connected by an edge. The original semantic information is retained among the nodes inside the subgraph according to the adjacency relationship.
[0099] For example, the conversion result of <Entity A, Relationship R, Entity B> in the semantic network can use "Relationship R" as an edge to connect the subgraphs of "Entity A" and "Entity B" in the topological network.
[0100] Step 2.2: Generation of the Hierarchical Topological Sequence
[0101] In the multi-layer topological network, generating the hierarchical topological sequence is a key step in analyzing the topological structure. Through the topological sorting algorithm, a strict hierarchical relationship is defined for the nodes in the network. In the subsequent semantic matching and structure matching processes, this sequence can be used to constrain the nodes to map only in the same or adjacent layers, thereby reducing unnecessary matching attempts and ensuring layer consistency. The specific steps are as follows:
[0102] Initialization: All nodes with an in-degree of 0 are used as the output of the first layer, and their parallel relationships are marked. These nodes and their associated edges are deleted, and the remaining network is updated.
[0103] Layer-by-layer recursion: Repeat the above operations. Each time, select a new set of nodes with in-degree 0 as the next layer until all nodes are output. The output sequence is represented in the following format: S = C1 / C2 / … / C n where C i is the set of nodes in the i-th layer, and " / " represents the inter-layer separator.
[0104] Symbol representation rules: Each layer of nodes is wrapped with "{}", and parallel nodes within a layer are separated by ",", e.g., {c1,c2}; the out-degree relationship of nodes is represented by "[", e.g., c1[c2,c3]; the inter-layer relationship is separated by " / ", e.g., {c1[c2],c3[c4]} / {c2[c5]}.
[0105] Step 3: Semantic matching (emphatically combined with GCN embedding).
[0106] 3.1: GCN representation of node semantic information.
[0107] Obtaining node vectors:
[0108] Through the trained GCN, obtain the vector o(c i ) of each node c i . Let the propagation rule of the l-th layer of the GCN model be where H (k) is the node feature matrix of the k-th layer, is the adjacency matrix after adding self-loops, is its corresponding degree matrix, W (k) is the weight matrix of the l-th layer, and σ is the activation function. After multiple layers of propagation, the finally obtained o(c i ) integrates text attributes (if any) and graph structure information. During the training process of the GCN model, it will learn an effective representation for each node based on the initial features of the nodes and the topological structure of the graph, so that o(c i ) can capture the semantic and structural characteristics of the nodes.
[0109] If a node contains a large amount of text description, BERT can first be used to represent the text as a vector e text . Assuming that the structural feature vector of the node is e struct , then the concatenated vector e = [e text ; e struct is input into the GCN. In this way, while learning the graph structure information, the GCN can make full use of the text semantic information extracted by BERT, so that the final o(c i ) has both rich language semantics and accurate graph structure information, and can more comprehensively describe the semantic features of the nodes.
[0110] Cosine similarity:
[0111] Query node q j and data node c i 's matching degree can be measured by . When query node q j and data node c i 's cosine similarity is greater than τ and they are at the same hierarchical topological level, then (q j , c i ) can be regarded as a preliminary candidate node pair; if not at the same level or the similarity is insufficient, it is directly excluded, that is, when cos(o(q j ), o(c i )) > τ, node pairs that are semantically more similar can be screened out as candidates for subsequent further analysis and matching.
[0112] Hierarchical topological sequence constraint:
[0113] If query node q j is located at the k-th layer, let the hierarchical topological sequence of the query graph be L q , and the hierarchical topological sequence of the data graph be L d . Determine the level of q q from L j as l q (q j ) = k, then only calculate cos(·) among the candidate nodes at the k-th layer determined from L d in the data graph. In this way, according to the constraint of the hierarchical topological sequence, the matching range can be restricted to nodes at the same level, avoiding unnecessary matching calculations between irrelevant levels, reducing a large number of unnecessary matching attempts, and improving the matching efficiency.
[0114] 3.2: Edge-level semantic matching (combining GCN / edge vectors).
[0115] Edge vector representation:
[0116] Use the edge information aggregation mechanism of GCN to obtain the edge vector o(e ij ). Assume that GCN obtains the edge vector by performing the following operations on the feature vectors o(c i ) and o(c j ) at both ends of the edge: o(e ij ) = σ(W1[o(c i );o(c j )]+b1), where W1 is the weight matrix, b1 is the bias vector, and σ is the activation function. GCN can generate a vector that can represent the semantic and structural information of the edge by combining, transforming, and aggregating the features of the nodes at both ends of the edge. Or use the simple method of "node vector splicing + MLP", and splice the vectors o(c i ) and o(cj ) are concatenated to obtain e concat = [o(c i ) ; o(c j )], and then through a multi-layer perceptron (MLP). Assuming the operation of the MLP is o(e ij ) = MLP(e concat ) = W2σ(W1e concat + b1) + b2 (where W1, W2 are weight matrices, and b1, b2 are bias vectors) for feature extraction and transformation to obtain the edge vector o(e ij ).
[0117] If the edge relationship has a time sequence t and attribute information a, the time sequence information is encoded as a vector e t , and the attribute information is encoded as a vector e a , which can be incorporated into the feature channels during aggregation. For example, in the GCN method, the input becomes e input = [o(c i ) ; o(c j ) ; e t ; e a , and then through the above GCN operation to obtain the edge vector; in the "node vector concatenation + MLP" method, after e concat = [o(c i ) ; o(c j ) ; e t ; e a , it is then passed through the MLP operation so that the edge vector can contain the time sequence and attribute features of the edge, more comprehensively describing the semantics of the edge.
[0118] Similarity measure:
[0119] The similarity between the query edge (q j , q k ) and the data edge (c i , c l ) is measured by . Only when the similarity (τ edge is the preset threshold for edge similarity) can it possibly become a matching edge.
[0120] After completing node matching, it is necessary to perform a semantic comparison between the edges (q j , q k ) in the query graph and the edges (c i , c l ) in the data graph. First, verify the connection consistency of the edges, that is, the starting point and ending point of the query edge must match the starting point and ending point of the data edge:
[0121] Let the query node q j match the data node c iDenoted as q j ~c i , the query node q k matches the data node c l Denoted as q k ~c l , then it is necessary to satisfy (q j ~c i ) ∧ (q k ~c l ).
[0122] Secondly, check whether the semantic labels of the edges are consistent: Let the semantic label of the query edge (q j , q k ) be r q , and the semantic label of the data edge (c i , c l ) be r d , then it is required that r q = r d .
[0123] If it meets the threshold and the temporal / type consistency of the edge passes the check, its semantic matching can be established.
[0124] 3.3: Local Consistency Verification.
[0125] Neighbor Coverage:
[0126] If the neighbor of the query node q j is and the neighbor of the data node c i is After the preliminary matching of the candidate nodes is completed, it should be ensured that such that n q ~n d (in the sense of matching mapping). This means that the neighbor nodes of the query node can all find corresponding matching neighbor nodes in the data graph. By checking the neighbor coverage, the rationality and accuracy of the matching can be further verified, and isolated matching nodes can be avoided.
[0127] Hierarchical Sequence Verification:
[0128] If (q j , q k ) points from the k-th layer to the k + 1-th layer, let the hierarchical topological sequence of the query graph be L q , determine from L q that q j is located in the k-th layer, denoted as l q (q j ) = k, q k is located in the k + 1-th layer, denoted as l q (q k ) = k + 1, then its matching edge (ci , c l ) should also satisfy the same hierarchical relationship. Let the hierarchical topological sequence of the data graph be L d , from L d , it is necessary to determine c i is located on the k-th layer, denoted as l d (c i ) = k, c l is located on the (k + 1)-th layer, denoted as l d (c l ) = k + 1. That is, the matching edges should also point from the k-th layer to the (k + 1)-th layer. Through this verification of the hierarchical sequence, it can be ensured that the matching edges are consistent in the hierarchical order and conform to the topological structure of the entire graph.
[0129] If the hierarchical order is violated or there is a loop conflict, it can be determined that its local consistency fails. If the hierarchical order of the matching edges is inconsistent with the query edges, or there is a loop conflict during the matching process, for example, an unreasonable circular dependency relationship is formed, it indicates that there may be a problem with the current matching, and the local consistency cannot be satisfied, and re-matching or adjustment is required.
[0130] 3.4: Semantic matching output.
[0131] The matching pairs obtained under the double constraints of "GCN embedding + hierarchical topological sequence" have higher accuracy. GCN embedding can accurately capture the semantic and structural information of nodes and edges, providing a more accurate matching basis; while the hierarchical topological sequence constrains the matching in terms of hierarchy and order, reducing the possibility of incorrect matching. The combination of the two can obtain more accurate and reliable matching results.
[0132] Output candidate node pairs and edge pairs to form candidate subgraphs, preparing for the structure matching stage. Subsequently, in the structure matching, it will be further verified based on the GCN vectors and hierarchical information whether the isomorphism requirements are satisfied.
[0133] Step 4: Subgraph matching based on GCN path embedding.
[0134] To improve the efficiency of subgraph matching, this section proposes a subgraph matching method based on GCN path embedding, which realizes the efficient retrieval of query subgraphs in large-scale data graphs through pruning strategies, indexing mechanisms, and subgraph matching algorithms.
[0135] 4.1: Pruning strategy.
[0136] 4.1.1: Path label pruning.
[0137] For a path p with length l z, use GCN to obtain its node embeddings, and then form the path label embedding vector o0(p by concatenating or aggregating node labels z ). During the pruning process, if the label embedding o0(p z ) of the data graph path p z does not match the label embedding o0(p z ) of the query path o0(p q ), that is, o0(p z ) ≠ o0(p q ), then this path can be safely pruned from the matching search space, thus effectively reducing unnecessary computational overhead.
[0138] 4.1.2: Path dominance pruning.
[0139] By performing specific aggregation operations on the node vectors output by GCN, such as maximization, minimization, or dominant vector concatenation, the path embedding o(p z ) is obtained. If the embedding o(p q ) corresponding to the query path p q does not dominate (or is not dominated by) the embedding o0(p z ) of the data path p z ), that is then this path is considered invalid in subgraph matching and can be quickly excluded. If the level of the nodes where the path belongs conflicts with the level of the corresponding nodes in the query path, hierarchical pruning can be directly performed, jointly exerting the'multiple pruning' effect with path label pruning and path dominance pruning.
[0140] 4.2: Index mechanism.
[0141] When facing ultra-large graphs, to significantly reduce the retrieval overhead, an index mechanism for path embeddings is carefully designed. First, the data graph is reasonably divided into several subgraph partitions G j , and this partitioning operation is based on factors such as the structural characteristics and node distribution of the graph, aiming to ensure that the paths within each partition have a certain similarity and relevance. Within each partition, all paths of length (l) are extracted, and then GCN is used to obtain the embeddings o(p z ) of these paths.
[0142] In terms of constructing the index structure, multi-dimensional indexes such as R*-trees are selected to organize these path vectors. R*-trees have efficient spatial data indexing capabilities and can well adapt to the multi-dimensional characteristics of path vectors. At the same time, the range of path label embeddings and the range of path dominance embeddings are stored, and these two range information provide a basis for quickly determining the overlap and dominance relationships in subsequent queries. For example, when the label embedding or dominance embedding of the query path falls within the range stored by a certain index node, the paths that may match can be quickly located, greatly reducing the retrieval time.
[0143] 4.3: Index-level pruning.
[0144] In the case of extremely large data scales, the index-level pruning strategy plays a crucial role. If the label or domination embedding of query path p q is completely disjoint from the range of a certain index node, whether in path label pruning (index-level) where o0(p q ) does not overlap with the path label range stored in the index node, or in path domination pruning (index-level) where the domination region of the query path embedding o0(p q ) has no intersection with the embedding range of the index node, it means that all paths under this index node cannot match the query path. Therefore, all paths under this index node can be pruned at once, and this batch pruning operation greatly improves the pruning efficiency and further reduces the matching search space.
[0145] 4.4: GCN-based subgraph matching algorithm.
[0146] Path extraction and embedding generation: For query graph q, all paths of length l are comprehensively extracted. These paths cover various possible connection relationships in the query graph and reflect the structural characteristics of the query graph. Then, use GCN to generate the domination embedding o(p q ) and label embedding o0(p q ) of these paths. The strength of GCN lies in its ability to deeply mine the features of nodes in the path and the relationships between them, providing accurate path representations for subsequent matching.
[0147] Index traversal and candidate path retrieval: In the data graph index, according to the previously set path label pruning and domination pruning rules, paths are quickly filtered. By comparing the embedding of the query path with the path embedding range stored in the index node, paths that do not meet the conditions are quickly excluded, and only candidate paths that meet the conditions are retained. This process greatly reduces the number of paths that need to be further processed and improves the retrieval efficiency.
[0148] Candidate subgraph generation and verification: Combine several candidate paths that have been screened into candidate subgraphs. These candidate subgraphs are the parts initially screened from a large amount of data that may match the query graph. Then, use global isomorphism verification and hierarchical topological sequence to perform exact matching with query graph q. Global isomorphism verification ensures that the candidate subgraph is exactly the same as the query graph in terms of the correspondence of nodes and edges and the topological structure, and the hierarchical topological sequence further ensures the accuracy of the matching from the perspective of the hierarchical structure. Once verified, it is determined that this candidate subgraph is isomorphic to q, completing the key step of subgraph matching.
[0149] Step 5: Structure Matching and Isomorphism Verification.
[0150] After completing path-based pruning and candidate subgraph screening, structure matching and isomorphism verification become the key steps in determining the final matching result. In this stage, strict structural verification is mainly carried out on the candidate subgraphs to ensure their topological consistency with the query graph, so as to achieve accurate subgraph matching.
[0151] 5.1: Global Isomorphism of Nodes and Edges.
[0152] Global isomorphism verification is one of the core steps in structure matching, aiming to ensure that the candidate subgraph and the query graph fully meet the isomorphism conditions in terms of the correspondence of nodes and edges. According to Whitney's isomorphism theorem or the subgraph isomorphism determination criterion, the following checks need to be carried out:
[0153] Node Bijection: Verify whether each node in the query graph can find a unique corresponding node in the candidate subgraph, and this correspondence is one-to-one. This means that the number of nodes in the two graphs is equal, and the mapping relationship of each node is clear and unique, without a situation where one query graph node corresponds to multiple candidate subgraph nodes, or vice versa. After completing local matching, each query node V q only corresponds to one data node V d .
[0154] Edge Correspondence: For each edge in the query graph, check whether the corresponding edge (V q , V d ) in the data graph meets the consistency of GCN vector / direction / label and hierarchical conditions to ensure double alignment of semantics and structure; confirm that there is an edge with the same connection relationship in the candidate subgraph. That is, the start and end points of the edge must correspond to the same nodes in the two graphs, and the attributes of the edge (such as weight, direction, etc., if any) should also be the same. For example, if the relationship type of the query edge is r q , and the relationship type of the data edge is r d , then it is required that r q = r d .
[0155] Topological Consistency: Check whether the overall topological structures of the two graphs are consistent. This requires that after the correspondence of nodes and edges is determined, the topological features such as the adjacency relationship, path length, and connectivity between nodes in the graph are consistent in the two graphs. For example, if there is a path of length 3 from node A to node B in the query graph, then there should also be a path of the same length from the corresponding nodes A' to B' in the candidate subgraph.
[0156] Through strict verification of these three aspects, the structural consistency between the candidate subgraph and the query graph is ensured, thus providing a guarantee for accurate subgraph matching.
[0157] 5.2: Hierarchical Topological Sequence Backtracking.
[0158] The hierarchical topological sequence generated in the preprocessing stage plays an important role here, providing additional constraints and verification basis for structure matching. By backtracking the hierarchical topological sequence, the following checks can be performed:
[0159] Node Level Check: Confirm whether the levels of the nodes in the query graph and the candidate subgraph are consistent. The level position of each node in its respective graph should match the hierarchical structure defined by the hierarchical topological sequence. For example, if node X in the query graph is at level 3, then the corresponding node X' in the candidate subgraph should also be at level 3.
[0160] In-degree / Out-degree Check: Compare the in-degrees and out-degrees of the nodes in the query graph and the candidate subgraph. The in-degree and out-degree of a node reflect its connection relationship and topological role in the graph, and the in-degrees and out-degrees of the corresponding nodes at the same level should be consistent. For each node v in the query graph G q and the data graph G d , extract the node degree: including the in-degree and out-degree, which are respectively defined as: deg in (v) = ∑ (u,v)∈E 1, deg out (v) = ∑ (v,u)∈E 1 where E is the set of edges in the graph. Ensure that the in-degree and out-degree of node v q in the query graph and the candidate node v d in the data graph are consistent:
[0161] deg in (v q ) = deg in (v d ), deg out (v q ) = deg out (v d )
[0162] Adjacent Layer Relationship Check: Check whether the adjacent relationships between nodes at different levels in the query graph and the candidate subgraph are consistent. Extract the characteristics of adjacent nodes, which are characterized by calculating the number of adjacent nodes and their average degree, and are defined as: where NS(v) = {u|(v,u) ∈ E} is the neighbor set of node v. Ensure that the number of neighbors |NS(v q )| of node v in the query graph is consistent with the number of neighbors |NS(v q )| of the candidate node v d : d )| = |NS(v
[0163] |NS(v q )| = |NS(v d )|
[0164] The hierarchical topological sequence not only defines the levels of nodes but also contains the connection relationship information between each level. For example, if it is queried that there is a connection relationship between node A in the second level and node B in the third level in the graph, then in the candidate subgraph, there should also be the same connection relationship between the corresponding node A' in the second level and the corresponding node B' in the third level.
[0165] By backtracking and checking the hierarchical topological sequence, the structural consistency between the query graph and the candidate subgraph is further ensured, effectively avoiding incorrect matching caused by inconsistent structures, thereby completing the final global isomorphism verification and determining the accurate subgraph matching result.
[0166] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A data query method based on a graph convolutional network, characterized in that, It includes the following steps: S1. Construct the extraction of heterogeneous data and multi-layer semantic network; S2. Generate the transformation of multi-layer topology network and hierarchical topology sequence; S3. Semantic matching; S4. Subgraph matching based on GCN path embedding; S5. Structure matching and isomorphism verification.
2. The data query method based on graph convolutional network according to claim 1, wherein The specific steps for constructing the extraction of heterogeneous data and multi-layer semantic network in step S1 include the following steps: S11. Extraction and preprocessing of heterogeneous data; S12. Construct a multi-layer semantic network; S13. Node and relationship mapping of the multi-layer semantic network; S14. Vector representation; S15. Strengthening of graph representation based on GCN.
3. The data query method based on a graph convolutional network according to claim 2, wherein The specific steps for strengthening the graph representation based on GCN in step S15 include the following steps: S151. Overview of the GCN model; S152. Data preparation and feature initialization; S153. Convolutional layer and model training; S154. GCN representation of the time series graph.
4. The data query method based on a graph convolutional network according to claim 1, wherein The specific steps for generating the transformation of multi-layer topology network and hierarchical topology sequence in step S2 include the following steps: S21. Transform the multi-layer topology network; S22. Generate a hierarchical topology sequence.
5. The data query method based on a graph convolutional network according to claim 4, wherein The specific steps for generating a hierarchical topology sequence in step S22 include the following steps: S221. Initialization: Take all nodes with in-degree 0 as the output of the first layer, mark their parallel relationships, and at the same time delete these nodes and their associated edges, and update the remaining network; S222. Layer-by-layer recursion: Repeat the above operations. Each time, select a new set of nodes with an in-degree of 0 as the next layer until all nodes are output. The output sequence is represented in the following format: S = C1 / C2 / … / C n where C i is the set of nodes in the i-th layer, and " / " represents the inter-layer separator; S223. Symbol representation rule: Each layer of nodes is wrapped with "{}", and parallel nodes within the layer are separated by ",", for example {S1, c2}; The out-degree relationship of nodes is represented by "[]".
6. The data query method based on a graph convolutional network according to claim 1, wherein The specific steps for the GCN representation of node semantic information in step S3 include the following steps: S31. GCN representation of node semantic information; S32. Edge-level semantic matching; S33. Local consistency verification; S34. Semantic matching output.
7. The data query method based on a graph convolutional network according to claim 1, wherein The specific steps for subgraph matching based on GCN path embedding in step S4 include the following steps: S41. Pruning strategy; S42. Index mechanism; S43. Index-level pruning; S44. Subgraph matching algorithm based on GCN.
8. The data query method based on a graph convolutional network according to claim 7, wherein The specific steps for the pruning strategy in step S41 include the following steps: S411. Path label pruning: For a path p of length l z , use GCN to obtain its node embedding, and then form a path label embedding vector o0(p z ) by concatenating or aggregating node labels. During the pruning process, if the label embedding o0(p z ) of the data graph path p z does not match the label embedding o0(p z ) of the query path o0(p q ), that is, o0(p z ) ≠ o0(p q ), then prune this path from the matching search space; S412. Path domination pruning: By performing specific aggregation operations on the node vectors output by the GCN, maximizing, minimizing, or dominating vector concatenation, the path embedding o(p z ) is obtained. If the embedding o(p q ) corresponding to the query path p q does not dominate (or is not dominated by) the embedding o0(p z ) of the data path p z ), that is, then it is determined that this path is invalid in subgraph matching and is quickly excluded. If the level of the node to which the path belongs conflicts with the level of the corresponding node in the query path, level pruning is directly performed.
9. The data query method based on a graph convolutional network according to claim 1, wherein The specific steps for structure matching and isomorphism verification in step S5 include the following steps: S51. Global isomorphism of nodes and edges; S52. Backtracking of the hierarchical topology sequence.
Citation Information
Patent Citations
Semantic-based data lake query system and method
CN114218400A
Unstructured data query method and device based on PostgreSQL
CN117216216A
Heterogeneous graph neural network node classification method of double-view normal form based on network mode and meta-path
CN118228103A
Electric power data management method based on heterogeneous data resource atlas
CN118535749A
Construction method and device of knowledge base question-answering system, equipment and storage medium
CN119293164A
Cited By
Multi-component object matching method and system based on graph matching
CN121095607A