A graph classification method based on fragmented data

By extracting and merging the metastructure models in the fragmented data graph, and using an adaptive graph classification method based on spatiotemporal information, the shortcomings of fragmented data graph classification in the prior art are solved, and more accurate and reliable graph classification results are achieved.

CN113627517BActive Publication Date: 2025-06-10XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110904130.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-06
Publication Date
2025-06-10
Estimated Expiration
2041-08-06

AI Technical Summary

Technical Problem

The existing technology cannot effectively model and classify fragmented data with huge amounts of content and complex content, resulting in uncertain and inaccurate perception results of the target system and inaccurate graph classification results, which affects staff's judgment of the current situation.

Method used

By extracting the metastructure model in the fragmented data graph, we can judge whether communication can be carried out between nodes, and merge the nodes and metastructures that can be merged to build new graph data. Graph classification is performed using an adaptive method based on spatiotemporal information, combining three-layer graph convolutional neural networks and recurrent neural networks to extract features and classify them.

Benefits of technology

Effectively utilize the information attributes and spatiotemporal attributes of fragmented data, reduce interference from redundant information and noise, improve the accuracy and reliability of graph classification, and enhance the ability to judge the current situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113627517B_ABST
    Figure CN113627517B_ABST
Patent Text Reader

Abstract

The present invention discloses a graph classification method based on fragmented data. The method includes: determining whether communication can be carried out between nodes according to the trajectory information and communication link information of the node fragmented data graph, and extracting the meta-structure model in the fragmented data graph; extracting the sub-graph structure from the extracted meta-structure model, and determining whether the meta-structures can be merged; finally, only retaining the nodes and meta-structures in the original graph that have mergeable nodes and meta-structures; the unmerged nodes and meta-structures are isolated nodes and isolated meta-structures, and the isolated nodes and isolated meta-structures are deleted from the original graph data set, and the remaining graph data is the constructed new graph data; the constructed new graph data is classified by using the self-adaptation based on spatio-temporal information. The fragmented data information attributes and spatio-temporal attributes are effectively utilized, and the interference of redundant information and noise is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and computer vision, and particularly relates to a graph classification method based on fragmented data. Background Art

[0002] In the modern information warfare environment, due to the defense and countermeasure measures of the target system, intelligence information presents fragmented and fragmentary characteristics, which affect the perception of the target system and the evaluation effect of combat capabilities, bring difficulties and risks to the analysis of the weak points of the target system, and cause problems such as uncertain and inaccurate real-time target system perception results and unreasonable selection of strike targets. At the same time, problems such as inaccurate perception and evaluation results, large-scale original intelligence analysis, chaotic organizational structure, lack of information association, and difficulty in constructing dynamic scenarios increase the difficulty of data visualization, data governance, and knowledge discovery in the analysis. Under the conditions of the informationized battlefield, due to the defense and countermeasure measures of the target system such as electromagnetic interference and deception, intelligence information presents fragmented and fragmentary characteristics, which affect the perception of the target system and the evaluation effect of combat capabilities, bring difficulties and risks to the analysis of the weak points of the target system, and cause problems such as uncertain and inaccurate real-time target system perception results and unreasonable selection of strike targets.

[0003] Traditional modeling methods cannot model the huge amount and complex content of fragmented data, cannot fully represent and utilize it, resulting in the inability to effectively obtain the characteristics of each node. Moreover, the traditional modeling method for graph classification of fragmented data does not consider the extraction of key nodes in the graph and does not utilize the sub-structure classification information of the graph, resulting in inaccurate graph classification results. It does not fully utilize the information attributes and spatio-temporal attributes of fragmented data, affecting the judgment of the current situation by the staff. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a graph classification method based on fragmented data to solve the problems in the prior art that the perception results of the target system are uncertain and inaccurate, the selection of strike targets is unreasonable, and traditional modeling methods cannot model and fully represent and utilize the huge amount and complex content of fragmented data, resulting in the inability to effectively obtain the characteristics of each node, making the graph classification results inaccurate, and causing inaccurate judgment of the current situation by the staff.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is a graph classification method based on fragmented data, characterized in that the method includes:

[0006] Judge whether communication can be carried out between nodes according to the trajectory information and communication link information of the node fragmented data graph, and extract the meta-structure model in the fragmented data graph;

[0007] Extract the sub-graph structure from the extracted meta-structure model and determine whether the meta-structures can be merged; finally, only retain the nodes and meta-structures in the original graph that can be merged; the unmerged nodes and meta-structures are isolated nodes and isolated meta-structures, and the isolated nodes and isolated meta-structures are deleted from the original graph data set, and the remaining graph data is the constructed new graph data;

[0008] The constructed new graph data is classified using adaption based on spatio-temporal information.

[0009] Furthermore, the fragmented data graph is the fragmented data of some vehicles, and each vehicle is regarded as a node;

[0010] The trajectory information includes the number, category, name, and coordinate information of each vehicle, where the coordinate information includes the time point when the vehicle is captured and the longitude and latitude information; the communication link information represents the communication relationship between vehicles, including the start and end times of communication, the type of communication link, and the maximum communication distance supported by the communication link type;

[0011] The node has information attributes and spatio-temporal attributes, and its model is:

[0012]

[0013] Among them, C is the set of information attributes and spatio-temporal attributes of the node from time 0 to t; is the set of information attributes and spatio-temporal attributes of the node at time T = t; is the information attribute of the fragment at time T = t, and can be described by a series of fragment attributes, denoted as Among them, is the set of information attributes of the fragment at time T = i; φ(Attr ij ) is the information attribute of the jth fragment at time T = i; Attr ij is the information attribute value of the jth fragment at time T = i; is the spatio-temporal attribute of the fragment at time T = t, and can be described by a series of fragment attributes, denoted as Among them, is the set of spatio-temporal attributes of the fragment at time T = i; φ′(Attr ij ) is the spatio-temporal attribute of the jth fragment at time T = i; Attr ij is the spatio-temporal attribute value of the jth fragment at time T = i.

[0014] Furthermore, the specific method for determining whether communication can occur between nodes is: within the start and end times, first determine whether the nodes have the same link type, and calculate the communication distance L 2 (a, b) between nodes using the Euclidean distance, specifically:

[0015]

[0016] Among them, a(x 1 , y 1 ) is the longitude information and latitude information of node a; b(x 2 , y 2 ) is the longitude information and latitude information of node b;

[0017] If the communication types between two nodes are the same and the communication distance is less than the maximum communication distance of this communication type, it means that there is a communication relationship between the two nodes.

[0018] Furthermore, the extraction of the meta-structure model in the fragmented data graph is specifically as follows:

[0019] Extract eight meta-structures of single-node shape, single-sided shape, double-sided star shape, three-sided star shape, triangle, triangle with a tail, double triangle, and quadrilateral from the meta-structures that often appear in the graph structure composed of fragmented data; the number of nodes controlled by each meta-structure is within 4.

[0020] Furthermore, the extraction of the sub-graph structure from the extracted meta-structure model is specifically as follows:

[0021] ① The graph structure G=(O, E) of the meta-structure to be extracted, where O is the node set of the meta-structure graph and E is the set of relationships between nodes in the meta-structure graph; first, extract the nodes O e in the meta-structure and store them in the set set;

[0022] ② Extract the next node O e+1 , and determine whether the relationship between node O e+1 and node O e satisfies the relationship constraint in the extracted meta-structure, that is, whether there is a communication relationship between node O e+1 and node O e . If the relationship between node O e+1 and node O e satisfies the relationship constraint in the extracted meta-structure, then store node O e+1 in the set set; otherwise, re-extract the next node O e+2 and make a judgment;

[0023] ③ Extract all the nodes O f in the meta-structure in turn, and determine whether the relationship between node O f and all the nodes already stored in the set set satisfies the relationship constraint in the extracted meta-structure. If the relationship between node O f and all the nodes already stored in the set set satisfies the relationship constraint in the extracted meta-structure, then store node O f in the set set;

[0024] ④ Determine whether the number of nodes in the set set is the same as the number of nodes in the meta-structure to be extracted. If they are the same, a sub-graph structure of a meta-structure is successfully extracted and saved to the sub-graph structure database of the meta-structure; if not, clear the set set and extract the next node O in the meta-structure e+1 Store it in the set set and return to ① to continue execution.

[0025] Further, the judgment of whether the meta-structures can be merged is as follows: First, perform a matching operation on the sub-graph structures of each meta-structure model. Specifically:[[]]

[0026] Select a random sub-graph structure from the sub-graph structure database of a meta-structure and match it with all sub-graph structures in the sub-graph structure database of another meta-structure. For meta-structure matching, it is necessary to calculate whether the nodes in the structure match. Its model is:[[]]

[0027] L 2 (V(c I ), V(c X )) ≤ th_obj (3)

[0028]

[0029] Among them, c I is the category of node O I ; c X is the category of node O X ; V(c I ) is the word vector corresponding to the category of c I ; V(c X ) is the word vector corresponding to the category of c X ; L 2 (V(c I ), V(c X )) is the Euclidean distance between the word vector V(c I ) corresponding to the category of c I and the word vector V(c X ) corresponding to the category of c X ; th_obj is the threshold of the Euclidean distance between word vectors, and the threshold is 2; A I,p is the p-th attribute of node O I ; V(A I,p ) is the word vector corresponding to the category attribute of A I,p ; A X,q is the q-th attribute of node O X ; V(A X,q ) is the word vector corresponding to the category attribute of A X,q ; Since a node may have 0 or more attributes, assume that node O I has n attributes, and node OX If there are m attributes, the distance between the attributes of two nodes is an n×m matrix; L 2 (V(A I,p ),V(A X,q )) is the Euclidean distance between the word vectors V(A I,p ) and V(A X,q ); For the query node O I in the n×m matrix, the attribute with the smallest distance between the attributes of the query node O X and the node O I is used as the distance between the attributes of the query node O X and the node O I . Finally, it is an n×1 matrix; ave(·) is the average attribute distance between the query node O X and the node O I ; th_attr is the threshold of the attribute distance between two nodes, and the threshold is 2;

[0030] For two meta-structures to match, the following constraints need to be satisfied:

[0031]

[0032] O J 、O X 、O Y are query nodes; E 1 (O I ,O J ) p is the p-th relationship between the query node O I and the node O J ; V(E 1 (O I ,O J ) p ) is the word vector corresponding to the relationship E 1 (O I ,O J ) p ; E 2 (O X ,O Y ) q is the q-th relationship between the query node O X and the node O Y ; V(E 2 (O X ,O Y ) q ) is the word vector corresponding to the relationship E 2 (O X ,O Y ) q ; L 2 (V(E 1 (OI ,O J ) p ),V(E 2 (O X ,O Y ) q ) is the Euclidean distance between the word vector V(E 1 (O I ,O J ) p ) and the word vector V(E 2 (O X ,O Y ) q ) The Euclidean distance between them; there may be zero or more relationships between nodes. Assume there are n relationships between node O I and node O J , and there are m relationships between node O X and node O Y . The distance between their pairwise relationships is an n×m matrix. For each relationship between node O I and node O J , among all the relationships between node O X and node O Y , the relationship closest to it, that is, is regarded as the relationship distance between this relationship and the matching node, obtaining an n×1 matrix. Finally, taking the maximum value of this matrix, that is, is the relationship distance between two pairs of matching nodes. Th_rel represents the threshold of the relationship distance between two pairs of matching nodes, and the threshold is 1.5;

[0033] After matching, judge whether there are the same nodes in the meta-structure. If there are the same nodes in the meta-structure, merge them; if there are no the same nodes, judge whether some nodes in the meta-structure have the same communication type and satisfy the communication distance threshold constraint in terms of distance, then there is a communication relationship between the nodes, and merge these two meta-structures; its model is:

[0034]

[0035] Among them, O I is the I-th node; O J is the J-th node; comm(·) is the set of node communication types, and each node can have multiple communication types; φ is an empty set; L 2 (O I ,O J ) is the Euclidean distance between node O I and node O J , comm k is the node O I and node O JShared communication link types, where two nodes share k identical communication types, and dis(comm k ) is the set of communication distances of all communication types in comm k ; max(dis(comm k )) represents the maximum communication distance among them. If the Euclidean distance between two nodes satisfies the above constraints, the binary structure is merged.

[0036] Furthermore, the specific process of using the constructed new graph data for graph classification based on spatio-temporal information is as follows:

[0037] Input the new graph data in step 2 into a three-layer graph convolutional neural network to extract features

[0038]

[0039] where X′ is the output feature of the graph data extracted by each layer of the graph convolutional neural network; X is the input feature of the new graph data at each layer of the graph convolutional neural network; A is the adjacency matrix of the new graph data; is the matrix after preprocessing the adjacency matrix of the new graph data, and where I is the identity matrix; is the degree matrix; θ is the weight matrix;

[0040] The final output feature X′ of the new graph data extracted by the three-layer graph convolutional neural network is passed through an adaptive pooling layer to complete the feature transfer of adjacent nodes in the new graph data. The adaptive pooling layer uses three pooling strategies, specifically:

[0041] S final =αS 1 +(1 - α)S 2 (8)

[0042] where S final is the node feature information score of the new graph data learned based on structural features; α and 1 - α are the weights of the node feature information for structural topology learning and feature topology learning respectively, and α = 0.6; S 1 is the node feature information score based on structural topology learning, and S 1 =σ(GNN(A,X)), where σ is the non-linear activation function, A is the adjacency matrix of the new graph data, X is the input feature of the new graph data at each layer of the graph convolutional neural network; GNN is the adaptive graph pooling neural network used to capture the structural features of nodes; S 2 is the node feature information based on feature topology learning, and S 2 =σ(MLP(X)), where X is the input feature of the new graph data at each layer of the graph convolutional neural network; MLP is a multi-layer perceptron composed of multiple stacked neural networks;

[0043] The model for updating node feature information by the adaptive graph pooling layer is as follows:

[0044] X final = αX 1 + (1 - α)X 2 (9)

[0045] Among them, X final is the node feature information of the new graph data calculated based on the structural features; α and 1 - α are the weights of the node feature information for structural topology learning and feature topology learning respectively, and α = 0.6; X 1 is the node structure information based on structural topology learning, and X 1 = GNN(A, X′), where A is the adjacency matrix of the new graph data, and X′ is the output feature of the graph data extracted by the graph convolutional neural network of each layer; GNN is to use the adaptive graph pooling neural network to capture the structural features of the nodes; X 2 is the node feature information based on feature-based topology learning, and X 2 = MLP(X′), where X′ is the output feature of the graph data extracted by the graph convolutional neural network of each layer; MLP is a multi-layer perceptron, which is stacked by multiple neural networks;

[0046] Perform feature aggregation on the nodes and the nodes in their first-order neighborhoods, specifically:

[0047]

[0048] Among them, X new is the aggregated node feature, N is the total number of first-order neighborhood nodes of the node, X n is the feature of the nth node in the first-order neighborhood, and f(X n ) is the feature aggregation function of the nth node;

[0049] Retain the nodes in the top 75% of the scores as the key nodes in the new graph; find the communication relationships of the key nodes in the new graph based on the key nodes, form the key structure of the new graph according to the communication relationships of the key nodes, and extract the key sub-structures from the key structure by using the method of extracting sub-graph structures from the meta-structure model;

[0050] Input the node feature information S final of the new graph data learned based on the structural features and the extracted key sub-structures into the recurrent neural network to classify the graph, specifically:

[0051] Use the recurrent neural network for spatio-temporal information fusion. At the beginning, the input layer is the kth key sub-structure Z k at the current time l, then the feature information of the kth key sub-structure is:

[0052]

[0053] Among them, is the sum of the adjacency matrix and the identity matrix of the k-th key sub-structure; X new is the aggregated node feature of the new graph data learned based on the structural features; GNN is to capture the structural features of the key sub-structures using a graph neural network; then the input layer of each layer is the hidden layer information at the previous moment of time l; then the classification category S of the key sub-structure at the current moment l is output through a recurrent neural network l is:

[0054] S l = RNN(Z k ′, H l-1 ) (12)

[0055] Among them, RNN is a recurrent graph neural network, and this network needs to use the output at the previous moment as the hidden layer information of the network at this moment; H l-1 is the hidden layer information at time l-1, that is, the hidden state H l-1 at the previous moment;

[0056] Classifying the new graph data is:

[0057]

[0058] G out is the category of the new graph data at time l; f(·) is to count the number of key sub-structures in each key sub-structure classification category, and return the category with the largest number of key sub-structures as the final classification category of the new graph data.

[0059] The beneficial effects of the present invention are: 1. The embodiments of the present invention use a typical meta-structure model to model the spatial evolution and temporal evolution characteristics of fragmented data, extract meta-structures with evolutionary characteristics, screen out isolated nodes and redundant structures composed of fragmented data, and effectively utilize the information attributes and spatio-temporal attributes of fragmented data.

[0060] 2. When merging meta-structures in the embodiments of the present invention, meta-structure matching is used. By using a word vector generation tool (Word2vec, ConceptNet), the node labels, node attributes, and relationship attributes in the graph structure are represented by word vectors, and similarity measurement is completed in the word vector space. When performing similarity matching based on meta-structures, a "structure-semantic" similarity matching method with attribute similarity and relationship similarity as constraint conditions is established, replacing the traditional method based on nodes and edges as the judgment basis.

[0061] 3. The graph classification method based on spatio-temporal information fragmented data in the embodiments of the present invention classifies graphs with evolution characteristics at different times, uses adaptive graph pooling to extract key nodes and form a key graph structure, extracts key sub-structures according to the meta-structure model, and reduces the interference of redundant information and noise. In addition, a recurrent neural network is used to fuse spatio-temporal information, and then the key sub-structures are classified. Based on the classification categories of the key sub-structures, the classification result of the large graph is calculated. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0063] Figure 1 is a schematic diagram of the overall framework of the graph classification method based on fragmented data;

[0064] Figure 2 is a schematic diagram of eight meta-structures;

[0065] Figure 3 is a schematic diagram of the adaptive pooling layer in the graph classification network. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0067] As Figure 1 shown is a graph classification method based on fragmented data provided by the embodiment, including the following steps:

[0068] Step S1: Construction of the meta-structure of the multi-dimensional fragmented data graph;

[0069] (1a) The fragmented data comes from battlefield entities, mainly the fragmented data of some vehicles, such as airplanes, vehicles, ships, etc. The information of these vehicles shows fragmented and piecemeal characteristics. The available information is generally the trajectory information and communication link information of the vehicles. Each vehicle is regarded as a node, and the input is the trajectory information and communication link information of the node set. The trajectory information of the node includes the number, category, name, and coordinate information of each vehicle, where the coordinate information includes the time point when the vehicle is captured and the longitude and latitude information. In addition, each vehicle has the communication link type it supports. The communication link information represents the communication relationship (edge) between vehicles, including the start and end time of communication, the communication link type, and the maximum communication distance supported by the communication link type; model the trajectory information and communication link information of the vehicle, where each vehicle represents a node, and each node has information attributes and spatio-temporal attributes, specifically as follows:

[0070] The model of fragmented information is:

[0071]

[0072] Among them, C is the set of information attributes and spatio-temporal attributes of the node from 0 to t moments; is the set of information attributes and spatio-temporal attributes of the node at T = t moment; is the information attribute of the fragment at T = t moment, and can be described by a series of fragment attributes, denoted as Among them, is the set of information attributes of the fragment at T = i moment; φ(Attr ij ) is the information attribute of the jth fragment at T = i moment; Attr ij is the information attribute value of the jth fragment at T = i moment; is the spatio-temporal attribute of the fragment at T = t moment, and can be described by a series of fragment attributes, denoted as Among them, is the set of spatio-temporal attributes of the fragment at T = i moment; φ′(Attr ij ) is the spatio-temporal attribute of the jth fragment at T = i moment; Attr ij is the spatio-temporal attribute value of the jth fragment at T = i moment;

[0073] (1b) Judge whether communication can be carried out between two nodes according to the trajectory information and communication link information of the node set;

[0074] Within the start and end time, first judge whether each node has the same link type, and use the Euclidean distance to calculate the communication distance L 2 (a, b), specifically as follows:

[0075]

[0076] Among them, a(x 1 , y 1 ) is the longitude information and latitude information of node a; b(x 2 , y 2 ) is the longitude information and latitude information of node b.

[0077] If the communication types between two nodes are the same and the communication distance is less than the maximum communication distance of this communication type, it indicates that there is a communication relationship (edge) between the two nodes;

[0078] (1c) Extract the fragmented data primitive structure model

[0079] Extract eight primitive structures such as Figure 2 shown in the single-node shape, single-sided shape, double-sided star shape, three-sided star shape, triangle, triangle with a tail, double triangle, and quadrilateral from the primitive structures that often appear in the fragmented data graph. The number of nodes controlled by each primitive structure is within 4. If the number is too large, it will affect the construction of fragmented data and the efficiency of graph data reconstruction based on the primitive structure model.

[0080] Step S2: Graph data reconstruction based on the primitive structure model;

[0081] (2a) Extract the sub-graph structure of the primitive structure model. The specific steps are as follows:

[0082] ① The graph structure G=(O, E) of the primitive structure to be extracted, where O is the node set of the primitive structure graph and E is the set of relationships between nodes in the primitive structure graph; first, extract the nodes O e in the primitive structure and store them in the set set;

[0083] ② Extract the next node O e+1 , and determine whether the relationship between node O e+1 and node O e satisfies the relationship constraints in the extracted primitive structure, that is, whether there is a communication relationship between node O e+1 and node O e . If the relationship between node O e+1 and node O e satisfies the relationship constraints in the extracted primitive structure, then store node O e+1 in the set set; otherwise, re-extract the next node O e+2 and make a judgment;

[0084] ③ Extract all the nodes O f in the primitive structure in turn, and determine whether the relationship between node O f and all the nodes already stored in the set set satisfies the relationship constraints in the extracted primitive structure. If node O fIf the relationships with all the nodes already stored in the set satisfy the relationship constraints in the extraction meta-structure, then node O f is stored in the set;

[0085] ④ Determine whether the number of nodes in the set is the same as the number of nodes in the meta-structure to be extracted. If they are the same, a sub-graph structure of a meta-structure is successfully extracted and saved to the sub-graph structure database of the meta-structure; if not, clear the set and extract the next node O e+1 store it in the set, and return to ① to continue execution;

[0086] (2b) Merge the meta-structure models

[0087] First, perform a matching operation on the sub-graph structures of the two meta-structure models. Specifically:

[0088] Select a random sub-graph structure from the sub-graph structure database of the meta-structure in step (2a) and match it with all the sub-graph structures in the sub-graph structure database of the other meta-structure. For meta-structure matching, it is necessary to calculate whether the nodes in the structure match. The model is:

[0089] L 2 (V(c I ), V(c X )) ≤ th_obj (3)

[0090]

[0091] where c I is the category of node O I ; c X is the category of node O X ; V(c I ) is the word vector corresponding to the category c I ; V(c X ) is the word vector corresponding to the category c X ; L 2 (V(c I ), V(c X )) is the Euclidean distance between the word vector V(c I ) corresponding to the category c I and the word vector V(c X ) corresponding to the category c X ; th_obj is the threshold of the Euclidean distance between word vectors, and the threshold is 2; A I,p is the p-th attribute of node O I ; V(A I,p ) is the word vector corresponding to the category attribute of A I,p ; A X,q is the node O XThe q-th attribute of; V(A X,q ) is the word vector corresponding to the category attribute of A X,q ; Since a node may have zero or more attributes, assume that node O I has n attributes, and node O X has m attributes, then the distance between the attributes of the two nodes is an n×m matrix; L 2 (V(A I,p ), V(A X,q )) is the Euclidean distance between the word vector V(A I,p ) and the word vector V(A X,q ); For the n×m matrix, the attribute with the minimum distance between the attributes of the query node O I and the attributes of node O X is used as the distance between the attributes of node O I and the attributes of node O X , and finally it is an n×1 matrix; ave(·) is the average attribute distance between node O I and node O X ; th_attr is the threshold of the attribute distance between the two nodes, and the threshold is 2.

[0092] For two meta-structures to match, the following constraints need to be satisfied:

[0093]

[0094] O I , O J , O X , O Y are query nodes; E 1 (O I , O J ) p is the p-th relationship between query node O I and node O J ; V(E 1 (O I , O J )) p is the word vector corresponding to the relationship E 1 (O I , O J ); p E 2 (O X , O Y ) q is the q-th relationship between query node O X and node O Y ; V(E 2 (O X , O Y )) q ) is the word vector corresponding to the relationship E2 (O X ,O Y ) q The corresponding word vector; L 2 (V(E 1 (O I ,O J ) p ),V(E 2 (O X ,O Y ) q )) is the word vector V(E 1 (O I ,O J ) p ) and the word vector V(E 2 (O X ,O Y ) q ) The Euclidean distance between them; There may be zero or more relationships between nodes. Assume there are n relationships between node O I and node O J , and there are m relationships between node O X and node O Y . The distance between their pairwise relationships is an n×m matrix. For each relationship between node O I and node O J , take the relationship closest to it among all the relationships between node O X and node O Y , that is, as the relationship distance between this relationship and the matching node, obtaining an n×1 matrix. Finally, find the maximum value of this matrix, that is, is the relationship distance between two pairs of matching nodes. th_rel represents the threshold of the relationship distance between two pairs of matching nodes, and the threshold is 1.5.

[0095] If there are identical nodes in the meta-structure model, merge them;

[0096] If there are no identical nodes, determine whether some nodes in the meta-structure have the same communication type and satisfy the communication distance threshold constraint in terms of distance. If there is a communication relationship between the nodes, then these two meta-structures can also be merged; Its model is:

[0097]

[0098] Among them, O I is the I-th node; O J is the J-th node; comm(·) is the set of node communication types, and each node can have multiple communication types; φ is an empty set; L 2 (O I ,OJ ) is the Euclidean distance between node O I and node O J , comm k is the communication link type shared by node O I and node O J . Among them, there are k identical communication types shared by two nodes. dis(comm k ) is the set of communication distances of all communication types in comm k ; max(dis(comm k )) represents the maximum communication distance among them. If the Euclidean distance between two nodes satisfies the above constraints, the binary structure can be merged.

[0099] After merging the meta-structure model, only the structures with common nodes and matching structures in the original graph are retained. The unmerged nodes and structures are isolated nodes and isolated structures. Based on the original graph data, the isolated nodes and isolated structures are deleted, and the remaining graph data is the newly constructed graph data;

[0100] Step 3: Classify the new graph data using self-adaptation based on spatio-temporal information

[0101] Input the new graph data from Step 2 into a three-layer graph convolutional neural network to extract features

[0102]

[0103] Among them, X′ is the output feature of the graph data extracted by each layer of the graph convolutional neural network; X is the input feature of each layer of the graph convolutional neural network in the new graph data; A is the adjacency matrix of the new graph data; is the matrix after preprocessing the adjacency matrix of the new graph data, and where I is the identity matrix; is the degree matrix; θ is the weight matrix.

[0104] For example Figure 3As shown in the figure, the final output feature X' of the new graph data extracted by the three-layer graph convolutional neural network is passed through the adaptive pooling layer to complete the feature transfer of adjacent nodes in the new graph data. The adaptive pooling layer uses three pooling strategies. The three pooling strategies calculate the importance scores for each node and update the feature information of each node: First, use structure-based topology learning and feature-based topology learning to learn the feature information of the final output feature X' of the new graph data extracted by the three-layer graph convolutional neural network. Among them, for structure-based topology learning, the graph convolutional network GCN is used to complete the core architecture. When designing the structure, the spectral graph method is used to complete the convolution operation of unstructured data (between nodes on the graph structure). On the one hand, it ensures that the local structures of different nodes are comprehensively considered during feature extraction, and at the same time, it takes into account that the convolutional kernel parameters after training can quickly extract feature data related to the graph structure. For feature-based topology learning, the multi-layer perceptron MLP in the neural network is used to complete. When designing the structure, the range of the perceptron input (receptive field) and the training optimization function need to be considered. Then, the feature information obtained from structure-based topology learning and feature-based topology learning is used for structure-feature-based topology learning to obtain more accurate graph feature information. Structure-feature-based topology learning is constructed using the graph attention network GAT, and the optimization function based on graph attention needs to be considered to improve the network performance and training convergence time to complete the screening task of key nodes and core local structures; specifically:

[0105] The model for the adaptive graph pooling layer to extract key node features is:

[0106] S final =αS 1 +(1 - α)S 2 (8)

[0107] Where S final is the score of the node feature information of the new graph data learned based on structural features; α and 1 - α are the weights of the node feature information of structure topology learning and feature topology learning respectively, and α = 0.6; S 1 is the score of the node feature information based on structure topology learning, and S 1 =σ(GNN(A,X)), where σ is the non-linear activation function, A is the adjacency matrix of the new graph data, and X is the input feature of each layer in the graph convolutional neural network of the new graph data; GNN is the adaptive graph pooling neural network used to capture the structural features of nodes; S 2 is the node feature information based on feature-based topology learning, and S 2 =σ(MLP(X)), where X is the input feature of each layer in the graph convolutional neural network of the new graph data; MLP is the multi-layer perceptron, which is stacked by multiple layers of neural networks;

[0108] The model for the adaptive graph pooling layer to update node feature information is:

[0109] X final =αX 1 +(1-α)X 2 (9)

[0110] Among them, X final is the node feature information of the new graph data calculated based on the structural features; α and 1-α are the weights of the node feature information of the structural topology learning and feature topology learning, respectively, and α=0.6; X 1 is the node structure information based on structural topology learning, and X 1 =GNN(A,X′), where A is the adjacency matrix of the new graph data, X′ is the output features of the graph data extracted by the graph convolutional neural network of each layer; GNN uses an adaptive graph pooling neural network to capture the structural features of the nodes; X 2 is the node feature information of feature-based topology learning, and X 2 =MLP(X′), where X′ is the graph data output feature extracted by the graph convolutional neural network at each layer; MLP is a multi-layer perceptron, which is composed of multiple layers of neural networks stacked together;

[0111] The adaptive pooling layer calculates the importance score of each node by combining the structure and feature information scores of the node. The adaptive pooling layer selects some important nodes as the pooling results, discards the nodes with lower scores, and performs feature aggregation on the nodes and their first-order neighboring nodes before discarding the nodes with lower scores. This allows the retained nodes to carry more information and reduce the feature loss caused by discarding nodes. Specifically:

[0112]

[0113] Among them, X new is the node feature after aggregation, N is the total number of first-order neighboring nodes of the node, X n is the feature of the nth node in the first-order neighborhood, f(X n ) is the feature aggregation function of the nth node. Update the features of each node to avoid a large number of feature losses due to discarding nodes. The top 75% of nodes with higher scores are retained as the key nodes in the new graph. Based on the key nodes, the communication relationship of these key nodes can be found in the new graph. According to the communication relationship of the key nodes, the key structure of the new graph can be constructed. From the key structure, the key substructure is extracted using the method of extracting subgraph structure from the meta-structure model;

[0114] The node feature information S of the new graph data learned based on the structural features final And extract the key substructures and input them into the recurrent neural network to classify the graph, specifically:

[0115] Use a recurrent neural network for spatio-temporal information fusion. At the beginning, the input layer is the k-th key sub-structure Z at the current moment l k , then the feature information of the k-th key sub-structure is:

[0116]

[0117] where is the sum of the adjacency matrix and the identity matrix of the k-th key sub-structure; X new is the aggregated node feature of the new graph data learned based on the structural features; GNN is to capture the structural features of the key sub-structures using a graph neural network; then the input layer of each layer is the hidden layer information at the previous moment of moment l; then the classification category S of the key sub-structure at the current moment l is output through the recurrent neural network l is:

[0118] S l = RNN(Z k ′, H l-1 ) (12)

[0119] where RNN is a recurrent graph neural network, and this network needs to use the output of the previous moment as the hidden layer information of the network at this moment; H l-1 is the hidden layer information at moment l-1, that is, the hidden state H l-1 at the previous moment.

[0120] Classify the new graph data as:

[0121]

[0122] G out is the category of the new graph data at moment l; f(·) is to count the number of key sub-structures in each key sub-structure classification category and return the category with the largest number of key sub-structures as the final classification category of the new graph data.

[0123] Each embodiment in this specification is described in a related manner. For the same and similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant part.

[0124] The above description is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A graph classification method based on fragmented data, characterized in that, the method includes: judging whether communication can be carried out between nodes according to the trajectory information and communication link information of the node fragmented data graph, and extracting the meta-structure model in the fragmented data graph; extracting the sub-graph structure from the extracted meta-structure model, and judging whether the meta-structures can be merged; finally, only the nodes and meta-structures with mergable ones in the original graph are retained; the unmerged nodes and meta-structures are isolated nodes and isolated meta-structures, and the isolated nodes and isolated meta-structures are deleted from the original graph data set, and the remaining graph data is the constructed new graph data; the constructed new graph data is classified by using self-adaptation based on spatio-temporal information; the fragmented data graph is the fragmented data of some vehicles, and each vehicle is regarded as a node; the trajectory information includes the number, category, name, and coordinate information of each vehicle, where the coordinate information includes the time point and longitude and latitude information when the captured vehicle appears; the communication link information represents the communication relationship between vehicles, including the start and end time of communication, the communication link type, and the maximum communication distance supported by the communication link type; the node has information attributes and spatio-temporal attributes, and its model is: Among them, C is the set of the information attributes and spatio-temporal attributes of the node from 0 to the t moment; is the set of the information attributes and spatio-temporal attributes of the node at the moment T = t; is the information attribute of the fragment at the moment T = t, and can be described by a series of fragment attributes, denoted as Among them, is the set of the information attributes of the fragment at the moment T = i; φ(Attr ij ) is the information attribute of the j-th fragment at the moment T = i; Attr ij is the information attribute value of the j-th fragment at the moment T = i; is the spatio-temporal attribute of the fragment at the moment T = t, and can be described by a series of fragment attributes, denoted as Among them, is the set of the spatio-temporal attributes of the fragment at the moment T = i; φ′(Attr ij ) is the spatio-temporal attribute of the j-th fragment at the moment T = i; Attr ij is the spatio-temporal attribute value of the j-th fragment at the moment T = i; The specific method for determining whether communication can be carried out between the judgment nodes is as follows: within the start and end times, first determine whether the nodes have the same link type, and calculate the communication distance L between the nodes using the Euclidean distance 2 (a, b), specifically: where a(x 1 , y 1 ) is the longitude and latitude information of node a; b(x 2 , y 2 ) is the longitude and latitude information of node b; if the communication types between two nodes are the same and the communication distance is less than the maximum communication distance of this communication type, it means that there is a communication relationship between the two nodes; the specific method for extracting the meta-structure model in the fragmented data graph is: extracting eight meta-structures of single-node shape, single-sided shape, double-sided star shape, triple-sided star shape, triangle, triangle with tail, double triangle, and quadrilateral from the meta-structures that often appear in the graph structure composed of fragmented data; the number of nodes controlled by each meta-structure is within 4; the specific method for extracting the sub-graph structure from the extracted meta-structure model is: ①The graph structure G=(O, E) of the meta-structure to be extracted, where O is the set of nodes of the meta-graph structure and E is the set of relationships between the nodes of the meta-graph structure; first, extract the nodes O in the meta-structure e and store them in the set set; ②Extract the next node O e+1 , and determine whether the relationship between node O e+1 and node O e satisfies the relationship constraints in the extracted meta-structure, that is, whether there is a communication relationship between node O e+1 and node O e . If there is a communication relationship between node O e+1 and node O e satisfies the relationship constraints in the extracted meta-structure, then store node O e+1 in the set set; otherwise, re-extract the next node O e+2 and make a judgment; ③ Sequentially extract all nodes O in the meta-structure f , and determine whether the relationship between node O f and all the nodes already stored in the set set satisfies the relationship constraints in the extraction meta-structure. If the relationship between node O f and all the nodes already stored in the set set satisfies the relationship constraints in the extraction meta-structure, then store node O f in the set set; ④ Determine whether the number of nodes in the set set is the same as the number of nodes in the meta-structure to be extracted. If they are the same, a sub-graph structure of a meta-structure is successfully extracted and saved to the sub-graph structure database of the meta-structure; if they are not the same, clear the set set and extract the next node O in the meta-structure e+1 Stored in the set set, return to ① and continue to execute; the judgment of whether the meta-structures can be merged is as follows: first, perform a matching operation on the sub-graph structures of each meta-structure model, specifically: selecting a random sub-graph structure in the sub-graph structure database of the meta-structure and matching it with all the sub-graph structures in the sub-graph structure database of another meta-structure. For meta-structure matching, it is necessary to calculate whether the nodes in the structure match, and its model is: L 2 (V(c I ),V(c X ))≤th_obj (3) Among them, c I is the category of node O I ; c X is the category of node O X ; V(c I ) is the word vector corresponding to the category of c I ; V(c X ) is the word vector corresponding to the category of c X ; L 2 (V(c I ), V(c X )) is the Euclidean distance between the word vector V(c I ) corresponding to the category of c I and the word vector V(c X ) corresponding to the category of c X ; th_obj is the threshold of the Euclidean distance between word vectors, and the threshold is 2; A I , p is the p-th attribute of node O I ; V(A I,p ) is the word vector corresponding to the category attribute of A I , p; A X , q is the q-th attribute of node O X ; V(A X,q ) is the word vector corresponding to the category attribute of A X , q; Since a node may have zero or more attributes, assuming that node O I has n attributes and node O X has m attributes, the distance between the attributes of the two nodes is an n×m matrix; L 2 (V(A I,p ), V(A X,q )) is the Euclidean distance between the word vector V(A I,p ) and the word vector V(A X,q ); is the attribute in the n×m matrix that has the smallest distance between the attributes of the query node O I and the attributes of node O X , and is used as the distance between the attributes of node O I and the attributes of node O X , and finally it is an n×1 matrix; ave(·) is the average attribute distance between node O I and node O X ; th_attr is the threshold of the attribute distance between two nodes, and the threshold is 2; for two meta-structures to be able to match, the following constraints need to be satisfied: O I 、O J 、O X 、O Y are query nodes; E 1 (O I ,O J ) p The p-th relationship between query node O I and node O J ; V(E 1 (O I ,O J )p) The relationship E 1 (O I ,O J ) p The corresponding word vector; E 2 (O X ,O Y ) q is query node O X and node O Y ; The q-th relationship between them; V(E 2 (O X ,O Y )q) is the relationship E 2 (O X ,O Y ) q The corresponding word vector; L 2 (V(E 1 (O I ,O J ) p ),V(E 2 (O X ,O Y ) q ) is the Euclidean distance between the word vectors V(E 1 (O I ,O J )p) and the word vector V(E 2 (O X ,O Y )q); there may be zero or more relationships between nodes. Assume there are n relationships between node O I and node O J , and m relationships between node O X and node O Y . The distance between their pairwise relationships is an n×m matrix. For each relationship between node O I and node O J , consider the relationship closest to it among all the relationships between node O X and node O Y , that is, as the relationship distance between this relationship and the matching node, obtaining an n×1 matrix. Finally, find the maximum value of this matrix, that is, is the relationship distance between two pairs of matching nodes. th_rel represents the threshold of the relationship distance between two matching node pairs, and the threshold is 1.5; after the matching, judge whether there are the same nodes in the meta-structure. If there are the same nodes in the meta-structure, merge them; if there are no the same nodes, judge whether some nodes of the meta-structure have the same communication type and satisfy the communication distance threshold constraint in terms of distance, then there is a communication relationship between the nodes, and these two meta-structures are merged; its model is: Among them, O I is the I-th node; O J is the J-th node; comm(·) is the set of node communication types, and each node can have multiple communication types; φ is an empty set; L 2 (O I , O J ) is the Euclidean distance between node O I and node O J ; comm k is the common communication link type between node O I and node O J . Among them, there are k identical communication types shared by two nodes, and dis(comm k ) is the set of communication distances of all communication types in comm k ; max(dis(comm k )) represents the maximum communication distance among them. If the Euclidean distance between two nodes satisfies the above constraints, the binary structure is merged.

2. The graph classification method based on fragmented data according to claim 1, characterized in that, the specific method for the constructed new graph data to be classified by using self-adaptation based on spatio-temporal information is: inputting the new graph data in step 2 into a three-layer graph convolutional neural network to extract features Among them, X′ is the graph data output feature extracted by the graph convolutional neural network for each layer; X is the input feature of the graph convolutional neural network for each layer in the new graph data; A is the adjacency matrix of the new graph data; is the matrix after preprocessing the adjacency matrix of the new graph data, and where I is the identity matrix; is the degree matrix; θ is the weight matrix; The final output feature X' of the new graph data extracted by the three-layer graph convolutional neural network completes the feature transfer of adjacent nodes of the new graph data through the adaptive pooling layer. The adaptive pooling layer uses three pooling strategies, specifically: S final = αS 1 + (1 - α)S 2 (8) Among them, S final is the node feature information score of the new graph data learned based on the structural features; α and 1 - α are the weights of the node feature information for structural topology learning and feature topology learning respectively, and α = 0.6; S 1 is the node feature information score based on structural topology learning, and S 1 = σ(GNN(A, X)), where σ is a non - linear activation function, A is the adjacency matrix of the new graph data, X is the input feature at each layer of the graph convolutional neural network in the new graph data; GNN is to capture the structural features of nodes using an adaptive graph pooling neural network; S 2 is the node feature information based on feature - based topology learning, and S 2 = σ(MLP(X)), where X is the input feature at each layer of the graph convolutional neural network in the new graph data; MLP is a multi - layer perceptron, stacked by multiple neural networks; The model for the adaptive graph pooling layer to update node feature information is: X final = αX 1 + (1 - α)X 2 (9) Among them, X final is the node feature information of the new graph data calculated based on the structural features; α and 1 - α are the weights of the node feature information for structural topology learning and feature topology learning respectively, and α = 0.6; X 1 is the node structure information based on structural topology learning, and X 1 = GNN(A, X′), where A is the adjacency matrix of the new graph data, and X′ is the output feature of the graph data extracted by the graph convolutional neural network at each layer; GNN is to capture the structural features of the nodes using an adaptive graph pooling neural network; X 2 is the node feature information based on feature - based topology learning, and X 2 = MLP(X′), where X′ is the output feature of the graph data extracted by the graph convolutional neural network at each layer; MLP is a multi - layer perceptron, which is stacked by multiple neural networks; Feature aggregation is performed on the nodes and the nodes in their first-order neighborhoods, specifically: Among them, X new is the aggregated node feature, N is the total number of first-order neighborhood nodes of the node, and X n is the feature of the nth node in the first-order neighborhood, and f(X n ) is the feature aggregation function of the nth node; Retain the nodes in the top 75% of the scores as the key nodes in the new graph; based on the key nodes, find the communication relationships of the key nodes in the new graph, form the key structure of the new graph according to the communication relationships of the key nodes, and extract the key sub-structures from the key structure using the method of extracting sub-graph structures from the meta-structure model; The node feature information S of the new graph data learned based on structural features final and the extracted key substructures are input into a recurrent neural network for graph classification, specifically as follows: Use a recurrent neural network for spatio-temporal information fusion. At the beginning, the input layer is the k-th key sub-structure Z at the current moment l k , then the feature information of the k-th key sub-structure is as follows: Among them, is the sum of the adjacency matrix of the k-th key sub-structure and the identity matrix; X new is the aggregated node feature of the new graph data learned based on the structural features; GNN is to capture the structural features of the key sub-structures using a graph neural network; then the input layer of each layer is the hidden layer information at the previous moment of time l; then the classification category S of the key sub-structure at the current moment l is output through a recurrent neural network l is: S l = RNN(Z k ′, H l-1 ) (12) Among them, RNN is a recurrent graph neural network, which needs to use the output of the previous moment as the hidden layer information of the network at this moment; H l-1 is the hidden layer information at time l-1, that is, the hidden state H at the previous moment l-1 ; The classification of the new graph data is: G out is the category of the new map data at time l; f(·) is to count the number of key sub-structures in each classification category of key sub-structures, and return the category with the largest number of key sub-structures as the final classification category of the new map data.

Citation Information

Patent Citations

  • Fragment fracture surface splicing method and system based on thermal kernel feature

    CN110047151A

  • Fragmented map splicing method based on road traffic markings

    CN112258391A