Table data learning method fusing multi-graph convolution
By introducing graph neural network models in table learning, feature embedding graphs and instance interaction graphs are constructed separately, and combining dual-core convolution and hierarchical pooling modules, the problems of insufficient balance of table data processing in the existing technology are solved, and more efficient information capture and prediction effects are achieved.
Patent Information
- Application Number
- CN202411826165.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-27
AI Technical Summary
The existing table learning methods are complex in balancing the relationship between features and instances, building graph representation, and focusing on a single angle, making it difficult to effectively capture the local and global information of table data.
A tabular learning model based on graph neural network is proposed. By initializing feature embedding graphs and instance interaction graphs from row and column angles of table data, combining graph convolution and graph attention, the dynamically gated hierarchical pooling module reduces graph complexity, and introduces the adaptive fusion module to balance the relationship between features and instances.
By integrating the characteristics and instance relationships of tabular data to improve the accuracy and prediction efficiency of information capture, the experimental results show that they are better than existing methods on multiple public data sets, and the advantages of the model in capturing feature/instance relationships and retaining important information are verified through ablation experiments.
Smart Images

Figure CN120047780A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tabular data processing, and particularly to a tabular data learning method integrating multi-graph convolution. Background Art
[0002] Existing tabular learning methods are complex in balancing the relationship between features and instances and constructing graph representations, and there are also problems such as a single focus angle. Therefore, the present invention proposes a new tabular learning model based on a graph neural network, which initializes a feature embedding graph and an instance interaction graph from the row and column perspectives of tabular data respectively, and integrates local and global information of the data. The model enhances the node embedding representation through a dual-core convolution module combining graph convolution and graph attention, uses a hierarchical pooling module based on dynamic gating to reduce the graph complexity and retain important node difference information, and at the same time introduces an adaptive fusion module to balance the relationship between features and instances and improve the model accuracy. Summary of the Invention
[0003] In view of the above technical problems, the present invention provides a tabular data learning method integrating multi-graph convolution.
[0004] The present invention is implemented by the following technical solutions: A tabular data learning method integrating multi-graph convolution, based on the GTabGNN model architecture, includes the following steps:
[0005] Step S1: Initialize tabular data as graph data suitable for graph neural network processing in the graph representation construction layer of the model, and construct two graph representations, namely a feature embedding graph based on columns and an instance interaction graph based on rows;
[0006] Step S2: Input them in parallel into the graph embedding representation learning layer, and enhance the embedding representation of nodes and retain important node differences through the ACGNN dual-core convolution module and the VG hierarchical pooling module of the model;
[0007] Step S3: Adaptively fuse the two graph representations through the feature fusion module of the model, and use them for the final task through the classification and regression layers.
[0008] Specifically, the graph representation construction layer includes a feature embedding encoder module, a feature embedding graph module, and an instance interaction graph module; the feature embedding encoder module maps the original features to the same and dense d-dimensional vector space, specifically including:
[0009] Let represent a tabular data set containing N data instances, where each instance X i consists of multiple continuous features and discrete features and the total number of features is m; that is And y 1 ∈Y represents the data instance X iLabel;
[0010] Process discrete and continuous features separately, and convert them into graph nodes suitable for GNN processing. The formula is:
[0011]
[0012] where ReLU is the non-linear activation function for continuous embedding, and b is j the j-th feature bias, which is a learnable vector. is a learnable lookup table, |j c | and are the size of the categorical feature and the one-hot encoded representation respectively;
[0013] After passing through the embedding encoder, the sample obtained is:
[0014]
[0015] where k num and k cat are the numbers of continuous features and discrete features respectively.
[0016] Specifically, the feature embedding graph module uses the multi-head attention mechanism Multi-Head Attention to capture the relationships between feature representations, specifically including:
[0017] For the input feature matrix H, obtain the query matrix Q, key matrix K, and value matrix V through linear transformation. The formula is:
[0018]
[0019] where W Q , W K and W V are learnable parameter matrices of the model;
[0020] Calculate the attention weights between any two nodes i and j of the two types of features. The calculation formula is:
[0021]
[0022] where k is the dimension of the key matrix, used for scaling to avoid overly large attention scores;
[0023] According to the attention weights, perform weighted summation on the value matrix V to obtain the output of a single attention head:
[0024] Z i = ∑ j a ij V j ;
[0025] The outputs of multiple heads are concatenated, fused, and linearly transformed to obtain the final feature embedding representation as follows:
[0026] Z = Concat(Z 1 , Z 2 ,..., Z k )W O .
[0027] Specifically, the instance interaction graph module constructs an instance interaction graph using the Euclidean distance to represent the dependence relationship between samples, specifically including:
[0028] The feature representations of individual samples are simply concatenated as the embedding representation of the samples, expressed as:
[0029]
[0030] The samples are regarded as independent nodes, and an initial graph is constructed. The similarity between any two samples i and j is measured using the Euclidean distance, expressed as:
[0031] K(H i , H j ) = exp(-‖H i - H i ‖ 2 );
[0032] where ‖Hi - H j ‖ 2 is the square of the Euclidean distance between two samples, and the value of k(H i , H j ) is between 0 and 1. The closer its value is to 1, the higher the similarity between the two samples; the closer it is to 0, the lower the similarity;
[0033] The value of the adjacency matrix A ij is set as:
[0034]
[0035] ReLU is the activation function; the larger the value of A ij , the higher the similarity between the two samples.
[0036] Specifically, the ACGNN dual - core convolution module includes a GCN module and a GAT module. The GCN module learns the global graph structure information. Based on the topological structure of the graph, it updates the node embedding representation by aggregating the information of neighbor nodes. The calculation of the (l + 1)-th layer is as follows:
[0037]
[0038] where, represents the adjacency matrix, Denotes the degree matrix, Denotes the output of node features at the l-th layer, Denotes the weight matrix at the l-th layer;
[0039] The final output layer is processed using the softmax function and is expressed as:
[0040] Specifically, the GAT module captures the local features of neighboring nodes, adopts a multi-head attention mechanism, learns the information of neighbor nodes through the attention mechanism, and enhances the embedded representation of nodes; the feature vector of node i is expressed as:
[0041] h i ′ = W a (∑ j∈N(i) a ij ·W·h j );
[0042] a ij Is the attention coefficient between nodes i and j, and the representation vectors h i And h j After linear transformation, they are concatenated and then subjected to an inner product operation with the vector a to obtain a ij The calculation formula is: σ is the LeakyReLU activation function;
[0043] The output representing the aggregation of K independent attention heads is:
[0044]
[0045] Finally, multi-class prediction is achieved through the softmax function:
[0046]
[0047] Specifically, the VG hierarchical pooling module performs multi-scale pooling based on a dynamic gating mechanism, which specifically includes:
[0048] The difference index of node i is expressed as:
[0049]
[0050] Among them, Is the k-th order neighbor of node i; the greater the feature difference between the node and its neighbor, the higher the difference index;
[0051] The gating value is dynamically adjusted using the difference index and is expressed as:
[0052]
[0053] Among them, K is the order of aggregated neighbors, and b g are learnable parameters;
[0054] Aggregate nodes and neighbors with the gating value, and the calculation formula is:
[0055]
[0056] where, is the average weight of neighbor nodes;
[0057] Normalize the difference index to represent the scoring coefficient, which is expressed as:
[0058]
[0059] where δ i is the difference index of node i, and max(δ) is the maximum value of the difference indices of all nodes;
[0060] Finally, select and retain the top KN nodes with the highest scoring coefficients:
[0061]
[0062] where, top-k is a function that sorts the scoring vector and returns the indices of the top KH nodes as the pooling result; k ∈ (0, 1] is the pooling rate, indicating the proportion of the retained nodes to the original nodes.
[0063] Specifically, the feature fusion module sets an adaptive fusion strategy and dynamically adjusts the weights of different features through a gating mechanism, which specifically includes:
[0064] Let H 1 and H 2 represent the features of the feature embedding graph and the instance interaction graph. After connecting them, input them into the gating function to generate a gating vector to determine the weight ratio of different features. The calculation formula is:
[0065] g = σ(W g · [H 1 ‖ H 2 );
[0066] where σ is the sigmoid activation function, W g is the learnable weight matrix, and || represents the feature connection operation;
[0067] Adaptively fuse the features according to the weight ratio, and the final output feature is expressed as:
[0068] H fused = g ⊙ H 1 + (1 - g) ⊙ H 2 ;
[0069] Among them, ⊙ represents element-wise multiplication;
[0070] Finally, the fused features are input into the fully connected layer for label prediction:
[0071] y i = softmax(W·H + b).
[0072] The beneficial effects of the present invention are as follows: The present invention proposes the GTabGNN framework, a model based on graph neural network, which integrates the features and instance relationships of tabular data to improve the accuracy of information capture and prediction efficiency. The model enhances the node embedding representation by constructing an instance interaction graph and a feature embedding graph respectively, combining dual-core convolution and hierarchical pooling modules, and balances the instance and feature relationships through an adaptive fusion mechanism. Experiments on 5 public datasets show that GTabGNN outperforms existing methods in multiple performance metrics. Ablation experiments further verify the advantages of the model in capturing feature / instance relationships and retaining important information. This research provides a new perspective for understanding the complex interactions between data and improves the accuracy and efficiency of prediction models. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0074] Figure 1 It is the architecture diagram of the GTabGNN model in the embodiment of the present invention;
[0075] Figure 2 It is the schematic diagram of the multi-scale pooling structure based on the dynamic gating mechanism in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown here can be arranged and designed in various different configurations.
[0077] It should be noted that: Similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0078] The following combines the attachedFigure 1-2 , some embodiments of the present invention will be described in detail. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0079] The present invention proposes a tabular data learning method integrating multi-graph convolution. Based on the GTabGNN model architecture, it includes the following steps:
[0080] Step S1: Initialize the tabular data as graph data suitable for graph neural network processing in the graph representation construction layer of the model, and construct two graph representations: a column-based feature embedding graph and a row-based instance interaction graph.
[0081] Step S2: Input them in parallel into the graph embedding representation learning layer, and enhance the node embedding representation through the ACGNN dual-core convolution module and the VG hierarchical pooling module of the model, and retain important node differences.
[0082] Step S3: Adaptively fuse the two graph representations through the feature fusion module of the model, and use them for the final task through the classification and regression layer.
[0083] In one embodiment, the GTabGNN model architecture is as Figure 1 shown. The model is divided into three modules. First, initialize the tabular data as graph data suitable for graph neural network processing in the graph representation construction layer, and construct two graph representations: a column-based feature embedding graph and a row-based instance interaction graph. Then input them in parallel into the graph embedding representation learning layer, and enhance the node embedding representation through the ACGNN dual-core convolution module and the VG hierarchical pooling module, and retain important node differences. Finally, adaptively fuse the two graph representations through the feature fusion module, and use them for the final task through the classification and regression layer. Its specific design and content are as follows:
[0084] I. Graph Representation Construction Layer
[0085] 1. Feature Embedding Encoder: To facilitate GNN processing, the feature embedding encoder maps the original features to the same and dense d-dimensional vector space. Let represent a tabular data set containing N data instances, where each instance X i consists of multiple continuous features and discrete features , and the total number of features is m. That is And y 1 ∈y represents the label of the data instance X i . In the feature embedding encoder, the discrete and continuous features are processed as follows respectively to convert them into graph nodes suitable for GNN processing:
[0086]
[0087] where ReLU is a non - linear activation function for continuous embedding, and it is b j the j - th feature bias, which is a learnable vector, is a learnable lookup table, |j c | and are the size of the categorical feature and the one - hot encoded representation respectively. The sample representation obtained after passing through the embedding encoder is:
[0088]
[0089] where k num and k cat are the numbers of continuous features and discrete features respectively.
[0090] 2. Feature Embedding Diagram: Use the multi - head attention mechanism Multi - Head Attention to capture the relationships between feature representations. Specifically, for the input feature matrix H, the query matrix Q, the key matrix K, and the value matrix V are obtained through different linear transformations:
[0091]
[0092] where, W Q 、W K and W V are learnable parameter matrices of the model. Then, the attention weights (here referring to two types of features) of any two nodes i and j are calculated through dot - product:
[0093]
[0094] where, k is the dimension of the key matrix, which is used for scaling to avoid overly large attention scores. According to the attention weights, a weighted sum of the value matrix V is performed to obtain the output of a single attention head:
[0095] Z i =∑ j a ij V j ;
[0096] In this embodiment, the multi - head attention mechanism is used, and each head has an independent parameter set. The outputs of multiple heads are concatenated, fused, and linearly transformed to obtain the final feature embedding representation:
[0097] Z=Concat(Z 1 ,Z 2 ,...,Z k )W O ;
[0098] 3. Instance Interaction Graph: Use Euclidean distance to construct an instance interaction graph to represent the dependencies between samples. Concatenate the feature representations of individual samples simply as the embedding representation of the sample, i.e.,
[0099]
[0100] Then regard the samples as independent nodes and perform initial graph construction. Specifically, use Euclidean distance to measure the similarity between any two samples i and j, which is expressed as:
[0101] K(H i ,H j ) = exp(-‖H i - H j ‖ 2 );
[0102] where ‖Hi - H j ‖ 2 is the square of the Euclidean distance between two samples, and the value of k(H i ,H j ) is between 0 and 1. The closer its value is to 1, the higher the similarity between the two samples; the closer it is to 0, the lower the similarity. Further, the value of the adjacency matrix A ij is set to:
[0103]
[0104] ReLU is the activation function. The larger the value of A ij , the higher the similarity between the two samples.
[0105] To learn the global structure information of the table more comprehensively, this embodiment takes the entire dataset as the input of the model. For a large dataset, design mini - batches to meet this requirement. In addition, the nodes in the graph construction process are position - independent, that is, when the order of feature / sample arrangement changes, the prediction result does not change. Similarly, it is also position - independent at other levels of the network.
[0106] II. Graph Embedding Representation Learning Layer
[0107] 1. ACGNN Dual - Core Convolution Module: This embodiment designs an ACGNN dual - core model, combining two modules, GCN and GAT. GCN is responsible for learning the global graph structure information, and GAT captures the local features of neighboring nodes, so as to understand the graph data features and structure more comprehensively.
[0108] The GCN module updates the node embedding representation based on the graph topology by aggregating the information of neighboring nodes. A single-layer GCN only learns the information of first-order neighboring nodes. As the number of layers increases, the receptive field gradually expands. The output of each layer passes through the BatchNormalization and ReLU activation functions to stabilize the training and enhance the non-linear expression ability. The calculation of the (l + 1)-th layer is as follows:
[0109]
[0110] where, represents the adjacency matrix, represents the degree matrix, represents the node feature output of the l-th layer, represents the weight matrix of the l-th layer. Dropout and residual connections are introduced in each layer to regularize and enhance the training stability. Dropout randomly discards some node features to reduce overfitting, and the residual connection ensures that the deep network can effectively learn the signal. The final output layer uses Softmax for processing:
[0111]
[0112] The GAT module learns the information of neighboring nodes through the attention mechanism to enhance the node embedding representation. After attention, the feature vector representation of node i is:
[0113] h i ′ = W a (∑ j∈N(i) a ij ·W·h j );
[0114] a ij is the attention coefficient between nodes i and j. During the calculation, the representation vectors h i and h j of nodes i and j are linearly transformed, then concatenated and subjected to an inner product operation with the vector a:
[0115]
[0116] σ is the LeakyReLU activation function. The GAT network of this module still adopts the multi-head attention mechanism. Therefore, the representation of node i aggregates the outputs of K independent attention heads, and its representation is:
[0117]
[0118] Finally, multi-class prediction is achieved through softmax:
[0119]
[0120] GCN and GAT learn and enhance features in parallel, and then perform feature fusion.
[0121] 2. VG hierarchical pooling module: To reduce the spatial size of the feature map and enhance feature extraction, pooling is required to aggregate the graph representation. Commonly used mean pooling, max pooling, and concatenation pooling usually only focus on the features within the local window and are difficult to capture the structural information of the graph. In addition, these methods lack adaptability and cannot be dynamically adjusted to adapt to the distribution and features of different data. Therefore, this embodiment proposes a multi-scale pooling method VGPooling based on a dynamic gating mechanism to better capture the graph structure information and node differences. Its structure is as Figure 2 shown.
[0122] To capture richer structural information, a multi-scale difference metric is introduced to learn the features of higher-order neighbors. Therefore, the difference metric of node i is expressed as
[0123]
[0124] where is the k-th order neighbor of node i. The greater the feature difference between the node and its neighbors, the higher the difference metric, which can reflect the uniqueness or heterogeneity of the node in the graph structure. Then, this metric is used to dynamically adjust the gating value:
[0125]
[0126] where K is the order of the aggregated neighbors, and b g are learnable parameters. Through the gating mechanism, the model can dynamically adjust the importance of nodes according to the structural information (i.e., difference). Combine the gating value to aggregate the node and its neighbors:
[0127]
[0128] is the average weight of the neighbor nodes. For node screening, first calculate the score coefficient to determine the importance of the node to the network, and then screen according to the score coefficient, only retaining the top KN nodes with the highest scores. This method can not only reduce the complexity of the graph but also improve the quality of feature representation. Then, directly normalize the difference metric to represent the score coefficient, that is
[0129]
[0130] where δ i is the difference metric of node i, and max(m) is the maximum value of the difference metrics of all nodes. Finally, select and retain the top KN nodes with the highest score coefficients:
[0131]
[0132] Here, top-k is a function that sorts the scoring vector and returns the indices of the top KN nodes as the pooling result. k ∈ (0, 1] is the pooling rate, indicating the proportion of nodes to be retained out of the original nodes.
[0133] III. Fusion Strategy
[0134] In this embodiment, an adaptive fusion strategy is proposed in the feature fusion module, and the weights of different features are dynamically adjusted through a gating mechanism. Specifically, let H 1 and H 2 represent the features of the feature embedding graph and the instance interaction graph. After connecting them, they are input into the gating function to generate a gating vector to determine the weight ratio of different features:
[0135] g = σ(W g · [H 1 ‖H 2 );
[0136] where σ is the sigmoid activation function, W g is the learnable weight matrix, and || represents the feature concatenation operation. Then, the features are adaptively fused according to the weight ratio, and the final output feature representation is as follows, where ⊙ is the element-wise multiplication.
[0137] H fused = g ⊙ H 1 + (1 - g) ⊙ H 2 ;
[0138] Finally, the fused features are input into the fully connected layer for label prediction:
[0139] y i = softmax(W · H + b);
[0140] This strategy is also used in the ACGNN module to fuse the output features of the GCN and GAT models.
[0141] For the foregoing embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to this application.
[0142] In the above embodiments, the basic principles, main features and advantages of the present invention are described. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, any modifications and changes made by those skilled in the art that do not deviate from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A tabular data learning method integrating multi-graph convolution, characterized in that: Based on the GTabGNN model architecture, the following steps are included: Step S1: Initialize the tabular data into graph data suitable for graph neural network processing in the graph representation construction layer of the model, and construct two graph representations: a column-based feature embedding graph and a row-based instance interaction graph; Step S2: Its parallel input graph is embedded into the representation learning layer, and the embedded representation of the node is enhanced and the important node differences are retained through the ACGNN dual-core convolution module and VG layer pooling module of the model; Step S3: The two graph representations are adaptively fused through the feature fusion module of the model and used for the final task through the classification and regression layers.
2. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 1, characterized in that: The graph representation construction layer includes a feature embedding encoder module, a feature embedding graph module and an instance interaction graph module; The feature embedding encoder module maps the original features to the same and dense d-dimensional vector space, specifically including: make Represents a tabular dataset containing N data instances, where each instance X i Multiple continuous features and discrete features The total number of features is m; that is And y1∈Y represents the data instance X i Labels; Discrete and continuous features are processed separately and converted into graph nodes suitable for GNN processing. The formula is: Where ReLU is a nonlinear activation function for continuous embedding, which is b j The jth feature deviation is a learnable vector, is a learnable lookup table, |j c | and They are the size of the categorical features and the one-hot encoding representation; After embedding the encoder, the sample is: where k num and k cat are the number of continuous and discrete features, respectively.
3. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 2, characterized in that: The feature embedding graph module uses the Multi-Head Attention mechanism to capture the relationship between feature representations, including: For the input feature matrix H, the query matrix Q, key matrix K and value matrix V are obtained through linear transformation. The formula is: Among them, W Q , W K and W V is the learnable parameter matrix of the model; Calculate the attention weights of any two nodes i and j of the two features, and the calculation formula is: Where k is the dimension of the key matrix, which is used for scaling to avoid the attention score being too large; According to the attention weights, the value matrix V is weighted and summed to obtain the output of a single attention head: From i =∑ j and ij In j ; The outputs of multiple heads are concatenated, fused, and linearly transformed to obtain the final feature embedding representation: Z=Concat(Z 1 ,WITH 2 ,...,WITH k )IN O 。 4. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 3, characterized in that: The instance interaction graph module uses Euclidean distance to construct an instance interaction graph to represent the dependency relationship between samples, specifically including: The feature representation of a single sample is simply concatenated as the embedded representation of the sample, expressed as: The samples are regarded as independent nodes, and the initial graph is constructed. The Euclidean distance is used to measure the similarity between any two samples i and j, which is expressed as: K(H i ,H j )=exp(-‖H i -H j ‖ 2 ); Among them ‖Hi-H j ‖ 2 is the square of the Euclidean distance between two samples, k(H j ,H j ) is between 0 and 1. The closer its value is to 1, the higher the similarity between the two samples is, and the closer it is to 0, the lower the similarity is. Adjacency Matrix A ij The value is set to: ReLU is the activation function; A ij The larger the value of , the higher the similarity between the two samples.
5. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 4, characterized in that: The ACGNN dual-core convolution module includes a GCN module and a GAT module. The GCN module learns the global graph structure information and updates the node embedding representation by aggregating the information of neighboring nodes based on the topological structure of the graph. The calculation of the l+1 layer is as follows: in, represents the adjacency matrix, represents the degree matrix, represents the feature output of the l-th layer node, represents the k-th layer weight matrix; The final output layer is processed using the softmax function and is expressed as:
6. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 5, characterized in that: The GAT module captures the local features of neighborhood nodes and adopts a multi-head attention mechanism to learn the information of neighboring nodes and enhance the embedded representation of nodes. The feature vector of node i is expressed as: h i ′=W a (∑ j∈N(i) a ij ·W·h j ); a ij is the attention coefficient of nodes i and j, and the representation vector h of nodes i and j is i and h j After linear transformation, concatenate and perform inner product operation on vector a to obtain a ij The calculation formula is: σ is the LeakyReLU activation function; The output of aggregating K independent attention heads is expressed as: Finally, multi-classification prediction is achieved through the softmax function:
7. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 6, characterized in that: The VG-level pooling module performs multi-scale pooling based on a dynamic gating mechanism, specifically including: The difference index of node i is expressed as: in, is the kth-order neighbor of node i; the greater the difference between the characteristics of a node and its neighbors, the higher the difference index; The gate value is dynamically adjusted using the difference index, expressed as: Where K is the neighbor order of aggregation, and b g is a learnable parameter; Combine the gated values to aggregate nodes and neighbors. The calculation formula is: in, is the average weight of neighbor nodes; The difference index is normalized to express the score coefficient, which is expressed as: Among them, δ i is the difference index of node i, max(δ) is the maximum value of the difference index of all nodes; Finally, select and retain the nodes with the top KN score coefficients: Among them, top-l is a function that sorts the score vector and returns the node indexes of the top KN as the pooling result; k∈(0,1] is the pooling rate, which means the ratio of the number of retained nodes to the original nodes.
8. The method for learning tabular data by integrating multi-graph convolution as claimed in claim 7, characterized in that: The feature fusion module sets an adaptive fusion strategy and dynamically adjusts the weights of different features through a gating mechanism, specifically including: Let H1 and H2 represent the features of the feature embedding graph and the instance interaction graph, connect them and input them into the gating function to generate a gating vector to determine the weight ratio of different features. The calculation formula is: g=σ(W g ·[H1‖H2]); Where σ is the sigmoid activation function, W g is a learnable weight matrix, || represents the feature connection operation; The features are adaptively fused according to the weight ratio, and the final output feature is expressed as: H fused =g⊙H1+(1-g)⊙H2; Among them, ⊙ is element-by-element multiplication; Finally, the fused features are input into the fully connected layer for label prediction: y i {softmax(W·H+b)。
Citation Information
Cited By
Partition graph and node graph fusion-based graph convolution complex equipment fault diagnosis method
CN120724105A