Active learning method based on enhanced graph convolutional network
By using an enhanced graph convolutional network as a sampler in the field of graph data, combining node feature similarity and multi-layer feature fusion, the cost of obtaining labeled data and data scarcity are solved, and the performance and generalization ability of active learning are improved.
Patent Information
- Application Number
- CN202510133948.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In machine learning, especially in the field of graph data, the cost of obtaining labeled data is high and data is scarce, resulting in limited active learning performance and generalization capabilities.
The enhanced graph convolution network is used as a sampler. By combining node feature similarity and adjacency matrix, the most informative sample is selected for marking, and the node representation ability is improved through multi-layer feature fusion.
Effectively reduce the need for labeled data, improve the performance and generalization capabilities of machine learning models, and ensure that the selected data is more representative.
Smart Images

Figure CN119990243A_ABST
Abstract
Description
Technical Field
[0003] The present invention relates to the field of machine learning, and more specifically to a method for improving active learning performance based on a graph convolutional network. Starting from the characteristics of the data itself, the present invention uses a graph convolutional network as a sampler to replace a target model for data selection, and selects the most informative samples for labeling to minimize the need for labeled data, thereby improving the performance and generalization ability of the machine learning model. Background Art
[0005] Effective analysis and learning of large-scale graph data usually requires a large amount of labeled data to train machine learning models. Due to the scarcity and high cost of labeled data, obtaining labeled data is often a huge challenge. Active learning minimizes the need for labeled data by selecting the most informative samples for labeling, thereby improving model performance and generalization. In the field of graph data, active learning can help optimize the selection strategy of labeled data and select the nodes that are most critical to the learning task for labeling, thereby improving training efficiency and model performance.
[0006] Graph Convolutional Networks (GCNs), as a powerful tool for processing graph-structured data, provide new possibilities for active learning. By combining active learning with graph convolutional networks, the complex relationships and topological structures between data can be better utilized, thereby improving the performance and efficiency of active learning. This combination can not only help the model select more representative nodes for labeling, but also dynamically adjust the sample selection strategy during the learning process to adapt to different data distributions and task characteristics. In the field of machine learning, feature selection is one of the key steps to optimize model performance and reduce the risk of overfitting. However, in high-dimensional data and complex feature spaces, traditional feature selection methods will lose important feature information. Especially in the context of active learning, how to select the most informative features to improve model performance is particularly critical.
[0007] In the current research field, research on feature selection and active learning using graph convolutional networks has attracted widespread attention. In 2016, Kipf et al. introduced a method for semi-supervised classification using graph convolutional networks. GCNs is a deep learning method based on graph-structured data that can effectively use the relationship information between nodes for learning and prediction. In 2020, Chen et al. proposed a graph convolutional network method based on node feature convolution. By introducing node feature convolution in GCNs, the relationship between node features can be better captured and the modeling performance of graph data can be improved. In 2023, Zhang et al. introduced a feature selection method based on graph convolutional networks for processing high-dimensional and low-sample-size data. This method uses GCNs to select the most representative features to improve the performance and generalization ability of the model. Summary of the invention
[0009] The present invention designs an active learning method based on an enhanced graph convolutional network. The method starts from the characteristics of the data itself and uses a graph convolutional network as a sampler instead of a target model to select data. By averaging the feature vectors output by multiple convolutional layers, the multi-level features of the nodes can be better captured. This helps to comprehensively consider the information of different convolutional layers and improve the representation ability of features. And by using the feature similarity of labeled data and unlabeled data, the graph convolutional network predicts whether the input data is labeled data, so that unlabeled data that is more different from labeled data can be selected, so that more feature representations can be learned with the least labeled data, ensuring that the selected data is more representative, and improving the effect and performance of active learning.
[0010] An active learning method based on enhanced graph convolutional networks includes three parts: calculating node feature similarity, multi-layer feature fusion and active learning based on enhanced graph convolutional networks.
[0011] (1) Enhanced graph convolutional network based on node feature similarity
[0012] Node features are vectors that describe node attributes or features in a graph, and are usually used to represent various attribute information of nodes in a graph. Node feature vectors can provide more key information about the node itself, help distinguish feature differences between different nodes, and help the model learn the relationship between nodes in graph data.
[0013] Although the adjacency matrix It can reflect the connection relationship between nodes, but when calculating the node correlation, not only the adjacency matrix is needed, but also the similarity of node features. The node feature vector reflects the attribute information of the node and can help distinguish the feature differences between different nodes, while the adjacency matrix describes the connection relationship between nodes and can reveal the topological structure between nodes.
[0014] Traditional graph convolutional networks only use the adjacency matrix The relationship between the reflection nodes is not considered the similarity of the node features. The present invention proposes a graph convolutional network based on feature correlation, and selects the unlabeled node with the least correlation with the labeled node by combining the feature similarity and adjacency matrix of the labeled nodes and the unlabeled nodes.
[0015] By combining feature similarity and adjacency, the correlation is used to determine the degree of association between labeled nodes and unlabeled nodes, and then the unlabeled nodes that are least correlated with the labeled nodes are selected. This method can mine the potential relationship between nodes in the graph and help select the unlabeled nodes with the largest feature difference from the labeled nodes.
[0016] (2) Multi-layer feature fusion
[0017] In graph convolutional networks (GCNs), node feature aggregation is a key process that aims to integrate the node's own features and neighbor node features to update the node's representation. By continuously aggregating information from neighbor nodes, the node's representation gradually incorporates more information about its surrounding structure and context, thereby improving the node feature representation. The nodes in the entire graph will go through this aggregation process, so that the representation of each node can learn global structural information.
[0018] In the traditional node feature aggregation process, after multiple information transmissions, the node will learn more information about neighboring nodes, resulting in the loss of key feature information of the node itself. This paper proposes a multi-layer feature fusion method, which improves the representation ability of nodes by aggregating node feature information output by different convolutional layers and combining local and global information.
[0019] By averaging the feature vectors of nodes in different convolutional layers, we are actually aggregating the feature information learned by each node in different convolutional layers. This operation enables the model to comprehensively consider the multi-level feature expressions from different convolutional layers, combining local and global information, and improving the model's performance in node feature extraction.
[0020] The advantage of the multi-layer feature fusion method is that by integrating the feature information learned from different convolutional layers, the model can better capture the complex relationships and structural features of graph data. By integrating local and global information, the multi-layer feature fusion method can enhance the model's ability to model node features, thereby improving the model's performance in node classification tasks.
[0021] (3) Active learning based on enhanced graph convolutional networks
[0022] The active learning method based on enhanced graph convolutional networks is a selection strategy based on the characteristics of the nodes themselves. By using graph convolutional networks as samplers, representative nodes can be selected more effectively, so that the model pays more attention to nodes that have an important impact on the model parameter update during training, thereby improving the performance of the model. By predicting based on the feature similarity of the nodes, representative nodes can be better selected and the efficiency of node selection can be improved.
[0023] In this method, the classification layer of the graph convolutional network is a binary classification, in which nodes with features similar to those of labeled data are likely to be predicted as labeled nodes, while the opposite is predicted as unlabeled nodes. The nodes that need to be manually labeled are selected by the confidence of the graph convolutional network model on the nodes, thereby reducing the cost of labeling nodes. Based on the predicted probability, it is judged whether the node has labeled data, and b nodes with the smallest probability of being predicted as unlabeled nodes are selected for manual labeling, and the labeled nodes are added to the set of labeled nodes middle.
[0024] The method of the present invention was tested on three datasets, Cora, CiteSeer and PubMed, and achieved accuracies of 94.47%, 92.86% and 91.51%, respectively, verifying the effectiveness of the proposed method. The experimental results show that the method of using unlabeled data feature information in active learning can improve the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 The active learning accuracy of the method of the present invention with different proportions of labeled data on different data sets;
[0027] Figure 2 Performance comparison of different active learning methods on different datasets. DETAILED DESCRIPTION
[0029] A member reasoning attack defense method based on global loss constraint, characterized by:
[0030] (1) Threshold-based global loss constraint calculation
[0031] (1.1) Cross entropy loss of the model is defined as:
[0032] (1)
[0033] Among them, the batch size is N, the total number of categories is C, The true label, is the predicted label.
[0034] (1.2) In order to implement global loss constraints in the output layer of the model, an indicator function I is introduced to measure the global loss.
[0035] The indicator function I is defined as: when p is true, I[p]=1; otherwise it is -1.
[0036] (1.3) Global loss The definition is as follows:
[0037] (2)
[0038] In the indicator function In, if , the output is 1, otherwise the output is -1. When the loss value of the model is higher than the preset threshold When , the model performs normal gradient descent to minimize the loss; otherwise, the loss is appropriately increased along the direction of the gradient.
[0039] (2) Loss function design based on global loss constraints
[0040] The total loss function of the model is defined as follows:
[0041] (3)
[0042] Among them: α is the weight coefficient of the loss term, balancing the contribution ratio of the loss term to the final total loss.
[0043] Validity Verification
[0044] The validity verification experiment of the present invention is as follows:
[0045] (1) Experimental setup
[0046] The present invention selects the standard dataset MNIST-FASHION which is widely used in the field of membership inference attack. The method of the present invention is implemented in pytorch and runs on a Windows PC with 32G RAM, Intel(R) Core(TM) i5-9300H CPU @ 2.40GHz and NVIDIA GeForce GTX 1650;
[0047] (2) Experimental results and analysis
[0048] Combination Figure 1 Figure 2It can be seen that with the increase of training rounds on both data sets, the classification accuracy continues to improve, while the member reasoning accuracy gradually decreases and approaches the random guessing level (about 50%). Moreover, the member reasoning accuracy on the MNIST-FASHION data set decreases very quickly, and is close to the random guessing level at 30 rounds. This difference reflects that LC-MID performs more efficiently on simple data sets, while complex data sets require longer training to show results. Whether it is a shadow model attack or a threshold attack, the present invention makes the accuracy of the member reasoning attack close to 50%, further proving that the method of the present invention can effectively resist different member reasoning attacks while maintaining the stability of the model utility.
Claims
1. An active learning method based on enhanced graph convolutional network, characterized by: (1) Graph Convolutional Network Based on Node Feature Similarity In graph convolutional networks (GCNs), the adjacency matrix Describe the connection relationship between nodes in the graph; Representation Node and nodes Is there an edge or connection between them? The feature matrix of the label node is , the feature matrix of unlabeled nodes is ,in is the number of labeled nodes, is the number of unlabeled nodes, is the dimension of the feature vector; (1.1) Calculation of feature similarity between nodes Feature similarity between nodes Calculated by the following formula: (1) in, Indicates a labeled node The characteristic vector of Represents an unlabeled node The eigenvector of (1.2) Calculation of feature similarity matrix Feature Similarity Matrix Calculated by the following formula: (2) (1.3) Calculation of adjacency matrix with feature similarity Calculate adjacency matrix with feature similarity : (3) in, , Represents the adjacency relationship between nodes. is the total number of nodes; and is a weight parameter used to balance the influence of feature similarity and adjacency; is the feature similarity matrix; (1.4) Graph Convolutional Network Feature Matrix Propagation The graph convolutional network feature matrix propagation is shown in the following formula: (4) (2) Multi-layer feature fusion (2.1) Calculation of feature vectors for one layer node In the Round, The feature vector calculation of the layer is shown in the following formula: (5) in Representation Node In the Round, The feature vector of the layer; and Representation Node node In the Round, The feature vector of the layer; It is Round, The weight matrix of the layer; Is a node The set of neighbor nodes of Representation Node and nodes relevance; is the activation function; (2.2) Multi-layer feature fusion The process of multi-layer feature fusion is shown in the following formula: (6) in Is a node In the The eigenvector of the wheel, Is a node In the Round, The feature vector of the convolutional layer, Indicates the number of convolutional layers; (2.3) Feature matrix after one round of aggregation No. The feature matrix after round average aggregation is shown in the following formula: (7) (3) Active learning based on enhanced graph convolutional networks By combining feature similarity and adjacency, the degree of association between labeled nodes and unlabeled nodes is determined, and then the unlabeled nodes that are least related to the labeled nodes are selected. This method can mine the potential relationship between nodes in the graph and help select the unlabeled nodes with the largest feature difference from the labeled nodes. (3.1) Node classification loss function calculation The node classification loss function is calculated as follows: (8) in and Is a node and nodes The real label, is a set of labeled nodes, is a set of unlabeled nodes, and Is a node and nodes Whether it is a labeled prediction result; (3.2) Node feature loss function calculation The node feature loss function is calculated as follows: (9) in It is a labeled node In the The aggregate feature vector of the wheel, Is an unlabeled node In the Aggregate feature vector of the wheel; (3.3) Enhanced total loss function of graph convolutional network By combining formula (8) and formula (9), the complete loss function for training the graph convolutional network model can be obtained as follows: (10) in and It is a hyperparameter that determines whether the graph convolutional network can effectively distinguish whether nodes have labels or not; (3.4) Loss function of active learning model The classification layer of the graph convolutional network is a binary classification, in which nodes with features similar to those of labeled data are likely to be predicted as labeled nodes, while the opposite is predicted as unlabeled nodes. The graph convolutional network model uses the confidence of the nodes to select unlabeled data that is as far away from the labeled data as possible and manually label them, thereby reducing the cost of labeling nodes. Based on the predicted probability, determine whether the node has label data, select b nodes with the smallest probability of being unlabeled nodes for manual labeling, and add the labeled nodes to the set of labeled nodes. middle; For classification tasks, active learning models use a collection of labeled nodes. Training is performed; the loss function is shown in formula (11): (11) in, is the number of samples in the training set; is the number of categories; represents the parameters of the active learning model; It is The node The true label value of each category; The model is The sample By minimizing the cross entropy loss function, the prediction results of the active learning model can be made as close to the true label as possible, thereby improving the classification accuracy of the active learning model.