Paper classification method and device based on local enhancement and depolarization comparison
By adopting local enhancement and debias comparison methods in the graph comparison model, the problem of insufficient consideration of node distribution in the existing technology is solved, which significantly improves the model's characterization ability and the accuracy of node classification.
Patent Information
- Application Number
- CN202510443645.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing graph comparison model has insufficient consideration of node distribution, resulting in sampling bias and model performance degradation.
The paper classification method based on local enhancement and debiased contrast is adopted to enhance structure and feature level by constructing local samples and using multivariate Bernoulli distributions, diversity views are generated, and the model is optimized through debiased contrast loss function.
It effectively reduces sampling deviation, improves the representation ability of graph comparison learning, and improves the accuracy of node classification tasks.
Smart Images

Figure CN119961457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of node classification, and in particular to a paper classification method and device based on local enhancement and debiasing comparison. Background Art
[0002] In recent years, self-supervised learning has attracted great attention from researchers. It eliminates the need for expensive annotated data and achieves comparable performance to supervised methods in many tasks. Contrastive learning is one of the main paradigms of self-supervised learning. It constructs multiple enhanced views of the input data through an enhancement strategy, and then maximizes the similarity between positive sample pairs while minimizing the similarity between negative sample pairs, thereby encouraging the model to learn meaningful embedding representations. Benefiting from some pioneering research in the fields of computer vision and natural language processing, contrastive learning has also made substantial progress in the application of graphs.
[0003] The performance of graph contrast models depends largely on the choice of graph enhancement strategies. Although data augmentation is mature enough in other fields, graph enhancement still needs further exploration due to the uniqueness of graph data. Most existing methods involve deleting existing edges and discarding node features according to a certain ratio. Researchers have also tried to combine expert knowledge or high-order relationships (communities) to guide graph enhancement. However, the enhancements in these methods do not consider the distribution of nodes. Exploring the distribution of nodes, edges, and features helps to generate diverse views without changing semantics as much as possible. But how to learn node distribution without label information is a challenge.
[0004] In addition, the sampling strategies adopted by existing node-level graph comparison models mostly regard the embeddings of the same node in different views as positive sample pairs, and the embeddings of other nodes in different views except the node itself as negative sample pairs. However, this approach may cause sampling bias, treating samples with the same node label as negative samples (false negative samples), resulting in nodes with the same label being pushed away in the representation space, ultimately affecting the performance of the model in downstream tasks. Summary of the invention
[0005] The purpose of this application is to propose a paper classification method and device based on local enhancement and debiasing comparison to address the above-mentioned technical problems.
[0006] In a first aspect, the present invention provides a paper classification method based on local enhancement and debiasing comparison, comprising the following steps:
[0007] A paper classification model is constructed and trained to obtain a trained paper classification model, which includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained through the following process: a graph dataset is constructed based on paper classification training data; a local sample is constructed for each node in the graph dataset; each node is enhanced at the structural level and at the feature level by using the local sample of each node and the multivariate Bernoulli distribution, respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; the enhanced adjacency matrix and the enhanced feature matrix are sampled twice to obtain a first enhanced view and a second enhanced view respectively; the feature matrix and the adjacency matrix of the first enhanced view are combined into a matrix. The feature matrix and the adjacency matrix of the first enhanced view and the second enhanced view are respectively input into two weight-sharing graph encoders to obtain the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view. The embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-sharing linear projection layers to obtain the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view; a debiasing contrast loss function is constructed according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view, and the graph encoder is trained based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast;
[0008] Obtain a collection of papers to be classified and construct a corresponding graph dataset, where nodes in the graph dataset represent papers or authors, an edge between two nodes represents a citation relationship between two papers or a cooperation relationship between two authors, and the feature vector of each node in the graph dataset corresponds to an element represented by a bag of words in each paper. Input the feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified into a trained paper classification model, first pass through a graph encoder learned by graph contrast to obtain the embedding of each node, and then pass the embedding of each node through a classifier to obtain a classification result, which includes the academic theme of the paper or the research field of the author.
[0009] Preferably, the graph encoder adopts a graph convolution structure, which is expressed as follows:
[0010] ;
[0011] in, represents the feature matrix, represents the adjacency matrix, is the adjacency matrix with self-loops added, represents the identity matrix, is the degree matrix, Indicates The adjacency matrix of nodes plus self-loops, is a nonlinear activation function, is the weight matrix in the graph convolution structure, represents the function corresponding to the graph encoder, Represents the embedding of each node output by the graph encoder.
[0012] Preferably, a local sample is constructed for each node in the graph dataset, specifically including:
[0013] Get the first-order neighbors of each node through the adjacency matrix;
[0014] Input the feature matrix and adjacency matrix of the graph dataset into the graph encoder to obtain the embedding of each node in the graph dataset;
[0015] Input the embedding of each node in the graph dataset into the k-NN algorithm to obtain the k nearest neighbors of each node;
[0016] The first-order neighbors and k nearest neighbors of each node constitute the local sample of each node.
[0017] Preferably, each node is enhanced at the structural level and the feature level by using the local sample of each node and the multivariate Bernoulli distribution, so as to obtain an enhanced adjacency matrix and an enhanced feature matrix, which specifically include:
[0018] The first modeling data is constructed according to the row vector corresponding to each node in each node and its local sample in the adjacency matrix, and the parameters of the Bernoulli distribution corresponding to each node are calculated according to the first modeling data. The likelihood function of the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data is shown as follows:
[0019] ;
[0020] in, Indicates The likelihood function of the parameters of the Bernoulli distribution corresponding to the nodes, The first modeling data The first The sample The value of the dimension, is the sample size, is the total number of nodes, Indicates The node Dimensional parameters;
[0021] The first modeling data is calculated from the The parameters of the Bernoulli distribution corresponding to the nodes are as follows:
[0022] ;
[0023] in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes;
[0024] Sampling the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data to obtain an enhanced adjacency matrix;
[0025] ;
[0026] in, represents a multivariate Bernoulli distribution, Indicates The enhanced adjacency matrix of nodes is " means to be subject to;
[0027] Constructing second modeling data according to the row vector corresponding to each node in each node and its local sample in the feature matrix;
[0028] In response to the feature vector of the node being binary data, the parameters of the Bernoulli distribution corresponding to each node are calculated according to the second modeling data, and the parameters of the Bernoulli distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula:
[0029] ;
[0030] in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes, Indicates The enhanced feature matrix of nodes;
[0031] In response to the feature vector of the node not being binary data, the parameters of the standard Gaussian distribution corresponding to each node are calculated according to the second modeling data. The likelihood function of the parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data is shown in the following formula:
[0032] ;
[0033] in, represents the first The likelihood function of the parameters of the standard Gaussian distribution corresponding to the nodes, represents the second modeling data, represents the value of the jth sample in the second modeling data, The second modeling data The mean vector of the values of all samples in the nodes, The second modeling data The covariance matrix of the values of all samples in nodes, represents the probability density function;
[0034] The parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula:
[0035] ;
[0036] in, represents a multivariate standard Gaussian distribution.
[0037] Preferably, the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-shared linear projection layers to obtain a node representation of each node in the first enhanced view and a node representation of each node in the second enhanced view, specifically including:
[0038] The node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view are calculated using the following formula:
[0039] ;
[0040] ;
[0041] in, represents a linear projection layer, represents the embedding of each node in the first augmented view, represents the embedding of each node in the second augmented view, a node representation representing each node in the first enhanced view, a node representation representing each node in the second enhanced view, Represents the weight matrix in the linear projection layer.
[0042] Preferably, constructing a debiasing contrast loss function according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view specifically includes:
[0043] The nodes in the local sample of each node are regarded as false negative samples and penalized by assigning weights as shown in the following formula:
[0044] ;
[0045] in, Indicates The local samples corresponding to the nodes are Indicates the first enhanced view Node representation of nodes or the second enhanced view Node representation of nodes , Indicates the first enhanced view or the second enhanced view The weight of each node;
[0046] The node representation of the same node in the first enhanced view and the node representation of the second enhanced view are regarded as positive sample pairs, and the node representation of other nodes in the first enhanced view and the node representation of the second enhanced view except the own node are regarded as negative sample pairs;
[0047] The debiasing loss function for the positive sample pair is:
[0048] ;
[0049] in, Indicates the first enhanced view The node representation of nodes is Indicates the first The node representation of nodes is , , Indicates the first The node representation of nodes is represents the temperature coefficient, represents the cosine similarity function;
[0050] The debiasing loss function for the positive sample pair is:
[0051] ;
[0052] in, , , Indicates the first enhanced view Node representation of nodes;
[0053] Debiasing Contrastive Loss Function It is expressed as:
[0054] ;
[0055] in, Indicates the total number of nodes.
[0056] In a second aspect, the present invention provides a paper classification device based on local enhancement and debiasing comparison, comprising:
[0057] The model building module is configured to build and train a paper classification model to obtain a trained paper classification model, wherein the trained paper classification model includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained by the following process: constructing a graph dataset based on paper classification training data; constructing a local sample for each node in the graph dataset; using the local sample of each node and the multivariate Bernoulli distribution to enhance each node at the structural level and the feature level, respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; sampling the enhanced adjacency matrix and the enhanced feature matrix twice to obtain a first enhanced view and a second enhanced view, respectively; and combining the features of the first enhanced view with the features of the first enhanced view. The matrix and the adjacency matrix of the first enhanced view and the feature matrix and the adjacency matrix of the second enhanced view are respectively input into two weight-sharing graph encoders to obtain the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view. The embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-sharing linear projection layers to obtain the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view; a debiasing contrast loss function is constructed according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view, and the graph encoder is trained based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast;
[0058] The classification module is configured to obtain the collection of papers to be classified and construct a corresponding graph dataset, wherein the nodes in the graph dataset represent papers or authors, the edge between two nodes represents the citation relationship between the two papers or the cooperation relationship between the two authors, and the feature vector of each node in the graph dataset corresponds to the element represented by the bag of words in each paper. The feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified are input into the trained paper classification model, and the embedding of each node is obtained by the graph encoder learned by graph contrast. The embedding of each node is passed through the classifier to obtain the classification result, which includes the academic theme of the paper or the research field of the author.
[0059] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0060] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0061] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.
[0062] Compared with the prior art, the present invention has the following beneficial effects:
[0063] (1) The paper classification method based on local enhancement and debiasing contrast proposed in the present invention constructs a local sample for each node in the graph dataset, estimates the distribution of the local sample using a statistical method, samples and generates a first enhanced view and a second enhanced view from the obtained distribution, sends the first enhanced view and the second enhanced view to a weight-shared graph encoder to learn the node representation under different views, generates positive and negative sample pairs, and uses the nodes used to construct the local samples as false negative samples of their corresponding nodes for weighted penalty to construct a debiasing contrast loss function; the model is continuously optimized based on the debiasing contrast loss function to encourage the model to learn more effective representations.
[0064] (2) The paper classification method based on local enhancement and debiased contrast proposed in the present invention adopts an enhancement strategy consisting of structure-level enhancement and feature-level enhancement to provide high-quality views for comparison in the image comparison learning process, and reduces sampling bias through weighting of false negative samples, thereby significantly improving the representation ability of image comparison learning.
[0065] (3) The paper classification method based on local enhancement and debiasing comparison proposed in this paper can design new enhancement strategies using local samples and reduce sampling bias by penalizing false negative samples, thereby improving the accuracy of node classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0067] Figure 1 A schematic diagram of a process flow of a paper classification method based on local enhancement and debiasing comparison according to an embodiment of the present application;
[0068] Figure 2 A schematic diagram of a framework of a paper classification method based on local enhancement and debiasing comparison according to an embodiment of the present application;
[0069] Figure 3 A schematic diagram of a paper classification device based on local enhancement and debiasing comparison according to an embodiment of the present application;
[0070] Figure 4A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0071] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0072] Figure 1 A paper classification method based on local enhancement and debiasing comparison provided by an embodiment of the present application is shown, comprising the following steps:
[0073] S1, construct and train a paper classification model to obtain a trained paper classification model, the trained paper classification model includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained by the following process: construct a graph dataset based on paper classification training data; construct a local sample for each node in the graph dataset; use the local sample of each node and the multivariate Bernoulli distribution to enhance each node at the structure level and the feature level respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; sample the enhanced adjacency matrix and the enhanced feature matrix twice to obtain a first enhanced view and a second enhanced view respectively; the feature matrix of the first enhanced view and the adjacency matrix of the enhanced feature matrix are combined into a matrix; the enhanced adjacency matrix and the enhanced feature matrix are sampled ... The adjacency matrix and the feature matrix and the adjacency matrix of the second enhanced view are respectively input into two weight-sharing graph encoders to obtain the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view. The embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-sharing linear projection layers to obtain the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view. A debiasing contrast loss function is constructed according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view. The graph encoder is trained based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast.
[0074] In a specific embodiment, the graph encoder adopts a graph convolution structure, which is expressed as follows:
[0075] ;
[0076] in, represents the feature matrix, represents the adjacency matrix, is the adjacency matrix with self-loops added, represents the identity matrix, is the degree matrix, Indicates The adjacency matrix of nodes plus self-loops, is a nonlinear activation function, is the weight matrix in the graph convolution structure, Represents the function corresponding to the graph encoder, Represents the embedding of each node output by the graph encoder.
[0077] Specifically, the graph encoder used in the embodiments of the present application adopts a graph convolution structure. The input of the graph convolution structure is the feature matrix and the adjacency matrix in the graph dataset, and the output is the embedding of each node in the graph dataset.
[0078] In a specific embodiment, constructing a local sample for each node in the graph dataset specifically includes:
[0079] Get the first-order neighbors of each node through the adjacency matrix;
[0080] Input the feature matrix and adjacency matrix of the graph dataset into the graph encoder to obtain the embedding of each node in the graph dataset;
[0081] Input the embedding of each node in the graph dataset into the k-NN algorithm to obtain the k nearest neighbors of each node;
[0082] The first-order neighbors and k nearest neighbors of each node constitute the local sample of each node.
[0083] Specifically, obtain a graph dataset. Define a graph dataset as ,in, represents the feature matrix, represents the adjacency matrix, Represents the nodes in the graph dataset. Each node has a feature vector , Represents the dimension of the feature vector. The feature matrix is composed of the feature vectors of all nodes in the graph dataset. The adjacency matrix represents the node and Is there an edge between them? If so, the corresponding element in the adjacency matrix During the training process of the graph encoder, the graph dataset is divided into training set, validation set and test set.
[0084] Based on the original graph data, a local sample is constructed for each node in the graph dataset. The local sample of a node consists of its first-order neighbors and k nearest neighbors. The first-order neighbors can be directly obtained from the neighbor matrix, and the k nearest neighbors are obtained by performing the k-nearest neighbor algorithm (k-NN) on the embedding obtained by inputting the feature matrix and adjacency matrix in the graph dataset into the graph encoder.
[0085] In a specific embodiment, each node is enhanced at the structural level and the feature level by using the local sample of each node and the multivariate Bernoulli distribution, so as to obtain an enhanced adjacency matrix and an enhanced feature matrix, which specifically include:
[0086] The first modeling data is constructed according to the row vector corresponding to each node in each node and its local sample in the adjacency matrix, and the parameters of the Bernoulli distribution corresponding to each node are calculated according to the first modeling data. The likelihood function of the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data is shown as follows:
[0087] ;
[0088] in, Indicates The likelihood function of the parameters of the Bernoulli distribution corresponding to the nodes, The first modeling data The first The sample The value of the dimension, is the sample size, is the total number of nodes, Indicates The node Dimensional parameters;
[0089] The first modeling data is calculated from the The parameters of the Bernoulli distribution corresponding to the nodes are as follows:
[0090] ;
[0091] in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes;
[0092] Sampling the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data to obtain an enhanced adjacency matrix;
[0093] ;
[0094] in, represents a multivariate Bernoulli distribution, Indicates The enhanced adjacency matrix of nodes is " means to be subject to;
[0095] Constructing second modeling data according to the row vector corresponding to each node in each node and its local sample in the feature matrix;
[0096] In response to the feature vector of the node being binary data, the parameters of the Bernoulli distribution corresponding to each node are calculated according to the second modeling data, and the parameters of the Bernoulli distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula:
[0097] ;
[0098] in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes, Indicates The enhanced feature matrix of nodes;
[0099] In response to the feature vector of the node not being binary data, the parameters of the standard Gaussian distribution corresponding to each node are calculated according to the second modeling data. The likelihood function of the parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data is shown in the following formula:
[0100] ;
[0101] in, represents the first The likelihood function of the parameters of the standard Gaussian distribution corresponding to the nodes, represents the second modeling data, represents the value of the jth sample in the second modeling data, The second modeling data The mean vector of the values of all samples in the nodes, The second modeling data The covariance matrix of the values of all samples in nodes, represents the probability density function;
[0102] The parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula:
[0103] ;
[0104] in, represents a multivariate standard Gaussian distribution.
[0105] Specifically, the performance of the graph encoder after graph contrast learning depends largely on the choice of graph enhancement strategy. Based on the unique status of local samples in graph datasets, a new graph enhancement strategy is designed using local samples. The graph enhancement strategy is: using local samples to design enhancement methods at both the structure level and the feature level, the first enhanced view ViewU and the second enhanced view View V are jointly generated by the enhancements at both the structure level and the feature level. The difference between the first enhanced view View U and the second enhanced view View V is that both are obtained by sampling twice based on the enhancement results at both the structure level and the feature level, and the different results of the two samplings result in the existence of two enhanced views for graph contrast learning.
[0106] (1) Structural-level enhancement: The topological structure in a graph dataset is usually represented by an adjacency matrix, which contains only binary data. Therefore, for structural-level enhancement, the embodiments of the present application use multivariate Bernoulli distribution to learn data distribution.
[0107] Constructing the first modeling data based on local samples . First modeling data It is the adjacency matrix between each node and each node in the corresponding local sample. The dimensions of these row vectors are , which is the same as the number of nodes in the graph dataset. The parameters of the multivariate Bernoulli distribution can be calculated. The parameters of the multivariate Bernoulli distribution are the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data. After obtaining the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data, directly sample to obtain the enhanced adjacency matrix .
[0108] (2) Feature-level enhancement: constructing the second modeling data based on local samples , the second modeling data is the feature matrix of each node and each node in the local sample corresponding to it. The dimension of each row vector is d. For the case where the feature vector of the node is also binary data, the embodiment of the present application also uses the multivariate Bernoulli distribution to learn the feature distribution, similar to the enhancement at the structural level, to calculate the enhanced feature matrix. For the case where the feature vector of the node is not binary data, the embodiment of the present application uses the multivariate standard Gaussian distribution to learn the feature distribution. The likelihood function of the parameters of the standard Gaussian distribution corresponding to each node is calculated from the second modeling data, and the parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data are solved to finally generate the enhanced feature matrix .
[0109] In a specific embodiment, the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-shared linear projection layers to obtain a node representation of each node in the first enhanced view and a node representation of each node in the second enhanced view, specifically including:
[0110] The node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view are calculated using the following formula:
[0111] ;
[0112] ;
[0113] in, represents a linear projection layer, represents the embedding of each node in the first augmented view, represents the embedding of each node in the second augmented view, represents the node representation of each node, represents the node representation of each node, Represents the weight matrix in the linear projection layer.
[0114] Specifically, refer to Figure 2 , the generated first enhanced view View U and the first View V are fed into two weight-shared graph encoders to learn node representations under different views and .Will and After a weight-sharing linear projection layer, we get and .according to and Construct a debiasing loss function and finally obtain a debiasing contrast loss function.
[0115] In a specific embodiment, constructing a debiasing contrast loss function according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view specifically includes:
[0116] The nodes in the local sample of each node are regarded as false negative samples and penalized by assigning weights as shown in the following formula:
[0117] ;
[0118] in, Indicates The local samples corresponding to the nodes are Indicates the first enhanced view Node representation of nodes or the second enhanced view Node representation of nodes , Indicates the first enhanced view or the second enhanced view The weight of each node;
[0119] The node representation of the same node in the first enhanced view and the node representation of the second enhanced view are regarded as positive sample pairs, and the node representation of other nodes in the first enhanced view and the node representation of the second enhanced view except the own node are regarded as negative sample pairs;
[0120] The debiasing loss function for the positive sample pair is:
[0121] ;
[0122] in, Indicates the first enhanced view The node representation of nodes is Indicates the first The node representation of nodes is , , Indicates the first The node representation of nodes is represents the temperature coefficient, represents the cosine similarity function;
[0123] The debiasing loss function for the positive sample pair is:
[0124] ;
[0125] in, , , Indicates the first enhanced view Node representation of nodes;
[0126] Debiasing Contrastive Loss Function It is expressed as:
[0127] ;
[0128] in, Indicates the total number of nodes.
[0129] Specifically, the representations of the same node in different views are regarded as positive sample pairs, while the representations of other nodes in different views except the node itself are regarded as negative sample pairs. The prior assumption in the enhancement strategy is that each node belongs to the same distribution as its corresponding local sample, so the nodes in the local sample of each node are directly regarded as false negative samples and penalized by assigning them a weight of 0. The debiasing loss function of each positive sample pair is calculated based on the weighted penalty. The debiasing loss function of the two positive sample pairs is calculated separately. and , and define the overall goal of maximization as the average of all positive sample pairs, thus constructing the debiased contrast loss function . Through the Adam gradient descent algorithm An optimization is performed to update the parameters of the graph encoder to obtain a graph encoder learned by graph contrast, and the graph encoder learned by graph contrast is used for the node classification task.
[0130] The node classification task in the embodiment of the present application is taken as an example of the academic theme classification of the paper or the research field classification of the paper author. In the academic theme classification task of the paper or the research field classification task of the paper author, the graph encoder learned by graph contrast is combined with the classifier to construct a paper classification model, and the paper classification model is trained to obtain a trained paper classification model. In the academic theme classification task of the paper or the research field classification task of the paper author, the classifier includes a logistic regression classifier and a softmax function layer. Specifically construct a graph data set according to the definition of nodes and edges, mark the labels corresponding to the nodes, and train the paper classification models in different classification tasks to meet the classification requirements of the academic theme classification task of the paper or the research field classification task of the paper author.
[0131] S2, obtain the collection of papers to be classified and construct the corresponding graph dataset, where the nodes in the graph dataset represent papers or authors; the edge between two nodes represents the citation relationship between two papers or the cooperation relationship between two authors, and the feature vector of each node in the graph dataset corresponds to the element represented by the bag of words in each paper. The feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified are input into the trained paper classification model, first through the graph encoder learned by graph contrast, to obtain the embedding of each node, and each node embedding passes through the classifier to obtain the classification result, which includes the academic theme of the paper or the research field of the author.
[0132] Specifically, the trained paper classification model is deployed. During the inference process, the feature matrix X and the adjacency matrix in the graph dataset corresponding to the paper set to be classified are firstly As input, we get the embedding through the graph encoder learned by graph contrast :
[0133] ;
[0134] Embed It is fed into a simple logistic regression classifier, first passing through a linear combination layer to obtain the node representation:
[0135] ;
[0136] in, and b are learnable parameters, and T represents the transposed matrix. The node representation is then converted into the predicted probability of each category through the softmax function layer, and the category with the highest probability is selected as the predicted field of the paper to be classified.
[0137] The technical effects of the present invention are further illustrated by means of specific embodiments below.
[0138] For comprehensive comparison, we evaluate the node classification performance on five open graph datasets, including Cora, CiteSeer, Wiki-CS, Coauthor-CS, and Coauthor-Physics, which are derived from real networks in different fields.
[0139] The method of the present invention is compared with classic unsupervised models including DeepWalk and Node2vec. In addition, the method of the present invention is also compared with excellent self-supervised models such as GAE, VGAE, DGI, MVGRL, GRACE, GCA, CCA-SSG, BGRL, COSTA and CSGCL. The method of the present invention is also compared with the supervised method GCN.
[0140] In the experimental case, each model is first trained in an unsupervised manner, and then the node representations output by the graph encoder are fed into a simple logistic regression classifier. A common split is used for the Wiki-CS dataset, and the remaining four datasets are randomly divided into 10%, 10%, and 80% for training, validation, and testing, respectively. Since the data partitions are mostly random, they are run 20 times on each dataset, and the average performance is used as the result.
[0141] Table 1 The number of nodes, edges, feature dimensions, and categories of the dataset.
[0142]
[0143] Table 2 Performance comparison of node classification tasks (average accuracy (%) ± standard deviation).
[0144]
[0145] Table 2 shows the performance of the node classification task of the proposed method on five datasets. It can be seen from Table 2 that compared with all baseline methods, the paper classification model proposed in the present invention achieves the best performance. It is worth noting that on the Cora dataset, compared with the second best baseline, the paper classification model proposed in the present invention has a significant improvement of nearly 3%. These findings provide convincing evidence that the proposed method is an effective framework that can effectively utilize local information.
[0146] Further references Figure 3 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a paper classification device based on local enhancement and debiasing comparison. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0147] The present application embodiment provides a paper classification device based on local enhancement and debiasing comparison, including:
[0148] The model building module 1 is configured to build and train a paper classification model to obtain a trained paper classification model, wherein the trained paper classification model includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained by the following process: constructing a graph dataset based on paper classification training data; constructing a local sample for each node in the graph dataset; using the local sample of each node and the multivariate Bernoulli distribution to perform structural enhancement and feature enhancement on each node respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; sampling the enhanced adjacency matrix and the enhanced feature matrix twice, to obtain a first enhanced view and a second enhanced view respectively; combining the features of the first enhanced view The matrix and the adjacency matrix of the first enhanced view and the feature matrix and the adjacency matrix of the second enhanced view are respectively input into two weight-sharing graph encoders to obtain the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view. The embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-sharing linear projection layers to obtain the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view; a debiasing contrast loss function is constructed according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view, and the graph encoder is trained based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast;
[0149] The classification module 2 is configured to obtain the collection of papers to be classified and construct a corresponding graph dataset, wherein the nodes in the graph dataset represent papers or authors, the edge between two nodes represents the citation relationship between the two papers or the cooperation relationship between the two authors, and the feature vector of each node in the graph dataset corresponds to the element represented by the bag of words in each paper. The feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified are input into the trained paper classification model, firstly passed through the graph encoder through graph contrast learning to obtain the embedding of each node, and each node embedding passes through the classifier to obtain the classification result, which includes the academic theme of the paper or the research field of the author.
[0150] Figure 4 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 4 As shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer-executable instructions; the processor 401 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.
[0151] Optionally, the memory 402 may be independent or integrated with the processor 401 .
[0152] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401 .
[0153] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 401 executes the computer execution instructions, the above method is implemented.
[0154] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.
[0155] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0156] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.
[0157] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0158] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 401 to perform some steps of the methods of various embodiments of the present application.
[0159] It should be understood that the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 401 may be any conventional processor 401, etc. The steps of the method disclosed in the invention may be directly embodied in the hardware processor 401 for execution, or may be executed by a combination of hardware and software modules in the processor 401.
[0160] The memory 402 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.
[0161] The bus 403 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 403 in the drawings of the present application is not limited to only one bus 403 or one type of bus 403.
[0162] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0163] An exemplary storage medium is coupled to the processor 401, so that the processor 401 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0164] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A paper classification method based on local enhancement and debiasing comparison, characterized in that: The following steps are involved: A paper classification model is constructed and trained to obtain a trained paper classification model, wherein the trained paper classification model includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained by the following process: a graph data set is constructed based on paper classification training data; Constructing a local sample for each node in the graph data set; using the local sample of each node and the multivariate Bernoulli distribution to perform structural enhancement and feature enhancement on each node, respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; Sampling the enhanced adjacency matrix and the enhanced feature matrix twice to obtain a first enhanced view and a second enhanced view respectively; Inputting the feature matrix and adjacency matrix of the first enhanced view and the feature matrix and adjacency matrix of the second enhanced view into two weight-sharing graph encoders respectively, obtaining the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view, respectively inputting the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view into two weight-sharing linear projection layers, obtaining the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view; Constructing a debiasing contrast loss function according to a node representation of each node in the first enhanced view and a node representation of each node in the second enhanced view, and training the graph encoder based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast; A collection of papers to be classified is obtained and a corresponding graph dataset is constructed, wherein a node in the graph dataset represents a paper or an author, an edge between two nodes represents a citation relationship between two papers or a cooperation relationship between two authors, and a feature vector of each node in the graph dataset corresponds to an element represented by a bag of words in each paper. The feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified are input into the trained paper classification model, firstly passed through the graph encoder learned by graph contrast learning to obtain an embedding corresponding to each node, and the embedding of each node is passed through the classifier to obtain a classification result, and the classification result includes the academic theme of the paper or the research field of the author.
2. The paper classification method based on local enhancement and debiasing contrast according to claim 1 is characterized in that: The graph encoder adopts a graph convolution structure, which is expressed as follows: ; in, represents the feature matrix, represents the adjacency matrix, is the adjacency matrix with self-loops added, represents the identity matrix, is the degree matrix, Indicates The adjacency matrix of nodes plus self-loops, is a nonlinear activation function, is the weight matrix in the graph convolution structure, represents the function corresponding to the graph encoder, represents the embedding of each node output by the graph encoder.
3. The paper classification method based on local enhancement and debiasing contrast according to claim 1 is characterized in that: Constructing a local sample for each node in the graph dataset, specifically including: Obtaining the first-order neighbors of each node through the adjacency matrix; Inputting the feature matrix and the adjacency matrix of the graph dataset into the graph encoder to obtain the embedding of each node in the graph dataset; Input the embedding of each node in the graph dataset into the k-NN algorithm to obtain the k nearest neighbors of each node; The first-order neighbors and k nearest neighbors of each node constitute the local sample of each node.
4. The paper classification method based on local enhancement and debiasing contrast according to claim 1 is characterized in that: The local samples of each node and the multivariate Bernoulli distribution are used to enhance the structure level and feature level of each node, respectively, to obtain the enhanced adjacency matrix and enhanced feature matrix, which specifically include: First modeling data is constructed according to the row vector corresponding to each node in each node and its local sample in the adjacency matrix, and the parameters of the Bernoulli distribution corresponding to each node are calculated according to the first modeling data. The likelihood function of the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data is shown in the following formula: ; in, Indicates The likelihood function of the parameters of the Bernoulli distribution corresponding to the nodes, The first modeling data is represented by The first The sample The value of the dimension, is the sample size, is the total number of nodes, Indicates The node Dimensional parameters; The first modeling data is calculated from the first The parameters of the Bernoulli distribution corresponding to the nodes are as follows: ; in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes; Sampling the parameters of the Bernoulli distribution corresponding to each node calculated by the first modeling data to obtain an enhanced adjacency matrix; ; in, represents a multivariate Bernoulli distribution, Indicates The enhanced adjacency matrix of nodes, " " means to be subject to; Constructing second modeling data according to the row vector corresponding to each node in each node and each node in its local sample in the feature matrix; In response to the feature vector of the node being binary data, the parameters of the Bernoulli distribution corresponding to each node are calculated according to the second modeling data, and the parameters of the Bernoulli distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula: ; in, represents the first The parameters of the Bernoulli distribution corresponding to the nodes, Indicates The enhanced feature matrix of nodes; In response to the feature vector of the node not being binary data, the parameters of the standard Gaussian distribution corresponding to each node are calculated according to the second modeling data, and the likelihood function of the parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data is shown as follows: ; in, represents the first The likelihood function of the parameters of the standard Gaussian distribution corresponding to the nodes, represents the second modeling data, represents the value of the jth sample in the second modeling data, The second modeling data The mean vector of the values of all samples in the nodes, The second modeling data The covariance matrix of the values of all samples in nodes, represents the probability density function; The parameters of the standard Gaussian distribution corresponding to each node calculated by the second modeling data are sampled to obtain an enhanced feature matrix, as shown in the following formula: ; in, represents a multivariate standard Gaussian distribution.
5. The paper classification method based on local enhancement and debiasing contrast according to claim 1 is characterized in that: The embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view are respectively input into two weight-shared linear projection layers to obtain a node representation of each node in the first enhanced view and a node representation of each node in the second enhanced view, specifically including: The node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view are calculated using the following formula: ; ; in, represents a linear projection layer, represents the embedding of each node in the first enhanced view, represents the embedding of each node in the second enhanced view, a node representation representing each node in the first enhanced view, a node representation representing each node in the second enhanced view, Represents the weight matrix in the linear projection layer.
6. The paper classification method based on local enhancement and debiasing contrast according to claim 1 is characterized in that: Constructing a debiasing contrast loss function according to the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view specifically includes: The nodes in the local sample of each node are regarded as false negative samples and penalized by assigning weights as shown in the following formula: ; in, Indicates The local samples corresponding to the nodes are Indicates the first enhanced view Node representation of nodes or the second enhanced view Node representation of nodes , Indicates the first enhanced view or the second enhanced view The weight of each node; The node representation of the same node in the first enhanced view and the node representation of the second enhanced view are regarded as positive sample pairs, and the node representation of other nodes in the first enhanced view and the node representation of the second enhanced view except the own node are regarded as negative sample pairs; The debiasing loss function for the positive sample pair is: ; in, Indicates the first enhanced view The node representation of nodes is Indicates the first The node representation of nodes is , , Indicates the first The node representation of nodes is represents the temperature coefficient, represents the cosine similarity function; The debiasing loss function for the positive sample pair is: ; in, , , Indicates the first enhanced view Node representation of nodes; The debiasing contrast loss function It is expressed as: ; in, Indicates the total number of nodes.
7. A paper classification device based on local enhancement and debiasing comparison, characterized in that: include: The model building module is configured to build and train a paper classification model to obtain a trained paper classification model, wherein the trained paper classification model includes a graph encoder and a classifier learned by graph contrast; the graph encoder learned by graph contrast is trained by the following process: building a graph dataset based on paper classification training data; Constructing a local sample for each node in the graph data set; using the local sample of each node and the multivariate Bernoulli distribution to perform structural enhancement and feature enhancement on each node, respectively, to obtain an enhanced adjacency matrix and an enhanced feature matrix; Sampling the enhanced adjacency matrix and the enhanced feature matrix twice to obtain a first enhanced view and a second enhanced view respectively; Inputting the feature matrix and adjacency matrix of the first enhanced view and the feature matrix and adjacency matrix of the second enhanced view into two weight-sharing graph encoders respectively, obtaining the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view, respectively inputting the embedding of each node in the first enhanced view and the embedding of each node in the second enhanced view into two weight-sharing linear projection layers, obtaining the node representation of each node in the first enhanced view and the node representation of each node in the second enhanced view; Constructing a debiasing contrast loss function according to a node representation of each node in the first enhanced view and a node representation of each node in the second enhanced view, and training the graph encoder based on the debiasing contrast loss function to obtain a graph encoder learned by graph contrast; The classification module is configured to obtain a collection of papers to be classified and construct a corresponding graph dataset, wherein the nodes in the graph dataset represent papers or authors, the edge between two nodes represents the citation relationship between the two papers or the cooperation relationship between the two authors, the feature vector of each node in the graph dataset corresponds to the element represented by the bag of words in each paper, the feature matrix and adjacency matrix in the graph dataset corresponding to the collection of papers to be classified are input into the trained paper classification model, first passed through the graph encoder learned by graph contrast learning to obtain the embedding of each node, and the embedding of each node is passed through the classifier to obtain the classification result, which includes the academic theme of the paper or the research field of the author.
8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Paper field classification method and device in citation network based on pseudo label depolarization and medium
CN118035448A
Graph contrast learning quotation network node classification method and system based on global and local information enhancement
CN118395116A
Graph contrast learning network node classification method and device based on denoising and mask reconstruction, and medium
CN118470411A
Dissertation classification method and apparatus based on classification model, and electronic device and medium
WO2021217930A1