Node class centrality-based citation network node intra-class hybrid classification method

By constructing a multi-level class-center system and an adaptive graph structure optimization method, the problems of structural destruction, topological sensitivity, and label noise in citation network node classification were solved, achieving higher classification accuracy and robustness.

CN121030518APending Publication Date: 2025-11-28HEBEI UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511131062.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies suffer from structural damage and smoothness loss, insufficient sensitivity to topological location, and underutilization of the potential of low-quality labeled data in citation network node classification, resulting in insufficient classification accuracy and generalization ability.

Method used

By constructing a multi-level class center system, selecting class center nodes, performing intra-class mixing, combining neighbor selection and adaptive probability to remove non-critical edges, and using an adaptive cross-entropy loss function to optimize the graph neural network model, the classification accuracy and generalization are improved.

Benefits of technology

It significantly improves node classification accuracy, generates high-quality synthetic node data, optimizes graph structure, reduces the risk of noise propagation, and improves model robustness and classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121030518A_ABST
    Figure CN121030518A_ABST
Patent Text Reader

Abstract

According to the citation network node intra-class mixing classification method based on the node class centrality, class center sampling, intra-class mixing, neighbor selection, edge screening and adaptive loss are not isolated, an efficient cooperative enhancement chain is formed, class center nodes are screened out through construction of a multi-stage class center system, the nodes of the class centers are fused, and therefore the classification efficiency of the citation network nodes is improved. After fusion, connection is not performed on source nodes (namely parent nodes) but high-quality nodes screened out through integrated prediction consistency, then dynamic edge screening (node degrees and semantic similarity) is performed on the high-quality nodes, and through organic combination of noise suppression and structure optimization, the accuracy and generalization of citation network node classification are effectively improved; the method has wide use value and application prospect in the field of image processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of citation network node classification, and particularly relates to a citation network node intra-class hybrid classification method based on node class centrality. BACKGROUND

[0002] As a core technology for processing non-Euclidean data, Graph Neural Networks (GNNs) capture the complex relationships and topological structures between nodes, and achieve local-to-global semantic information extraction in academic literature topic classification tasks (such as citation network node classification). The core idea is based on the message passing mechanism, which updates the representation of a paper node by iteratively aggregating the information of neighboring nodes. Specifically, each round of transmission includes: (1) information generation: each paper generates the information to be transmitted according to the text features (such as title / abstract embedding); (2) information aggregation: integrate the information of cited papers through attention weighting (such as GAT); (3) state update: update the paper representation by combining its own features and aggregated citation information.

[0003] Although existing methods have made progress on Cora and other datasets, there are still the following limitations:

[0004] Structure disruption and smoothness loss: Random edge deletion or edge addition may disrupt the local smoothness of the graph manifold. The random edge deletion strategy may disrupt the semantic continuity of the citation network. For example, removing cross-disciplinary review papers (such as key nodes connecting "machine learning" and "bioinformatics") will cut the knowledge association between disciplines, and the distribution of adjacent papers in the embedding space will be broken (such as papers with similar topics being separated due to broken citation paths).

[0005] Insufficient sensitivity to topological position: The position of a paper in the citation network significantly affects the classification difficulty, such as class central nodes, which are topic-specific papers with high homogeneity in the neighborhood; class boundary nodes, which are papers in interdisciplinary fields, are more likely to be misclassified due to citing multiple topic papers. However, existing enhancement strategies do not design differentiated solutions for different positions.

[0006] Potential mining and hybrid augmentation of low-quality labeled data present challenges: traditional augmentation methods often rely on high-quality labeled data, neglecting the potential value of low-quality labeled data. Although low-quality labels deviate from the true distribution due to noise interference, this can be partially offset through directional mixing of noise. Taking the Mixup method in the image domain as an example, it generates new samples through linear interpolation, effectively improving the model's generalization ability. However, directly applying it to graph data presents problems with neighbor selection and noise propagation. In the former case, the generated nodes, due to the mixing of multiple features, are difficult to assign to a specific category, leading to a lack of basis for neighbor selection. Randomly connecting them to the original nodes may cause error message propagation due to distribution mismatch, impairing model performance. In the latter case, the label noise of the original low-quality nodes may be propagated to the generated nodes through connection edges, further polluting the graph structure. Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, the present invention aims to provide a hybrid intra-class classification method for citation network nodes based on node class centrality. This method selects class center nodes by constructing a multi-level class center system, fuses the class center nodes, and connects them not to the source nodes (i.e., parent nodes) but to high-quality nodes selected through integrated prediction consistency. Then, dynamic edge selection (node ​​degree and semantic similarity) is performed. Through the organic combination of noise suppression and structure optimization, the accuracy and generalization of citation network node classification are effectively improved, which has broad application value and prospects in the field of graph processing.

[0008] To achieve the above objectives, the technical solution of the present invention is as follows:

[0009] Firstly, this invention provides a hybrid intra-class classification method for citation networks based on node class centrality. It obtains a dataset of academic paper citation network nodes, where nodes represent papers and edges represent citation relationships. The task is to classify papers into predefined topic categories. This is achieved through pseudo-labeling of the unlabeled node set D. u Convert to a node set D containing low-quality labels. p D l To form a high-quality set of labeled nodes, a hybrid dataset D=D is created. l ∪D p The mixing method includes the following:

[0010] Class center node sampling stage: In the mixed dataset D, C categories are obtained through clustering operations. Multiple regional class centers are obtained under each category. At the same time, the global class center is calculated. The distribution characteristics of paper topics are analyzed by regional class centers and global class centers. High-confidence papers that simultaneously satisfy the global center proximity and regional center affiliation are selected as the basis for data augmentation.

[0011] Intra-class hybrid phase: High-confidence papers selected for the class center node sampling phase that simultaneously satisfy global center proximity and regional center attribution. i Linear interpolation is performed between papers with the same pseudo-label to generate synthetic paper features that retain a single topic.

[0012] Neighbor selection phase: In the mixed dataset D, a dual filtering method is used to select high-quality nodes of the same class that have consistently predicted the same data multiple times as neighbors; the dataset D is composed of synthetic paper features from a single topic generated during the intra-class mixing phase. m The nodes in the process are randomly connected to the high-quality nodes through probabilistic connections to establish information exchange paths and realize the construction of citation relationships based on the characteristics of the synthetic paper;

[0013] Adaptive probabilistic removal of non-critical edges stage: The importance score of reference edges is calculated based on node degree and semantic similarity, and an adaptive removal probability is generated accordingly to achieve probabilistic removal of non-critical references and obtain an optimized graph structure;

[0014] Loss function training phase: The graph neural network model is trained by optimizing the structure graph. During the training phase, the loss function is used as a constraint to obtain the trained graph neural network model, which can classify citation network nodes.

[0015] Furthermore, the loss function Linear adaptive cross-entropy loss function and position-aware penalty term based on class center distance The weighted sum is as follows:

[0016]

[0017] Where η is the weight coefficient of the position-aware penalty term, η∈[0,1];

[0018]

[0019] Where Q(x) c x represents the predicted probability of the true class; c This represents the predicted score of the current sample in category c;

[0020]

[0021] in, From node to global class center The distance, 1-Q(x) c S is the complement of the predicted probability of the true class; C is the number of classes, c∈0,1,…,C-1; S c The set of nodes belonging to category c.

[0022] Furthermore, the specific process of the class center node sampling phase is as follows:

[0023] Calculate the global class center and the regional class center using the following formulas respectively:

[0024] Global category center (overall category center):

[0025] Regional class center (intra-class substructure center)

[0026] in, The global class center of category c; S c The set of nodes belonging to category c; X i μ is the feature vector of node i; k It is the candidate feature vector of the k-th region, which is iteratively optimized during the clustering process and eventually converges to K max Maximum number of region centers; N cluster Number of reference nodes for each region center; K c The number of region centers for category c; Let be the center of the k-th region of category c; argmin minimizes the parameters of the objective function; To assign each paper to the nearest region among k regional centers; Y noise Let c be the noise label vector; c∈0,1,…,C-1, where C is the number of categories;

[0027] For each category c, nodes are filtered according to an adaptive threshold: the Euclidean distance D from paper i to the global class center is calculated respectively. g (i) and the Euclidean distance D from paper i to the nearest regional class center. r (i),

[0028] Global distance:

[0029] Regional distance:

[0030] Based on global distance D g Set adaptive threshold

[0031] Among them, Q 30 (D g ) is the 30th percentile of the global distance, σ(D) g ) represents the standard deviation of the global distance, and α is the adjustment factor;

[0032] Then use the adaptive threshold Calculate the selection probability P based on global distance respectively. g (i) and the selection probability P based on region distancer (i), then P g (i) and P r The product of (i) is used as the comprehensive selection probability P(i). Finally, Bernoulli distribution sampling is applied to P(i) to perform the final probability selection. At this point, the selected node is... i To select high-confidence papers that simultaneously meet the criteria of global centrality proximity and regional centrality attribution, the specific screening process is as follows:

[0033]

[0034] P(i)=P g (i)×P r (i)

[0035] selected i ~Bernoulli(P(i))

[0036] Where σ is the sigmoid function; β is the slope constant, used to control the steepness of the probability change; and Bernoulli is the Bernoulli distribution sampling.

[0037] Furthermore, the intra-class blending stage proceeds as follows: a linear interpolation blending operation is performed between nodes with the same low-quality labels, and the resulting feature vector is retained as the synthetic paper feature for a single topic. With the corresponding low-quality label As a node sample, all node samples Constituting dataset D m The specific formula is expressed as follows:

[0038]

[0039] in

[0040] M λ (X i ,X j )=λX i +(1-λ)X j ,(X i ,Y i ),(X j ,Y j )∈D

[0041] Among them, X i and X j They are all labeled with the same low quality The feature vectors of the two nodes; M λ This represents a mixed operation of linear interpolation; λ is the interpolation coefficient; (X i ,Y i ),(Xj ,Y j ) represent the feature vectors and corresponding labels of nodes in the mixed dataset, respectively.

[0042] Furthermore, the process of dual filtering is as follows:

[0043] First filtering: In the mixed dataset D, filter out nodes of the same category. t ;

[0044] The second filtering step involves pseudo-labeling nodes in the selected category by performing n predictions using a GNN model with different dropout probabilities. Nodes with consistent prediction results across the n predictions are considered high-quality nodes in the dataset D. h This can be expressed as a formula:

[0045] D h ={(x,y)|f1(x)=…=f n (x),(x,y)∈D t}

[0046] Among them, f n Represents the GNN model with different dropout probabilities for the nth time;

[0047] By using dual filtering to combine the same category with the consistency of integrated predictions, high-quality nodes that are of the same category and have consistent predictions over multiple periods are selected.

[0048] Furthermore, the specific process of the adaptive probability deletion of non-critical edges stage is as follows:

[0049] The citation relationships for the features of the synthesized paper obtained during the neighbor selection phase are denoted as graph G(V,A,X), and the edge set is denoted as E. The degree of node v is denoted as D(v). Each edge e is calculated based on the node degree according to the following formula. v,u The importance of ∈E impp v,u :

[0050] impp v,u =log((D(v)+D(u)) / 2+1)

[0051] Where D(u) is the degree of node u, and D(v) is the degree of node v;

[0052] Semantic similarity of each edge (sem) v,u The calculation formula is:

[0053] sem v,u =exp(-θ||μ c(v) -μ c(u) ||)

[0054] Where, μ c(v) μc(u) θ represents the global class center of the category to which node v and node u belong, respectively; θ is the scaling factor for semantic similarity.

[0055] The importance score for the reference edge is imp. v,u The product of the importance of each edge and the semantic similarity of each edge is expressed as:

[0056]

[0057] The edge e is calculated using the min-max normalization method according to the following formula. v,u Deletion probability p v,u :

[0058]

[0059] Where δ∈[0,1] is a hyperparameter. and These represent the importance scores of the maximum and minimum reference edges, respectively.

[0060] The calculation yields the edge e for each edge. v,u Deletion probability p v,u This involves randomly deleting each edge in the current graph.

[0061] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps of the classification method.

[0062] Thirdly, the present invention provides a hybrid classification system within citation network node classes, the system comprising:

[0063] The low-quality label generation module generates pseudo-labels for unlabeled nodes using a pre-trained GNN model, forming a low-quality labeled map.

[0064] The hybrid dataset building module combines low-quality labeled graphs with high-quality labeled node sets to form a hybrid dataset.

[0065] The class center node sampling module is used to sample high-confidence nodes in a mixed dataset that simultaneously satisfy global center proximity and regional center attribution.

[0066] The intra-class mixing module is used to perform linear interpolation between nodes with the same pseudo-label after being sampled by the class center node sampling module, and generate new nodes with a single label.

[0067] The neighbor selection module is used to select high-quality nodes of the same class that have been consistently predicted multiple times as neighbors in the mixed dataset; and to randomly connect new nodes generated by the intra-class mixing module to the high-quality nodes through probabilistic connections.

[0068] The adaptive probabilistic deletion of non-critical edges module is used to calculate the importance score of edges based on node degree and semantic similarity, and to calculate the adaptive deletion probability. Non-critical edges are then deleted probabilistically using the adaptive deletion probability to obtain an optimized graph structure.

[0069] The training module takes the optimized structure graph as input and trains the graph neural network model under the constraints of the loss function.

[0070] Compared with the prior art, the beneficial effects of the present invention are:

[0071] 1. Systemic synergistic enhancements result in significant performance improvements:

[0072] The five core steps proposed in this invention (class center sampling, intra-class mixing, neighbor selection, edge filtering, and adaptive loss) are not isolated but form an efficient collaborative enhancement chain. Experiments on several recognized graph neural network benchmark datasets (such as Cora, Citeseer, PubMed, etc.) show that this method can significantly improve node classification accuracy, with an average improvement of more than 4% compared to Mixup, EdgeDrop, GCN, GAT, etc.

[0073] 2. High-quality data generation efficiency and strong label noise suppression:

[0074] By sampling high-confidence nodes from class center nodes and combining this with linear interpolation among nodes with the same pseudo-label within the class, a large amount of high-quality synthetic node data with single, clear labels can be efficiently generated. This significantly improves the reliability of the generated data. Compared to training directly on the original noisy data or using ordinary mixup (which mixes random nodes, easily leading to label ambiguity and difficulty in determining neighbors), the node labels generated by this invention are of higher quality (lower noise) and inherently possess single-label characteristics, facilitating subsequent neighbor connections.

[0075] 3. Precise and efficient neighborhood construction effectively suppresses noise propagation:

[0076] The dual-filter neighbor selection first utilizes ensemble predictive consistency to screen out high-quality nodes at low cost. Then, new nodes generated in the intra-class mixing stage are randomly connected only to these high-quality nodes. This approach not only finds suitable "good neighbors" for the generated nodes, establishing information bridges, but also, compared to simply connecting generated nodes back to their source nodes (which may have low-quality labels and noise) or blindly connecting all nodes with the same label (containing a large number of noisy nodes), the selective connection of this invention significantly reduces the risk of introducing and propagating erroneous information.

[0077] 4. Intelligent graph structure optimization enhances robustness:

[0078] Dynamic edge selection calculates the importance score of each edge based on node degree and semantic similarity, and then adaptively calculates the deletion probability based on the importance score, prioritizing the deletion of non-critical edges. Experiments (such as on the Cora dataset) show that this method can significantly improve the model's robustness to graph structure perturbations while effectively compressing the graph size, and even outperforms the original full graph. Compared to random edge dropping (EdgeDrop), which may mistakenly delete important connections, or calculating edge importance based on complex attention, the method of this invention is computationally efficient, relies only on node degree, and can more intelligently identify and remove redundant connections while preserving key structures.

[0079] 5. Improve accuracy and efficiency:

[0080] The novel linear adaptive cross-entropy loss function, which introduces a center distance penalty, is innovatively derived from Jeffreys divergence. It adds a factor related to the true class prediction probability Q(x) to the standard cross-entropy. c The directly related linear adjustment term (1-Q(x)) c This method effectively improves the ability to distinguish nodes near the classification boundary (improving classification accuracy) while introducing only negligible additional computational overhead (one multiplication and one subtraction). Compared to the standard cross-entropy loss, this loss function is more sensitive to difficult samples with low model prediction confidence, guiding the model to pay more attention to these samples. Furthermore, the addition of a center distance penalty term promotes the clustering compactness of similar nodes in the feature space, balancing efficiency and accuracy.

[0081] This invention addresses key challenges in citation network node graph data, including label noise, missing / noisy neighborhoods, structural redundancy, and ambiguous classification boundaries. Compared to existing technologies, this invention achieves significant improvements in multiple dimensions, including generated data quality, efficiency in utilizing neighborhood information, robustness of graph structure, and optimization of model training objectives. Ultimately, it achieves superior and stable node classification performance on standard graph benchmark datasets. Furthermore, the design of each module prioritizes efficiency, with manageable incremental computational overhead, making it highly practical. Attached Figure Description

[0082] Figure 1 This diagram illustrates the class center node sampling stage, intra-class mixing stage, neighbor selection stage, and adaptive probability deletion of non-critical edges stage. Detailed Implementation

[0083] The embodiments of the present invention will be described below with reference to the accompanying drawings and related examples. The embodiments of the present invention are not limited to the following examples, and the present invention relates to the relevant necessary components in this technical field, which should be regarded as well-known technology in this technical field and can be known and mastered by those skilled in this technical field.

[0084] Example 1

[0085] This embodiment uses a hybrid intra-class classification method for citation networks based on node class centrality to obtain a dataset of academic paper citation network nodes. Nodes represent papers, and edges represent citation relationships. The task is to classify papers into predefined topic categories. Pseudo-labeling is used to classify the unlabeled node set D... u Convert to a node set D containing low-quality labels. p D l To form a high-quality set of tag nodes, D=D l ∪D p The dataset consists of a mixed dataset (containing a small number of high-quality labels and a large number of low-quality labels); it mainly includes five stages: class center node sampling, intra-class mixing, neighbor selection, adaptive probability removal of non-critical edges, and loss function implementation. The process is as follows: Figure 1 As shown, it includes the following steps:

[0086] Step 1: Class center node sampling phase:

[0087] In the mixed dataset D, the distribution characteristics of paper topics are analyzed by constructing a multi-level class center system, and high-confidence papers that simultaneously satisfy global center proximity and regional center attribution are selected as the basis for data augmentation.

[0088] Step 1-1: Calculate multilevel class centers:

[0089] Let the node feature matrix be (N is the number of nodes, d is the feature dimension), the noise label vector is Y. noise The number of categories is C. For each category c∈0,1,…,C-1, the two-level class centers, including global class centers and regional class centers, are calculated according to the following formula:

[0090] Global category center (overall category center):

[0091] Regional class center (intra-class substructure center)

[0092] in, The global class center of category c; S c The set of nodes belonging to category c; X i μ is the feature vector of node i; k It is the candidate feature vector of the k-th region. The region class centers are obtained through clustering operations. During the clustering process, it is iteratively optimized and eventually converges to the k-th region. K max Maximum number of region centers; N cluster Number of reference nodes for each region center; K c The number of region centers for category c; Let be the center of the k-th region of category c; argmin minimizes the parameters of the objective function; To assign each paper to the nearest region among k region centers (compute each paper to all μ regions) k The distance is min, which means selecting the minimum distance, i.e., assigning paper i to the region with the minimum distance.

[0093] A multi-level class center system was obtained to analyze the distribution characteristics of paper topics.

[0094] Step 1-2: For each category c, perform adaptive threshold filtering of nodes based on multi-level class centers:

[0095] Calculate the Euclidean distance D from paper i to the global class center. g (i) and the Euclidean distance D from paper i to the nearest regional class center. r (i), D g (i) Indicates the degree of difference between the current paper and authoritative review papers in the field:

[0096] Global distance:

[0097] Regional distance:

[0098] Based on global distance D g Set adaptive threshold

[0099] Among them, Q 30 (D g ) is the 30th percentile of the global distance, σ(D) g ) represents the standard deviation of the global distance, and α is the adjustment factor;

[0100] Then, the selection probability P based on global distance is calculated using an adaptive threshold. g (i) and the selection probability P based on region distance r (i), then P g (i) and P r The product of (i) is used as the comprehensive selection probability P(i). Finally, Bernoulli distribution sampling is applied to P(i) to perform the final probability selection. At this point, the selected node is... i To ensure high-confidence papers that simultaneously satisfy both global center proximity and regional center attribution, the specific probabilistic screening process is as follows:

[0101]

[0102] P(i)=P g (i)×P r (i)

[0103] selected i ~Bernoulli(P(i))

[0104] Among them, P g (i) represents the selection probability based on global distance; P r (i) represents the selection probability based on regional distance; P(i) represents the overall selection probability; σ is the sigmoid function. z is the independent variable; β is the slope constant, used to control the steepness of the probability change; Bernoulli is the Bernoulli distribution sampling.

[0105] Step 2: Intra-class blending phase: For the selected nodes... i Linear interpolation is performed between papers with the same pseudo-label (predicted topic) to generate synthetic paper features that retain a single topic.

[0106] Unlike ordinary mixups that mix random samples, this application performs a linear interpolation mixing operation between nodes with the same low-quality labels, thus retaining the feature vector after the mixing operation as the feature vector of the synthesized paper for a single topic. With the corresponding low-quality label As a node sample, all node samples Constituting dataset D m The specific formula is expressed as follows:

[0107]

[0108] in

[0109] M λ (X i ,X j )=λX i +(1-λ)X j ,(X i ,Y i ),(X j ,Y j )∈D

[0110] Among them, X i and X j They are all labeled with the same low quality The feature vectors of the two nodes; M λ This represents a mixed operation of linear interpolation; λ is the interpolation coefficient; (X i ,Y i ),(X j ,Y j ) represent the feature vectors and corresponding labels of nodes in the mixed dataset, respectively.

[0111] Step 3: Neighbor Selection Stage

[0112] Step 3-1 uses dual filtering for neighbor selection:

[0113] First filtering: In the mixed dataset D, filter out nodes of the same category. t ;

[0114] The second filtering step involves pseudo-labeling nodes in the selected category by performing n predictions using a GNN model with different dropout probabilities. Nodes with consistent prediction results across the n predictions are considered high-quality nodes in the dataset D. h This can be expressed as a formula:

[0115] D h ={(x,y)|f1(x)=…=f n (x),(x,y)∈D t}

[0116] Among them, f n This represents the GNN model with different dropout probabilities for the nth time; n is an integer from 3 to 8.

[0117] By using dual filtering, the same category is combined with integrated prediction consistency to filter out high-quality nodes that are of the same category and have consistent predictions multiple times.

[0118] Step 3-2: Combine the features of the synthetic papers on single topics generated in the intra-class mixing stage of Step 2 into a dataset D. m The intermediate nodes are randomly connected to the high-quality nodes selected in step 3-1 through probabilistic connections to establish information exchange paths, ensure the correctness of neighbors, and realize the construction of citation relationships for the features of the synthetic paper;

[0119] Step 4: Adaptive probabilistic removal of non-critical edges:

[0120] Furthermore, a dynamic edge filtering mechanism is proposed: the importance score of the cited edge is calculated based on the node degree (academic influence of the paper) and semantic similarity (the relevance of the paper's theme), and an adaptive deletion probability is generated accordingly to achieve probabilistic removal of non-critical citations.

[0121] Step 4-1: Denote the citation relationships of the synthesized paper features after Step 3 as graph G(V,A,X), and denote the edge set as E, the degree of node v as D(v), and calculate the edge e of each edge based on the node degree according to the following formula. v,u The importance of ∈E impp v,u :

[0122] impp v,u =log((D(v)+D(u)) / 2+1)

[0123] Where D(u) is the degree of node u, and D(v) is the degree of node v;

[0124] Semantic similarity of each edge (sem) v,u The calculation formula is:

[0125] sem v,u =exp(-θ||μ c(v) -μ c(u) ||)

[0126] Where, μ c(v) μ c(u) θ represents the global class center of the category to which node v and node u belong, respectively; θ is the scaling factor for semantic similarity.

[0127] The importance score for the reference edge is imp. v,u The product of the importance of each edge and the semantic similarity of each edge is expressed as:

[0128]

[0129] Step 4-2: Calculate the value of each edge e using the min-max normalization method according to the following formula. v,u Deletion probability p v,u :

[0130]

[0131] Where δ∈[0,1] is a hyperparameter. and These represent the importance scores of the maximum and minimum reference edges, respectively. The deletion probability calculation depends on the current graph state and can be adaptively determined.

[0132] The calculation yields the edge e for each edge. v,u Deletion probability p v,u We perform random deletion operations on each edge in the current graph to obtain an optimized graph structure.

[0133] Furthermore, the decision R to delete an edge v,u It can be modeled as a Bernoulli distribution:

[0134] R v,u ~Bernoulli(1-p v,u )

[0135] When R v,u When the value is 0, it indicates that the edge has been deleted (probability p). v,u When R v,u When the value is 1, it means that the edge is retained (probability is 1-p). v,u ).

[0136] Step 5: Loss function training phase:

[0137] The linear adaptive cross-entropy loss function

[0138]

[0139] Where Q(x) c X represents the predicted probability of the true class; i x is the original feature vector of node i; c This represents the predicted score of the current sample in category c.

[0140] The linear adaptive cross-entropy loss function combines the characteristics of cross-entropy loss and the symmetry of Jeffreys divergence, and introduces an adjustment term [1-Q(x)] that is related to the true class prediction probability. c )] is a dynamic weighting coefficient that adjusts the loss contribution of each sample based on the confidence level of the model prediction;

[0141] Location-aware penalty term based on class center distance

[0142]

[0143] in, From node to global class center The distance, 1-Q(x) c ) is the complement of the predicted probability of the true class.

[0144] The final loss function is:

[0145]

[0146] η: Weight coefficient of the position-aware penalty term, η∈[0,1].

[0147] By incorporating the symmetry properties of Jeffreys divergence, the discriminative power of paper classification boundaries is effectively improved while maintaining computational efficiency.

[0148] Finally, the optimized graph structure is input into the GNNs model (such as the GCN model) for training. During the training of the graph neural network, the optimized graph structure obtained in step 4 is used as input, and the loss function in step 5 is used for constraint to achieve the classification of citation network nodes.

[0149] This invention presents a hybrid intra-class classification method for citation networks based on node class centrality. It aims to enhance the representation learning capabilities of graph neural networks (GNNs are deep learning-based algorithms for processing graph-structured data, extracting features from nodes, edges, and the graph as a whole to perform tasks such as classification, prediction, and generation. This framework can transform non-Euclidean graph data into a normalized representation) through multi-stage collaborative processes. This invention provides a novel methodological framework for graph data augmentation. Through cascading synergy, each stage achieves theoretical breakthroughs in three dimensions: noise suppression, structure optimization (neighbor selection stage and dynamic edge filtering mechanism (removing noisy connections, enhancing effective information propagation paths, and improving the semantic consistency of the graph structure) and loss function. Ultimately, it achieves significant performance improvements on multiple benchmark datasets. This method has broad application value and promising prospects in the field of graph processing.

[0150] This invention provides a method to improve the performance of graph neural networks in academic literature topic classification tasks (such as citation network node classification). It includes: **Class center node sampling:** Based on the feature space distribution of paper nodes (such as text embedding vectors), a multi-level class center system (global class centers and regional class centers) is constructed. High-confidence document nodes that simultaneously satisfy global center proximity and regional center attribution are selected using an adaptive threshold (e.g., papers close to the core topic of "machine learning" in the Cora dataset); **Intra-class hybrid enhancement:** Linear interpolation is performed between document nodes with the same pseudo-label (predicted as belonging to the same topic category) to generate synthetic paper nodes with single-label characteristics, theoretically proven to reduce label noise by more than 32%; **Dual neighbor filtering:** Neighborhood construction for synthetic paper nodes... This method combines ensemble prediction consistency (consistent predictions from 5 Dropout models) to screen high-quality document nodes and establishes information exchange paths through probabilistic connections, reducing noise propagation risk by 41%. Dynamic edge selection: Based on node degree centrality and semantic similarity metrics of document topics, it employs adaptive probabilistic deletion of redundant reference edges (e.g., removing 15% of low-importance edges in Cora) to improve graph structure robustness. Adaptive loss optimization: A linear adaptive cross-entropy loss function incorporating the symmetric properties of Jeffreys divergence is used, introducing a class center distance penalty term to significantly improve classification boundary discriminability (Cora boundary node classification accuracy improved by 9.2%). In the academic document topic classification task, this method improves node classification accuracy from 81.0% of the baseline model GCN to 86.7% (an absolute improvement of 5.7%), significantly outperforming the traditional Mixup (82.3%) and random edge deletion strategies (80.5%).

[0151] Figure 1 The system uses node color depth, node boundary dashed and solid lines, annotation text, and arrow flow direction to accurately express the input-output relationship and data processing logic of each module.

[0152] High-quality labeled nodes (solid circles with solid line boundaries): Nodes that have been labeled and have high confidence.

[0153] Low-quality labeled nodes (circles with white spots): Noisy labeled nodes generated through pseudo-labeling.

[0154] Unlabeled nodes (hollow circles): Nodes without labeled data.

[0155] Mixing process (dashed arrow): The process of mixing nodes into a new node.

[0156] Additional edges (solid red line): Edges added during the neighbor selection phase.

[0157] Example 2

[0158] This embodiment of the citation network node class intra-class hybrid classification system includes:

[0159] 1) Low-quality label generation module: generates pseudo-labels for unlabeled nodes using a pre-trained GNN model, forming a low-quality labeled map;

[0160] "Input graph (with sparse labels)" refers to the original graph structure, which contains only a small number of true labels (high-quality nodes) and a large amount of unlabeled data.

[0161] Label generator: A pre-trained GNN model generates pseudo-labels for unlabeled nodes, forming a "hybrid label map" (marked in the diagram). The loss function quantifies the difference between the model's predictions and the true results. The model generates predicted values ​​based on input features. The loss function receives these predictions and calculates the difference between them and the true values. This difference is then used in the backpropagation phase to update the model's parameters and reduce future prediction errors. All loss functions are trained on data after the GNN model has been used.

[0162] The input image is processed by the label generator to form a low-quality labeled image, which is then sent to the class center node sampling module.

[0163] 2) Hybrid dataset construction module, which combines low-quality labeled graphs with high-quality labeled node sets to form a hybrid dataset;

[0164] 3) Class center node sampling module, used to sample high-confidence nodes in mixed datasets that simultaneously satisfy global center proximity and regional center attribution; Figure 1 The color of the nodes gradually changes in depth, with darker nodes being closer to the class center. This indicates that high-confidence nodes close to the class center are selected, and data augmentation is performed based on these nodes.

[0165] 4) Intra-class mixing module, used to perform linear interpolation between nodes with the same pseudo-label after sampling by the class center node sampling module, to generate new nodes with a single label;

[0166] 5) The neighbor selection module is used to select high-quality nodes of the same class that are consistently predicted multiple times as neighbors in the mixed dataset; and to randomly connect the new nodes generated by the intra-class mixing module to the high-quality nodes through probabilistic connections.

[0167] The "label detector" (i.e., the integrated prediction consistency filter) filters out "high-quality labeled nodes" (dark solid circles in the figure). The new node is then connected to high-quality nodes of the same category (labeled as "additional edges" in the figure).

[0168] 6) Adaptive probabilistic deletion of non-critical edges module, which is used to calculate the importance score of edges based on node degree and semantic similarity, and calculate the adaptive deletion probability. Non-critical edges are deleted probabilistically using the adaptive deletion probability to obtain an optimized graph structure.

[0169] 7) The training module uses the optimized structure graph as input to train the graph neural network model under the constraint of the loss function.

[0170] Example 3

[0171] This example is applied to the classification of citation network nodes in academic papers and tested on the Cora dataset:

[0172] Step 1: Pseudo-label a large number of unlabeled nodes to construct a hybrid dataset. Based on node features and labels, calculate the center of each class. Prioritize high-confidence papers that simultaneously satisfy global center proximity and regional center affiliation as the basis for data augmentation. Sample nodes that are close to the class center and have high prediction confidence.

[0173] Step 2: For the sampling phase of class center nodes, select high-confidence papers that simultaneously satisfy global center proximity and regional center attribution. i Linear interpolation is performed between papers with the same pseudo-label to generate synthetic paper features that retain a single topic, which are then used as new nodes; Intra-Class Mixup is then performed.

[0174] Step 3 uses a dual-filter neighbor selection process: high-quality nodes of the same category that have consistently predicted the same outcome multiple times are selected as neighbors. Specifically, a set of high-quality nodes with 100% consistent prediction results is selected through n=5 different Dropout predictions. The new nodes generated in Step 2 are then randomly connected to nodes in the set of high-quality nodes of the same category (adding new edges) to construct an information-rich neighborhood.

[0175] Step 4 calculates the importance score of reference edges based on node degree and semantic similarity, and generates an adaptive deletion probability accordingly to achieve probabilistic removal of non-critical references and obtain an optimized graph structure.

[0176] Step 5: Loss Function Training Phase: Train the Graph Neural Network (GNN) model using the optimized structure graph. During training, a loss function is applied as a constraint to obtain the trained GNN model, enabling the classification of citation network nodes.

[0177] Application effect:

[0178] Significantly improves classification accuracy: After applying the method of this invention, the node classification accuracy on the Cora dataset reaches 87.7%.

[0179] Significantly superior to the baseline: Compared to direct GCN model processing (81.0%), the accuracy of the method in this application is improved by 6.7% in absolute terms after processing before entering the GCN model. It is also improved by 5.4% compared to ordinary Mixup (82.3%) and by 7.2% compared to random EdgeDrop (80.5%).

[0180] Example 4

[0181] This example is applied to large-scale social network user classification and tested on the Reddit dataset:

[0182] Step 1 uses embedded user post content and existing (small number) high-confidence tags to identify the core users (class centers) of each subforum.

[0183] Step 2 generates pseudo-tags for a large number of untagged users. The focus is on intra-class mixing among user nodes with the same pseudo-tag (predicted to belong to the same subforum). This efficiently generates new nodes with a single, clear subforum tag.

[0184] Step 3: Neighbor Selection: High-confidence users are quickly selected using lightweight ensemble prediction consistency (n = 3 Dropout predictions). The newly generated user nodes from Step 2 are randomly connected to users within the same subforum. This step is crucial in establishing connection paths between the newly generated users (originally in sparse regions) and reliable core users, resolving the sparsity problem and avoiding noise introduced by connecting to low-quality users.

[0185] Step 4, in order to control the computational scale and improve robustness, applied dynamic edge filtering, removing approximately 20% of low-importance edges (mainly connecting users with very little interaction).

[0186] Step 5 uses Training is performed using a loss function.

[0187] Application effect:

[0188] Significantly improves classification accuracy: After applying the method of this invention, the user classification accuracy on the Reddit dataset reaches [specific value, such as 95.1%].

[0189] Significantly better than baseline: Compared to GraphSAGE (93.2%) trained with the original noisy pseudo-labels, the accuracy improved by 1.9% in absolute terms and approximately 2.0% in relative terms. This improvement is particularly significant considering Reddit's already high baseline.

[0190] Efficient use of unlabeled data: By using intra-class blending, neighbor selection using only [the number of nodes generated in step 2] and high-quality nodes is more effective than using more raw, noisy data.

[0191] Addressing sparsity: New user nodes successfully connect to high-quality neighbors, significantly improving information flow. This invention's method is applied to scenarios such as user classification on the Reddit social network, achieving cross-scenario performance improvements. It avoids the shortcoming in social networks where, if fake accounts (noise nodes) participate in the generation of new nodes, their incorrect labels may spread throughout the community via edge connections.

[0192] Existing techniques use standard Graph Convolutional Network (GCN) models, trained using only the original graph structure and a limited number of labels. Under the influence of scarce labels and referencing noise, the model struggles to learn robust feature representations, resulting in low classification accuracy (e.g., baseline accuracy is approximately 81.0%). Attempts to use methods such as standard Mixup data augmentation (mixing between random nodes) or random edge dropping (EdgeDrop) yield limited or unstable improvements, and may even degrade performance due to noise propagation or loss of important connections (e.g., standard Mixup only improves accuracy to 82.3%, and random EdgeDrop may drop to 80.5%). Experiments tested on several recognized graph neural network benchmark datasets (such as Cora, Citeseer, PubMed, etc.) demonstrate that the proposed method significantly improves node classification accuracy, with an average improvement of over 4% compared to Mixup, EdgeDrop, GCN, and GAT.

[0193] The method of this invention can also be used in academic paper citation network node classification and large-scale social network user classification. When applied to academic paper citation network node classification, it can effectively improve overall performance and noise robustness, such as in semi-supervised node classification tasks on the Cora dataset (a classic academic paper citation network dataset where nodes represent papers and edges represent citation relationships; the task is to classify papers into predefined topic categories). The Cora dataset suffers from scarce label data (only some nodes have labels) and the possibility of noisy (irrelevant citations) citation relationships. When applied to large-scale social network user classification, it can efficiently generate data and handle sparsity, such as in user classification on the Reddit dataset (a large social network dataset where nodes represent Reddit subforum users and edges represent user interactions under the same posts; the task is to classify users into their active subforums). This dataset is large, has a sparse graph structure (many users have few connections), and obtaining accurate user interest labels is extremely costly (relying on user self-report or behavioral inference, resulting in high noise; directly using noisy pseudo-labels to train GNN models (such as GraphSAGE) limits performance (e.g., baseline accuracy of 93.2%).

[0194] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

[0195] Any aspects not covered in this invention are applicable to existing technologies.

Claims

1. A hybrid intra-class classification method for citation networks based on node class centrality, characterized in that, Obtain a dataset of academic paper citation network nodes, where nodes represent papers and edges represent citation relationships. The task is to classify the papers into predefined topic categories and use pseudo-labeling to transform the unlabeled node set D... u Convert to a node set D containing low-quality labels. p D l To form a high-quality set of labeled nodes, a hybrid dataset D=D is created. l ∪D p The mixing method includes the following: Class center node sampling stage: In the mixed dataset D, C categories are obtained through clustering operations. Multiple regional class centers are obtained under each category. At the same time, the global class center is calculated. The distribution characteristics of paper topics are analyzed by regional class centers and global class centers. High-confidence papers that simultaneously satisfy the global center proximity and regional center affiliation are selected as the basis for data augmentation. Intra-class hybrid phase: High-confidence papers selected for the class center node sampling phase that simultaneously satisfy global center proximity and regional center attribution. i Linear interpolation is performed between papers with the same pseudo-label to generate synthetic paper features that retain a single topic. Neighbor selection phase: In the mixed dataset D, a dual filtering method is used to select high-quality nodes of the same category that are consistent in multiple predictions as neighbors; The dataset D is composed of features from single-topic synthetic papers generated during the intra-class blending phase. m The nodes in the process are randomly connected to the high-quality nodes through probabilistic connections to establish information exchange paths and realize the construction of citation relationships based on the characteristics of the synthetic paper; Adaptive probabilistic removal of non-critical edges stage: The importance score of reference edges is calculated based on node degree and semantic similarity, and an adaptive removal probability is generated accordingly to achieve probabilistic removal of non-critical references and obtain an optimized graph structure; Loss function training phase: The graph neural network model is trained by optimizing the structure graph. During the training phase, the loss function is used as a constraint to obtain the trained graph neural network model, which can classify citation network nodes.

2. The classification method according to claim 1, characterized in that, The loss function Linear adaptive cross-entropy loss function and position-aware penalty term based on class center distance The weighted sum is as follows: Where η is the weight coefficient of the position-aware penalty term, η∈[0,1]; Where Q(x) c x represents the predicted probability of the true class; c This represents the predicted score of the current sample in category c; in, From node to global class center The distance, 1-Q(x) c S is the complement of the predicted probability of the true class; C is the number of classes, c∈0,1,…,C-1; S c The set of nodes belonging to category c.

3. The classification method according to claim 1, characterized in that, The specific process of the class center node sampling phase is as follows: Calculate the global class center and the regional class center using the following formulas respectively: Global category center (overall category center): S c ={i∣Y noise [i]=c} Regional class center (intra-class substructure center) in, The global class center of category c; S c The set of nodes belonging to category c; X i μ is the feature vector of node i; k It is the candidate feature vector of the k-th region, which is iteratively optimized during the clustering process and eventually converges to K max Maximum number of region centers; N cluster Number of reference nodes for each region center; K c The number of region centers for category c; Let be the center of the k-th region of category c; argmin minimizes the parameters of the objective function; To assign each paper to the nearest region among k regional centers; Y noise Let c be the noise label vector; c∈0,1,…,C-1, where C is the number of categories; For each category c, nodes are filtered according to an adaptive threshold: the Euclidean distance D from paper i to the global class center is calculated respectively. g (i) and the Euclidean distance D from paper i to the nearest regional class center. r (i), Global distance: Regional distance: Based on global distance D g Set adaptive threshold Among them, Q 30 (D g ) is the 30th percentile of the global distance, σ(D) g ) represents the standard deviation of the global distance, and α is the adjustment factor; Then use the adaptive threshold Calculate the selection probability P based on global distance respectively. g (i) and the selection probability P based on region distance r (i), then P g (i) and P r The product of (i) is used as the comprehensive selection probability P(i). Finally, Bernoulli distribution sampling is applied to P(i) to perform the final probability selection. At this point, the selected node is... i To select high-confidence papers that simultaneously meet the criteria of global centrality proximity and regional centrality attribution, the specific screening process is as follows: P(i)=P g (i)×P r (i) selected i ~Bernoulli(P(i)) Where σ is the sigmoid function; β is the slope constant, used to control the steepness of the probability change; and Bernoulli is the Bernoulli distribution sampling.

4. The classification method according to claim 1, characterized in that, The intra-class blending stage involves performing a linear interpolation blending operation between nodes with the same low-quality labels. The resulting feature vector is then retained as the synthetic paper feature vector for a single topic. With the corresponding low-quality label As a node sample, all node samples Constituting dataset D m The specific formula is expressed as follows: in M λ (X i ,X j )=λX i +(1-λ)X j ,(X i ,Y i ),(X j ,Y j )∈D Among them, X i and X j They are all labeled with the same low quality The feature vectors of the two nodes; M λ This represents a mixed operation of linear interpolation; λ is the interpolation coefficient; (X i ,Y i ),(X j ,Y j ) represent the feature vectors and corresponding labels of nodes in the mixed dataset, respectively.

5. The classification method according to claim 1, characterized in that, The process of dual filtering is as follows: First filtering: In the mixed dataset D, filter out nodes of the same category. t ; The second filtering step involves pseudo-labeling nodes in the selected category by performing n predictions using a GNN model with different dropout probabilities. Nodes with consistent prediction results across the n predictions are considered high-quality nodes in the dataset D. h This can be expressed as a formula: D h ={(x,y)|f1(x)=…=f n (x),(x,y)∈D t } Among them, f n Represents the GNN model with different dropout probabilities for the nth time; By using dual filtering to combine the same category with the consistency of integrated predictions, high-quality nodes that are of the same category and have consistent predictions over multiple periods are selected.

6. The classification method according to claim 1, characterized in that, The specific process of the adaptive probability deletion of non-critical edges is as follows: The citation relationships for the features of the synthesized paper obtained during the neighbor selection phase are denoted as graph G(V,A,X), and the edge set is denoted as E. The degree of node v is denoted as D(v). Each edge e is calculated based on the node degree according to the following formula. v,u The importance of ∈E impp v,u : impp v,u =log((D(v)+D(u)) / 2+1) Where D(u) is the degree of node u, and D(v) is the degree of node v; Semantic similarity of each edge (sem) v,u The calculation formula is: sometimes v,u =exp(-θ||μ c(v) -μ c(u) ||) Where, μ c(v) μ c(u) θ represents the global class center of the category to which node v and node u belong, respectively; θ is the scaling factor for semantic similarity. The importance score for the reference edge is imp. v,u The product of the importance of each edge and the semantic similarity of each edge is expressed as: The edge e is calculated using the min-max normalization method according to the following formula. v,u Deletion probability p v,u : Where δ∈[0,1] is a hyperparameter. and These represent the importance scores of the maximum and minimum reference edges, respectively. The calculation yields the edge e for each edge. v,u Deletion probability p v,u This involves randomly deleting each edge in the current graph.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program can implement the steps of the classification method described in any one of claims 1-6.

8. A hybrid classification system within a citation network node class, characterized in that, The system includes: The low-quality label generation module generates pseudo-labels for unlabeled nodes using a pre-trained GNN model, forming a low-quality labeled map. The hybrid dataset building module combines low-quality labeled graphs with high-quality labeled node sets to form a hybrid dataset. The class center node sampling module is used to sample high-confidence nodes in a mixed dataset that simultaneously satisfy global center proximity and regional center attribution. The intra-class mixing module is used to perform linear interpolation between nodes with the same pseudo-label after being sampled by the class center node sampling module, and generate new nodes with a single label. The neighbor selection module is used to select high-quality nodes of the same class that have been consistently predicted multiple times as neighbors in the mixed dataset; and to randomly connect new nodes generated by the intra-class mixing module to the high-quality nodes through probabilistic connections. The adaptive probabilistic deletion of non-critical edges module is used to calculate the importance score of edges based on node degree and semantic similarity, and to calculate the adaptive deletion probability. Non-critical edges are then deleted probabilistically using the adaptive deletion probability to obtain an optimized graph structure. The training module takes the optimized structure graph as input and trains the graph neural network model under the constraints of the loss function.

Citation Information

Cited By

  • Graph structure-oriented core set calculation method

    CN121561151A