Robustness data classification system and method based on graph enhancement and dual-channel graph convolution
By introducing graph enhancement technology and dual-channel graph convolution network into the text classification system, and combining the attention mechanism for feature fusion, the problem of insufficient robustness of text classification models in industrial scenarios is solved, and higher classification accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510244994.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
The existing text classification methods cannot effectively deal with the problems of noise interference, outliers and data distribution changes in complex industrial scenarios, resulting in insufficient robustness.
A robust data classification system based on graph enhancement and dual-channel graph convolution is adopted to significantly improve the robustness and classification accuracy of the model through the diversified enhancement of graph structure and node features, and combined with dual-channel graph convolution and attention mechanism.
It significantly improves the classification performance and robustness of the model in complex industrial scenarios, can effectively process large-scale text data, provide efficient and stable classification results, and is suitable for equipment fault diagnosis and production quality control and other fields.
Smart Images

Figure CN120179825A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data mining and machine learning. More specifically, it relates to a robust data text classification system and method based on graph augmentation and dual-channel graph convolution. Background Art
[0002] With the rapid development of industrial Internet and intelligent manufacturing, enterprises will generate a large amount of text data during the production process, such as equipment logs, fault reports, process documents, and operation records. These text data are not only huge in quantity but also complex in content, often containing multi-source heterogeneous information. Efficient and accurate classification and analysis of them are of great significance for fault diagnosis, preventive maintenance, and production process optimization.
[0003] Existing text classification methods mainly rely on deep learning or traditional machine learning algorithms, usually only focusing on the semantic information or context information of the text, and it is difficult to effectively utilize the potential correlation structure between data. In addition, data in industrial environments are often affected by factors such as noise interference, outliers, and distribution changes, making traditional text classification models lack robustness when dealing with complex and changing industrial scenarios. Although some studies based on graph neural networks attempt to transform text data into graph structures for modeling, most of them only consider a single graph structure or feature representation and are difficult to adapt to the diverse data distributions under multiple perturbation scenarios.
[0004] Therefore, how to make full use of the structural correlation information of text data in industrial environments and improve the accuracy and stability of text classification through effective graph augmentation strategies and robust graph convolution models has become an urgent problem to be solved in the current technical field. Based on this, the present invention proposes an industrial robust text classification system based on graph augmentation and dual-channel graph convolution. By diversifying the enhancement at the graph structure and node feature levels and combining dual-channel graph convolution with attention mechanisms for feature fusion, the classification performance and robustness of the model in complex industrial scenarios can be significantly improved. Summary of the Invention
[0005] The present invention aims to provide a robust data classification system and method based on graph augmentation and dual-channel graph convolution to solve the problems that traditional text classification methods in the prior art cannot effectively cope with noise interference, outliers, and data distribution changes in complex industrial scenarios. By introducing graph augmentation technology and dual-channel graph convolution networks, the system can better capture the multi-level correlation information of text data and improve the robustness and classification accuracy of the model.
[0006] The technical solution of the present invention is as follows:
[0007] A robust data classification system based on graph augmentation and dual-channel graph convolution, the system includes a graph data generation module, an initialization module, a graph data augmentation module, a dual-channel graph convolution module, and an inference verification module;
[0008] The graph data generation module includes a data preprocessing module, a semantic feature extraction module, an adaptive graph construction module, and a data storage and loading module, which are used to preprocess industrial data and transmit the data to the graph data augmentation module;
[0009] The initialization module includes a model initialization module and a parameter initialization module, which are used to complete the standardized initialization of the model structure and parameters. The initialized model structure and parameter weights serve as the basis for the dual-channel graph convolution module;
[0010] The graph augmentation module includes a graph structure feature augmentation sub-module and a graph node feature augmentation sub-module, which are used to diversely augment the graph structure and node features of the data generated by the graph data generation module, and improve the model's adaptability to data changes in complex environments;
[0011] The dual-channel graph convolution module includes a first channel, a second channel, and an attention fusion module. The first channel is used to extract graph structure augmentation features, the second channel is used to extract graph node augmentation features, and a comprehensive feature representation is generated by combining the attention mechanism;
[0012] The inference verification module includes a classifier module and a verification module, which are used to complete the classification task of the input data and evaluate the model performance. A robust data classification method based on graph augmentation and dual-channel graph convolution, using the above system, includes the following steps:
[0013] (2-1) First, the graph data generation module preprocesses industrial data. For text data, cleaning, tokenization, and stop-word removal are performed, and semantic embeddings are extracted using a pre-trained BERT model, that is, the text is converted into a vector representation; for numerical data, normalization and feature engineering extraction are performed. Subsequently, the text and numerical data are fused to generate a unified node feature matrix X, where each row represents the features of a node; based on X, the similarity between nodes is calculated, and the edge set E is constructed using a dynamic neighborhood selection algorithm, that is, the connection relationship between nodes in the graph, and a normalized adjacency matrix A is generated. The node feature matrix X and the adjacency matrix A generated by the graph data generation module are used as the input of the graph augmentation module for subsequent augmentation processing;
[0014] (2-2) The initialization stage includes two parts: model structure initialization and parameter weight initialization:
[0015] Model structure initialization: Set the number of layers \(L\), the hidden layer dimension \(d\), and the activation function \(\sigma\) of the graph convolutional network to ensure that the network structure adapts to the input data \(X\) and task requirements;
[0016] Parameter weight initialization: Initialize the weight matrices in the network, including the weight matrix \(W_1\) of the first graph convolutional layer, the weight matrix \(W_2\) of the second graph convolutional layer, and the weight matrix \(W\) of the classifier. c These matrices are initialized using a uniform distribution or a normal distribution, or pre-trained parameters are loaded to improve the stability and efficiency of training;
[0017] (2-3) The graph augmentation module uses the graph data generated by the graph data generation module to augment the graph data \(G=(V, E)\) from two aspects, where \(V\) represents the set of nodes and \(E\) represents the set of edges;
[0018] Topological augmentation: Improve connectivity and diversity by modifying the structure of the graph, including randomly selecting pairs of nodes \((v\) i , v\) j ) that are not directly connected and adding new edges; randomly removing existing edges to generate a new edge set \(E'\);
[0019] Feature augmentation: Modify the node feature matrix \(X\) to enhance the feature expression ability, including adding random noise to generate enhanced features \(X' = X+\eta\), where \(\eta\) represents noise following a normal distribution; adding adversarial perturbations to generate perturbations \(\delta\) in a specific direction to obtain enhanced features \(X'' = X+\delta\); adding other transformations, such as scaling the features \((sX)\), translating the features \((X + t)\), or applying a non-linear transformation \(g(X)\);
[0020] (2-4) Use the dual-channel graph convolutional module to extract and fuse features from the augmented graph data, and perform graph convolution based on the model structure and parameter weights generated by the initialization module:
[0021] First channel: For the topologically augmented \(E'\), use the weight matrix \(W_1\) of the first graph convolutional layer, the weight matrix \(W_2\) of the second graph convolutional layer, and the weight matrix \(W\) of the classifier c to perform graph convolution operations, extract core information, and generate first-channel features \(H_1\);
[0022] Second channel: For the feature-augmented \(X'\) or \(X''\), use the weight matrix \(W_1\) of the first graph convolutional layer, the weight matrix \(W_2\) of the second graph convolutional layer, and the weight matrix \(W\) of the classifier c to perform graph convolution operations, extract core information, and generate first-channel features \(H_2\);
[0023] The outputs of the two channels are dynamically weighted and fused through the attention fusion module to generate comprehensive features \(H\);
[0024] (2-5) Finally, the fused comprehensive feature H is input into the classifier with the weight matrix W", thus completing the classification task and outputting the prediction result y. At the same time, the performance of the model is evaluated through the verification module, and the metrics include:
[0025] Verification accuracy R v : Measure the classification accuracy of the model on the test set;
[0026] Anti-interference ability R a : Measure the stability of the model when the input data undergoes small changes or noise is added;
[0027] Finally, calculate the comprehensive score. According to the comprehensive score, further adjust the model parameters and enhancement strategies to improve the overall robustness and generalization ability of the model. Preferably, the specific steps of the above step (2-1) are as follows:
[0028] (3-1) Data preprocessing
[0029] (3-1-1) Data cleaning
[0030] Clean the input industrial data (such as equipment logs, process documents, sensor data, etc.) to remove noise and redundant information; for text data, use regular expressions to remove special characters, punctuation marks, and stop words; for numerical data, perform missing value filling and outlier handling;
[0031] (3-1-2) Data standardization
[0032] Normalize the numerical data to ensure the consistency of the data distribution; for text data, perform word segmentation and stemming to generate a standardized text representation;
[0033] (3-1-3) Class label processing
[0034] For data with class imbalance, use stratified random sampling to ensure an equal number of samples in each class;
[0035] (3-2) Node feature extraction
[0036] (3-2-1) Text feature extraction
[0037] Use the pre-trained BERT model to extract the semantic embeddings of the text data and generate high-dimensional feature vectors. For long texts, use a sliding window or segmented processing to ensure the integrity of semantic information;
[0038] (3-2-2) Numerical feature extraction
[0039] For sensor data, extract statistical features (such as mean, variance, maximum, minimum) and temporal features (such as difference, sliding window statistics). For multi-source data, perform feature fusion to generate a unified node feature matrix X. For high-dimensional features (such as 768-dimensional BERT embeddings), use PCA to reduce the dimension to 256 dimensions to reduce computational complexity. Standardize the reduced-dimensional features to ensure the consistency of the feature distribution, and obtain the node set V and the node feature matrix X;
[0040] (3-3) Similarity calculation
[0041] (3-3-1) Based on the node feature matrix X, calculate the cosine similarity matrix S, where S ij represents the similarity between node i and node j;
[0042] (3-3-2) Dynamic neighborhood selection. For each node i, calculate the local density ρ i (such as the median of the similarities). Dynamically determine the number of neighbors k i :
[0043] k i =max(k min ,min(k max ,ρ i ×γ)) (1)
[0044] where k min and k max are the lower and upper limits of the number of neighbors respectively, and γ is the scaling factor;
[0045] (3-3-3) Edge set generation. For each node i, select the k i neighbors with the highest similarity to generate the edge set E. Adopt a two-way connection strategy to ensure the connectivity and information propagation ability of the graph;
[0046] (3-3-3) Edge weight normalization. Normalize the edge weights to generate the adjacency matrix A, where A ij represents the connection strength between node i and node j;
[0047] (3-4) Data storage and loading
[0048] (3-4-1) Data storage: Store the generated graph data G(V, E, X) in a standardized format (such as NPZ). Record the feature engineering parameters (such as PCA variance ratio, standardization coefficient, etc.) for subsequent reproduction and optimization;
[0049] (3-4-2) Data loading: Use a general data loading method in a standardized format such as NPZ for data loading.
[0050] Preferably, the specific steps of the above step (2-2) are as follows:
[0051] (4-1) First, the model initialization module defines the structure of the graph convolutional network and sets the parameters adapted to the task requirements: the number of network layers L, the hidden layer dimension d, and the activation function σ. The number of network layers L determines the depth of the graph convolution and selects an appropriate value according to the data complexity; the hidden layer dimension d controls the feature representation ability of each layer to ensure that the model can effectively learn the features of the input data; the activation function σ uses a non-linear function (such as ReLU or LeakyReLU) to enhance the non-linear expression ability of the network.
[0052] (4-2) The output dimension C of the classifier is set to the number of categories of the target task to ensure that the network can output results consistent with the classification target. Next, the parameter initialization module initializes the network weights and biases. The weight matrices W1, W2, W of the network c are randomly initialized using a uniform distribution or a normal distribution to ensure the rationality of the parameter distribution and avoid the situation of gradient vanishing or explosion; the bias vectors b1, b2, b c are initialized to zero to ensure the stability of the model output in the initial state. If there is a pre-trained model, load the pre-trained parameters for initialization, which can effectively accelerate the convergence of the model and improve the initial performance.
[0053] (4-3) After completing (4-2), the model structure and parameters are initialized, and the initialized model M is output. The model M contains the network weights W1, W2, W c and biases b1, b2, b c , and can accept the enhanced data as input to prepare for subsequent training and inference.
[0054] Preferably, the specific steps of the above step (2-3) are as follows:
[0055] (5-1) First, standardize the feature matrix X and the adjacency matrix A of the input graph data G=(V, E, X). The node feature matrix X is normalized to make the feature distributions uniform, which is beneficial to the optimization of the model; the adjacency matrix A is standardized to generate A' to balance the connection weights between nodes and ensure the stability and effectiveness of the graph convolution calculation.
[0056] (5-2) Subsequently, the topological enhancement sub-module operates on the edge set E of the input graph, aiming to increase the diversity and robustness of the data by adjusting the graph structure. The topological enhancement includes the following three specific steps: Adding edges is to randomly select unconnected node pairs (v i , v j ), according to the connection probability P(v i , v j)Add new edges to enhance the connectivity of the graph and the modeling ability of potential relationships; delete edges according to the edge weight w ij , remove edges with lower weights, optimize the graph structure, and reduce redundant information; rearrange edges is to randomly adjust the connection relationships of some edges to generate a new graph topology E′, thereby increasing the diversity of data;
[0057] (5-3) Secondly, the feature enhancement sub-module diversifies the node feature matrix X of the input graph, simulates the uncertainty of data and environmental changes. Adding noise is to add random noise to the node feature matrix to obtain enhanced features X′, such as noise with normal distribution or uniform distribution, to generate enhanced features X′, so as to improve the adaptability of the model to data fluctuations; Adversarial perturbation generation is to obtain enhanced features X′′ = X + δ by optimizing the generation of perturbations δ in a specific direction, which is used to test the robustness of the model;
[0058] (5-4) After completing the topology enhancement and feature enhancement, the generated enhanced graph includes graph structure feature enhanced data G = (V, E′, X), graph node feature enhanced data G = (V, E, X′) or G = (V, E, X″″). The enhanced graph data provides diverse inputs for the subsequent dual-channel graph convolution module, ensuring that the model has stronger adaptability to complex scenarios.
[0059] Preferably, the specific steps of the above step (2-4) are as follows:
[0060] (6-1) First, the dual-channel graph convolution module receives the feature enhanced graph data generated by the graph enhancement module and extracts features through the first channel and the second channel respectively; the first channel is used to process the graph structure feature enhanced data G = (V, E′, X), and extracts the core information in the original features through the adjacency matrix A′ and the weight matrix W1. Subsequently, the weight matrices W2 and W3 are continued to be used to extract features. The output feature H1 of the first channel retains the main structural information of the input data and ensures that the model can accurately capture the basic structural characteristics of the data and can adapt to the perturbations of the graph structure;
[0061] (6-2) Secondly, the second channel is used to process the graph node feature enhanced data G = (V, E, X′) or G = (V, E, X″). Similarly, the adjacency matrix A′ and the weight matrix W1 are used for feature extraction. Subsequently, the weight matrices W2 and W3 are continued to be used to extract features. The output feature H2 of the second channel mainly focuses on the change features of the graph nodes under specific conditions, enhancing the adaptability of the model to node perturbations; the output feature of the second channel provides supplementary information, which helps to capture the node-level perturbation characteristics that the first channel cannot fully handle;
[0062] (6-3). Subsequently, the dual-channel features are dynamically weighted and fused through the attention fusion module to generate a comprehensive feature representation H. The attention mechanism assigns weights α according to the importance of the output features of the two channels and calculates the fused feature H. The weights contributed by the first channel and the second channel are α and 1-α respectively, and the fusion formula is
[0063] H = αH1+(1-α)H2 (2)
[0064] The weight α is dynamically learned through the attention mechanism and can automatically adjust the contribution ratio of the two channels according to different task requirements. Finally, the fused comprehensive feature H is used as the input of the subsequent inference and verification module for classification tasks and performance verification. The dual-channel graph convolution module realizes the full modeling of diverse inputs through feature extraction and fusion, improving the classification ability and robustness of the model.
[0065] Preferably, the specific steps of the above steps (2-5) are as follows:
[0066] (7-1). First, the comprehensive feature representation H output by the dual-channel graph convolution module is input into the classifier module to complete the classification task. The classifier uses the weight matrix W c and the bias b c to calculate the classification result y and output the prediction probability of each sample belonging to different classes. The output dimension of the classifier is consistent with the number of classes C of the target task, ensuring that the model can accurately classify multi-class tasks.
[0067] (7-2). Subsequently, the verification module comprehensively evaluates the performance of the model, including two key indicators: verification accuracy and anti-interference ability. The verification accuracy R v is used to measure the classification accuracy of the model on the test set, that is, the proportion of correctly predicted samples; the anti-interference ability R a is used to test the stability of the model when the input data undergoes small changes or noise is added, reflecting the robustness of the model to uncertain data. The verification module comprehensively analyzes the performance of the model in different scenarios through these two indicators;
[0068] (7-3). Finally, the verification module calculates the comprehensive performance score R according to the verification results, by weighted combination of the verification accuracy R v and the anti-interference ability R a , and the calculation formula of the comprehensive score is
[0069] R = ω v R v + ω a R a (3)
[0070] where ω v and ω arespectively represent the weights of two indicators. The verification module dynamically optimizes the enhancement strategy and model parameters according to the scoring results, and adjusts the weights of the adjacency matrix, the feature enhancement intensity, and the channel weight ratio to further improve the robustness and generalization ability of the model.
[0071] The beneficial effects of the present invention are:
[0072] The present invention can effectively improve the robustness and adaptability of the industrial text classification system, and solves the problems of handling noise, outliers, and data distribution changes in the existing methods in complex environments; a system based on graph enhancement and dual-channel graph convolution is designed, and through multi-level feature extraction and fusion, the classification accuracy is significantly improved. This system can automatically process large-scale text data in industrial applications, provide efficient and stable classification results, and is widely applicable to fields such as equipment fault diagnosis and production quality control. Description of the Drawings
[0073] Figure 1 is the algorithm flowchart of the present invention. Detailed Embodiments
[0074] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.
[0075] As Figure 1 shown, a robust data classification system based on graph enhancement and dual-channel graph convolution, the system includes a graph data generation module, an initialization module, a graph data enhancement module, a dual-channel graph convolution module, and an inference verification module;
[0076] The graph data generation module includes a data preprocessing module, a semantic feature extraction module, an adaptive graph construction module, and a data storage and loading module, which are used to preprocess industrial data and transmit the data to the graph data enhancement module;
[0077] The initialization module includes a model initialization module and a parameter initialization module, which are used to complete the standardized initialization of the model structure and parameters, and the initialized model structure and parameter weights serve as the basis for the dual-channel graph convolution module;
[0078] The graph enhancement module includes a graph structure feature enhancement sub-module and a graph node feature enhancement sub-module, which are used to diversely enhance the graph structure and node features of the data generated by the graph data generation module, and improve the adaptability of the model to data changes in complex environments;
[0079] The dual-channel graph convolution module includes a first channel, a second channel, and an attention fusion module. The first channel is used to extract graph structure enhanced features, the second channel is used to extract graph node enhanced features, and a comprehensive feature representation is generated by combining the attention mechanism;
[0080] The inference and verification module includes a classifier module and a verification module, which are used to complete the classification task of the input data and evaluate the model performance. A robust data classification method based on graph augmentation and dual-channel graph convolution, using the above system, includes the following steps:
[0081] (2-1) First, the graph data generation module preprocesses the industrial data. For text data, cleaning, tokenization, and stop-word removal are performed, and the pre-trained BERT model is used to extract semantic embeddings, that is, converting the text into a vector representation; for numerical data, normalization and feature engineering extraction are performed. Subsequently, the text and numerical data are fused to generate a unified node feature matrix X, where each row represents the features of a node; based on X, the similarity between nodes is calculated, and the edge set E is constructed using the dynamic neighborhood selection algorithm, that is, the connection relationship between nodes in the graph, and a normalized adjacency matrix A is generated. The node feature matrix X and the adjacency matrix A generated by the graph data generation module are used as the input of the graph augmentation module for subsequent augmentation processing;
[0082] (2-2) The initialization stage includes two parts: model structure initialization and parameter weight initialization:
[0083] Model structure initialization: Set the number of layers L, the hidden layer dimension d, and the activation function σ of the graph convolutional network to ensure that the network structure adapts to the input data X and the task requirements;
[0084] Parameter weight initialization: Initialize the weight matrices in the network, including the weight matrix W1 of the first graph convolutional layer, the weight matrix W2 of the second graph convolutional layer, and the weight matrix W of the classifier c , and these matrices are initialized using a uniform distribution or a normal distribution, or pre-trained parameters are loaded to improve the stability and efficiency of training;
[0085] (2-3) The graph augmentation module uses the graph data generated by the graph data generation module to perform data augmentation on the graph data G=(V, E) from two aspects, where V represents the node set and E represents the edge set;
[0086] Topological augmentation: Improve connectivity and diversity by modifying the graph structure, including randomly selecting unconnected node pairs (v i , v j ) and adding new edges; randomly removing existing edges to generate a new edge set E′;
[0087] Feature Enhancement: Transform the node feature matrix X to improve the feature expression ability, including adding random noise to generate the enhanced feature X′ = X + η, where η represents noise following a normal distribution; adding adversarial perturbations to generate a perturbation δ in a specific direction to obtain the enhanced feature X′′ = X + δ; adding other transformations, such as scaling the features (sX), translating (X + t), or applying a non - linear transformation g(X);
[0088] (2 - 4) Use the dual - channel graph convolution module to extract and fuse features from the enhanced graph data, and perform graph convolution based on the model structure and parameter weights generated by the initialization module:
[0089] First channel: For the topologically enhanced E′, use the weight matrix W1 of the first graph convolution layer, the weight matrix W2 of the second graph convolution layer, and the weight matrix W of the classifier c Perform graph convolution operations to extract core information and generate the first - channel feature H1;
[0090] Second channel: For the feature - enhanced X′ or X′′, use the weight matrix W1 of the first graph convolution layer, the weight matrix W2 of the second graph convolution layer, and the weight matrix W of the classifier c Perform graph convolution operations to extract core information and generate the first - channel feature H2;
[0091] The outputs of the two channels are dynamically weighted and fused through the attention fusion module to generate the comprehensive feature H;
[0092] (2 - 5) Finally, input the fused comprehensive feature H into the classifier with the weight matrix W" to complete the classification task and output the prediction result y. At the same time, evaluate the model performance through the verification module, and the metrics include:
[0093] Verification accuracy R v : Measure the classification accuracy of the model on the test set;
[0094] Anti - interference ability R a : Measure the stability of the model when the input data undergoes small changes or noise is added;
[0095] Finally, calculate the comprehensive score. According to the comprehensive score, further adjust the model parameters and enhancement strategies to improve the overall robustness and generalization ability of the model. Preferably, the specific steps of the above step (2 - 1) are as follows:
[0096] (3 - 1) Data pre - processing
[0097] (3 - 1 - 1) Data cleaning
[0098] Clean the input industrial data (such as equipment logs, process documents, sensor data, etc.) to remove noise and redundant information; for text data, use regular expressions to remove special characters, punctuation marks, and stop words; for numerical data, perform missing value filling and outlier handling;
[0099] (3-1-2) Data Standardization
[0100] Normalize numerical data to ensure the consistency of data distribution; for text data, perform word segmentation and stemming to generate a standardized text representation;
[0101] (3-1-3) Class Label Processing
[0102] For imbalanced class data, use stratified random sampling to ensure an equal number of samples in each class;
[0103] (3-2) Node Feature Extraction
[0104] (3-2-1) Text Feature Extraction
[0105] Use a pre-trained BERT model to extract semantic embeddings of text data and generate high-dimensional feature vectors. For long texts, use a sliding window or segment processing to ensure the integrity of semantic information;
[0106] (3-2-2) Numerical Feature Extraction
[0107] For sensor data, extract statistical features (such as mean, variance, maximum, minimum) and temporal features (such as difference, sliding window statistics). For multi-source data, perform feature fusion to generate a unified node feature matrix X. For high-dimensional features (such as 768-dimensional BERT embeddings), use PCA to reduce the dimension to 256 dimensions to reduce computational complexity, and perform standardization processing on the reduced-dimensional features to ensure the consistency of feature distribution, obtaining the node set V and the node feature matrix X;
[0108] (3-3) Similarity Calculation
[0109] (3-3-1) Based on the node feature matrix X, calculate the cosine similarity matrix S, where S ij represents the similarity between node i and node j;
[0110] (3-3-2) Dynamic Neighbor Selection. For each node i, calculate the local density ρ i (such as the median of similarities). Dynamically determine the number of neighbors k i :
[0111] k i =max(k min ,min(k max ,ρi × γ)) (1)
[0112] where k min and k max are the lower and upper bounds of the number of neighbors respectively, and γ is the scaling factor;
[0113] (3 - 3 - 3) Edge set generation. For each node i, select the k i neighbors with the highest similarity to generate the edge set E. Adopt a two - way connection strategy to ensure the connectivity and information propagation ability of the graph;
[0114] (3 - 3 - 3) Edge weight normalization. Normalize the edge weights to generate the adjacency matrix A, where A ij represents the connection strength between node i and node j;
[0115] (3 - 4) Data storage and loading
[0116] (3 - 4 - 1) Data storage: Store the generated graph data G(V, E, X) in a standardized format (such as NPZ). Record the feature engineering parameters (such as PCA variance ratio, standardization coefficients, etc.) for subsequent reproduction and optimization;
[0117] (3 - 4 - 2) Data loading: Use a general data loading method for standardized formats such as NPZ to load the data.
[0118] Preferably, the specific steps of the above step (2 - 2) are as follows:
[0119] (4 - 1). First, the model initialization module defines the structure of the graph convolutional network and sets the parameters adapted to the task requirements: the number of network layers L, the hidden layer dimension d, and the activation function σ. The number of network layers L determines the depth of the graph convolution and an appropriate value is selected according to the data complexity; the hidden layer dimension d controls the feature representation ability of each layer to ensure that the model can effectively learn the features of the input data; the activation function σ uses a non - linear function (such as ReLU or LeakyReLU, etc.) to enhance the non - linear expression ability of the network;
[0120] (4 - 2). Set the output dimension C of the classifier to the number of classes of the target task to ensure that the network can output results consistent with the classification target. Next, the parameter initialization module initializes the network weights and biases. The weight matrices W1, W2, W c of the network are randomly initialized using a uniform distribution or a normal distribution to ensure the rationality of the parameter distribution and avoid the situation of gradient vanishing or explosion; the bias vectors b1, b2, b c are initialized to zero to ensure the stability of the model output in the initial state. If there is a pre - trained model, load the pre - trained parameters for initialization, which can effectively accelerate the convergence of the model and improve the initial performance.
[0121] (4-3) After completing (4-2), the model structure and parameters are initialized, and the initialized model M is output. The model M includes network weights W1, W2, W c and biases b1, b2, b c , and it can accept the enhanced data as input to prepare for subsequent training and inference.
[0122] Preferably, the specific steps of the above step (2-3) are as follows:
[0123] (5-1) First, standardize the feature matrix X and the adjacency matrix A of the input graph data G=(V, E, X). The node feature matrix X is normalized to make the feature distributions uniform, which is beneficial to the optimization of the model; the adjacency matrix A is standardized to generate A' to balance the connection weights between nodes and ensure the stability and effectiveness of graph convolution calculations;
[0124] (5-2) Subsequently, the topological enhancement sub-module operates on the edge set E of the input graph, aiming to increase the diversity and robustness of the data by adjusting the graph structure. The topological enhancement includes the following three specific steps: Adding edges is by randomly selecting unconnected node pairs (v i , v j ), and adding new edges according to the connection probability P(v i , v j ) to enhance the connectivity of the graph and the modeling ability of potential relationships; Deleting edges is based on the edge weight w ij , removing the edges with lower weights to optimize the graph structure and reduce redundant information; Rearranging edges is to randomly adjust the connection relationships of some edges to generate a new graph topology E' to increase the data diversity;
[0125] (5-3) Secondly, the feature enhancement sub-module diversifies the node feature matrix X of the input graph, simulating the uncertainty of the data and environmental changes. Adding noise is to add random noise to the node feature matrix to obtain the enhanced feature X', such as noise with a normal distribution or a uniform distribution, to generate the enhanced feature X' to improve the model's adaptability to data fluctuations; Adversarial perturbation generation is to optimize and generate a perturbation δ in a specific direction to obtain the enhanced feature X'' = X + δ for testing the robustness of the model;
[0126] (5-4) After completing the topological enhancement and feature enhancement, the generated enhanced graph includes the graph structure feature enhanced data G=(V, E', X), the graph node feature enhanced data G=(V, E, X') or G=(V, E, X''), and the enhanced graph data provides diverse inputs for the subsequent dual-channel graph convolution module, ensuring that the model has stronger adaptability to complex scenarios.
[0127] Preferably, the specific steps of the above steps (2-4) are as follows:
[0128] (6-1) First, the dual-channel graph convolution module receives the feature-enhanced graph data generated by the graph enhancement module and performs feature extraction through the first channel and the second channel respectively; the first channel is used to process the graph structure feature-enhanced data G=(V, E′, X), and extracts the core information in the original features through the adjacency matrix A′ and the weight matrix W1. Subsequently, the weight matrices W2 and W3 are continued to be used to extract features. The output feature H1 of the first channel retains the main structure information of the input data and ensures that the model can accurately capture the basic structural characteristics of the data and can adapt to the perturbations of the graph structure;
[0129] (6-2) Secondly, the second channel is used to process the graph node feature-enhanced data G=(V, E, X′) or G=(V, E, X″). Similarly, the adjacency matrix A′ and the weight matrix W1 are used for feature extraction. Subsequently, the weight matrices W2 and W3 are continued to be used to extract features. The output feature H2 of the second channel mainly focuses on the change features of the graph nodes under specific conditions and enhances the adaptability of the model to node perturbations; the output feature of the second channel provides supplementary information, which helps to capture the node-level perturbation characteristics that cannot be fully processed by the first channel;
[0130] (6-3) Subsequently, the dual-channel features are dynamically weighted and fused through the attention fusion module to generate a comprehensive feature representation H; the attention mechanism assigns weights α according to the importance of the output features of the two channels and calculates the fused feature H; the weights contributed by the first channel and the second channel are α and 1-α respectively, and the fusion formula is
[0131] H = αH1+(1-α)H2 (2)
[0132] The weight α is dynamically learned through the attention mechanism and can automatically adjust the contribution ratio of the two channels according to different task requirements. Finally, the fused comprehensive feature H is used as the input of the subsequent inference and verification module for classification tasks and performance verification. The dual-channel graph convolution module realizes the full modeling of diverse inputs through feature extraction and fusion, improving the classification ability and robustness of the model.
[0133] Preferably, the specific steps of the above steps (2-5) are as follows:
[0134] (7-1) First, the comprehensive feature representation H output by the dual-channel graph convolution module is input into the classifier module to complete the classification task. The classifier uses the weight matrix W c and the bias b c to calculate the classification result y and output the prediction probability that each sample belongs to different classes. The output dimension of the classifier is consistent with the number of classes C of the target task, ensuring that the model can accurately classify multi-class tasks.
[0135] (7-2), Subsequently, the verification module comprehensively evaluates the performance of the model, including two key indicators: verification accuracy and anti-interference ability. The verification accuracy R v is used to measure the classification accuracy of the model on the test set, that is, the proportion of correctly predicted samples; the anti-interference ability R a is used to test the stability of the model when small changes occur in the input data or noise is added, reflecting the robustness of the model to uncertain data. The verification module comprehensively analyzes the performance of the model in different scenarios through these two indicators;
[0136] (7-3), Finally, the verification module calculates the comprehensive performance score R according to the verification results, by weighted combination of the verification accuracy R v and the anti-interference ability R a , The calculation formula of the comprehensive score is
[0137] R = ω v R v + ω a R a (3)
[0138] where ω v and ω a respectively represent the weights of the two indicators. The verification module dynamically optimizes the enhancement strategy and model parameters according to the scoring results, adjusts the weights of the adjacency matrix, the intensity of feature enhancement, and the channel weight ratio to further improve the robustness and generalization ability of the model.
Claims
1. A robust data classification system based on graph augmentation and dual-channel graph convolution, characterized in that: The system includes a graph data generation module, an initialization module, a graph data enhancement module, a dual-channel graph convolution module, and an inference verification module; The graph data generation module includes a data preprocessing module, a semantic feature extraction module, an adaptive graph construction module, and a data storage loading module, which are used to preprocess the industrial data and transmit the data to the graph data enhancement module; The initialization module includes a model initialization module and a parameter initialization module, which are used to complete the standardized initialization of the model structure and parameters. The initialized model structure and parameter weights serve as the basis of the dual-channel graph convolution module; The graph enhancement module includes a graph structure feature enhancement submodule and a graph node feature enhancement submodule, which are used to perform diversified enhancement of graph structure and node features on the data generated by the graph data generation module, thereby improving the model's ability to adapt to data changes in complex environments; The dual-channel graph convolution module includes a first channel, a second channel, and an attention fusion module. The first channel is used to extract graph structure enhancement features, and the second channel is used to extract graph node enhancement features, and generates a comprehensive feature representation in combination with the attention mechanism. The reasoning verification module includes a classifier module and a verification module, which are used to complete the classification task of input data and evaluate the model performance.
2. A robust data classification method based on graph enhancement and dual-channel graph convolution, characterized in that Utilizing the system of claim 1, comprising the steps of: (2-1) First, the graph data generation module preprocesses the industrial data. For text data, it cleans, segments, and removes stop words, and uses the pre-trained BERT model to extract semantic embedding, that is, convert the text into a vector representation; for numerical data, it normalizes and extracts feature engineering, and then fuses the text and numerical data to generate a unified node feature matrix X, where each row represents the feature of a node; Based on X, the similarity between nodes is calculated, and the edge set E, that is, the connection relationship between nodes in the graph, is constructed using the dynamic neighborhood selection algorithm. The normalized adjacency matrix A is generated, and the node feature matrix X and the adjacency matrix A generated by the graph data generation module are used as the input of the graph enhancement module for subsequent enhancement processing; (2-2) The initialization phase includes two parts: model structure initialization and parameter weight initialization: Model structure initialization: Set the number of layers L, hidden layer dimension d, and activation function σ of the graph convolutional network to ensure that the network structure adapts to the input data X and task requirements; Parameter weight initialization: Initialize the weight matrix in the network, including the weight matrix W1 of the first graph convolution layer, the weight matrix W2 of the second graph convolution layer, and the weight matrix W of the classifier c ,These matrices are initialized with uniform or normal distribution, or loaded with pre-trained parameters, thereby improving the stability and efficiency of training; (2-3) The graph enhancement module uses the graph data generated by the graph data generation module to perform data enhancement on the graph data G = (V, E) from two aspects, where V represents a node set and E represents an edge set; Topological enhancement: Improve connectivity and diversity by modifying the structure of the graph, including randomly selecting pairs of nodes that are not directly connected (v i ,v j ) and add new edges; randomly remove connected edges to generate a new edge set E′; Feature enhancement: Transform the node feature matrix X to improve the feature expression ability, including adding random noise to generate enhanced features X′=X+η, where η represents noise that obeys the normal distribution; add adversarial perturbations to generate perturbations δ in a specific direction to obtain enhanced features X″=X+δ; add other transformations, such as scaling (sX), translating (X+t) or nonlinear transformation g(X) to the features; (2-4) Use the dual-channel graph convolution module to extract and fuse the enhanced graph data, and perform graph convolution based on the model structure and parameter weights generated by the initialization module: First channel: For topologically enhanced E′, use the weight matrix W1 of the first graph convolution layer, the weight matrix W2 of the second graph convolution layer, and the weight matrix W of the classifier c Perform graph convolution operation, extract core information, and generate the first channel feature H1; Second channel: For feature-enhanced X′ or X″, use the weight matrix W1 of the first graph convolution layer, the weight matrix W2 of the second graph convolution layer, and the weight matrix W of the classifier c Perform graph convolution operation, extract core information, and generate the first channel feature H2; The outputs of the two channels are dynamically weighted fused through the attention fusion module to generate the comprehensive feature H; (2-5) Finally, the integrated feature H is input into the classifier, and the weight matrix of the classifier is W", so as to complete the classification task and output the prediction result y. At the same time, the model performance is evaluated through the verification module. The indicators include: Verification accuracy R v :Measure the classification accuracy of the model on the test set; Anti-interference ability R a : Measures the stability of the model when small changes or noise are added to the input data; Finally, the comprehensive score is calculated, and the model parameters and enhancement strategies are further adjusted based on the comprehensive score to improve the overall robustness and generalization ability of the model.
3. The robust data classification method based on graph enhancement and dual-channel graph convolution according to claim 2, characterized in that: The specific steps of step (2-1) are as follows: (3-1) Data preprocessing (3-1-1) Data cleaning Clean the input industrial data to remove noise and redundant information; for text data, use regular expressions to remove special characters, punctuation marks, and stop words; For numerical data, fill missing values and handle outliers; (3-1-2) Data Standardization Normalize numerical data to ensure the consistency of data distribution; perform word segmentation and stem extraction on text data to generate standardized text representation; (3-1-3) Category label processing For data with imbalanced categories, stratified random sampling was used to ensure a balanced number of samples in each category; (3-2) Node feature extraction (3-2-1) Text feature extraction Use the pre-trained BERT model to extract the semantic embedding of text data and generate high-dimensional feature vectors. For long texts, use sliding windows or segmentation processing to ensure the integrity of semantic information; (3-2-2) Numerical feature extraction Extract statistical features and time series features from sensor data, perform feature fusion on multi-source data, generate a unified node feature matrix X, use PCA to reduce the dimension of high-dimensional features to 256 dimensions to reduce computational complexity, standardize the reduced-dimensional features to ensure the consistency of feature distribution, and obtain the node set V, the node feature matrix X; (3-3) Similarity calculation (3-3-1) Based on the node feature matrix X, calculate the cosine similarity matrix S, where S ij Represents the similarity between node i and node j; (3-3-2) Dynamic neighborhood selection, for each node i, calculate the local density ρ i ; Dynamically determine the number of neighbors k based on local density i : k i =max(k min ,min(k max ,ρ i ×γ))(1) where k min and k max are the lower and upper bounds of the number of neighbors, respectively, and γ is the scaling factor; (3-3-3) Edge set generation: for each node i, select the k nodes with the highest similarity. i neighbors, generate edge set E, and adopt a bidirectional connection strategy to ensure the connectivity and information dissemination ability of the graph; (3-3-3) Normalize edge weights. Normalize edge weights to generate an adjacency matrix A, where A ij Represents the connection strength between node i and node j; (3-4) Data storage and loading (3-4-1) Data storage: Store the generated graph data G(V, E, X) in a standardized format. Record feature engineering parameters to facilitate subsequent reproduction and optimization; (3-4-2) Data loading: Data is loaded using a common data loading method such as NPZ or other standardized formats.
4. The robust data classification method based on graph enhancement and dual-channel graph convolution according to claim 2, characterized in that: The specific steps of step (2-2) are as follows: (4-1) First, the model initialization module defines the structure of the graph convolutional network and sets the parameters that meet the task requirements: the number of network layers L, the hidden layer dimension d, and the activation function σ. The number of network layers L determines the depth of the graph convolution, and an appropriate value is selected according to the complexity of the data; the hidden layer dimension d controls the feature representation capability of each layer to ensure that the model can effectively learn the features of the input data; the activation function σ uses a nonlinear function to enhance the nonlinear expression capability of the network; (4-2) The output dimension C of the classifier is set to the number of categories of the target task to ensure that the network can output results consistent with the classification target. Next, the parameter initialization module initializes the network weights and biases. The network weight matrices W1, W2, W c Use uniform distribution or normal distribution for random initialization to ensure the rationality of parameter distribution and avoid gradient disappearance or explosion; bias vector b1, b2, b c Initialize to zero to ensure that the model output is stable in the initial state. If there is a pre-trained model, load the pre-trained parameters for initialization; (4-3) After completing (4-2), the model structure and parameter initialization are completed, and the initialized model M is output. The model M contains the network weights W1, W2, W c and bias b1, b2, b c , can accept enhanced data as input and prepare for subsequent training and reasoning.
5. The robust data classification method based on graph enhancement and dual-channel graph convolution according to claim 2, characterized in that: The specific steps of step (2-3) are as follows: (5-1) First, the feature matrix X and adjacency matrix A of the input graph data G = (V, E, X) are standardized. The node feature matrix X is normalized to make the features evenly distributed, which is conducive to model optimization; the adjacency matrix A is standardized to generate A′ to balance the connection weights between nodes and ensure the stability and effectiveness of graph convolution calculation; (5-2) Then, the topology enhancement submodule operates on the edge set E of the input graph, aiming to increase the diversity and robustness of the data by adjusting the structure of the graph. The topology enhancement includes the following three specific steps: adding edges is done by randomly selecting unconnected node pairs (v i ,v j ), according to the connection probability P(v i ,v j ) Add new edges to enhance the connectivity of the graph and the modeling ability of potential relationships; delete edges based on the edge weight w ij ,Remove the edges with low weights, optimize the graph structure, and reduce redundant information; ,rearrange the edges to randomly adjust the connection relationships of some edges to generate a new graph topology E′, thereby increasing the diversity of the data; (5-3) Secondly, the feature enhancement submodule performs diversified processing on the node feature matrix X of the input graph to simulate the uncertainty of data and environmental changes. Adding noise is to add random noise to the node feature matrix to obtain enhanced features X′; Adversarial perturbation generation is to optimize the perturbation δ in a specific direction to obtain the enhanced feature X″=X+δ, which is used to test the robustness of the model; (5-4) After completing topology enhancement and feature enhancement, the generated enhanced graph includes graph structure feature enhancement data G = (V, E′, X) and graph node feature enhancement data G = (V, E, X′) or G = (V, E, X″). The enhanced graph data provides diverse inputs for the subsequent dual-channel graph convolution module, ensuring that the model has stronger adaptability to complex scenarios.
6. The robust data classification method based on graph enhancement and dual-channel graph convolution according to claim 2, characterized in that: The specific steps of the step (2-4) are as follows: (6-1) First, the dual-channel graph convolution module receives the feature-enhanced graph data generated by the graph enhancement module, and extracts features through the first channel and the second channel respectively; the first channel is used to process the graph structure feature enhancement data G = (V, E', X), extracts the core information in the original features through the adjacency matrix A' and the weight matrix W1, and then continues to use the weight matrices W2 and W3 to extract features. The output feature H1 of the first channel retains the main structural information and of the input data, ensuring that the model can accurately capture the basic structural characteristics of the data and can adapt to the perturbations of the graph structure; (6-2) Secondly, the second channel is used to process the graph node feature enhancement data G = (V, E, X′) or G = (V, E, X″). It also uses the adjacency matrix A′ and the weight matrix W1 for feature extraction, and then continues to use the weight matrices W2 and W3 to extract features. The output feature H2 of the second channel mainly focuses on the change characteristics of the graph nodes under specific conditions, which enhances the adaptability of the model to node perturbations; the output features of the second channel provide supplementary information, which helps to capture the node-level perturbation characteristics that the first channel cannot fully process; (6-3) Then, the dual-channel features are dynamically weighted fused through the attention fusion module to generate a comprehensive feature representation H; The attention mechanism assigns weight α according to the importance of the output features of the two channels and calculates the fusion feature H; the weights contributed by the first channel and the second channel are α and 1-α respectively, and the fusion formula is H=αH1+(1-α)H2 (2) The weight α is dynamically learned through the attention mechanism, which can automatically adjust the contribution ratio of the two channels according to different task requirements. Finally, the fused comprehensive feature H is used as the input of the subsequent reasoning verification module for classification tasks and performance verification. The dual-channel graph convolution module realizes full modeling of diversified inputs through feature extraction and fusion, thereby improving the classification ability and robustness of the model.
7. The robust data classification method based on graph enhancement and dual-channel graph convolution according to claim 2, characterized in that: The specific steps of the step (2-5) are as follows: (7-1) First, the comprehensive feature representation H output by the dual-channel graph convolution module is input into the classifier module to complete the classification task. The classifier uses the weight matrix W c and bias b c Calculate the classification result y and output the predicted probability that each sample belongs to different categories. The output dimension of the classifier is consistent with the number of categories C of the target task, ensuring that the model can accurately classify multi-category tasks. ; (7-2) Then, the verification module conducts a comprehensive evaluation of the performance of the model, including two key indicators: verification accuracy and anti-interference ability. The verification accuracy R v It is used to measure the classification accuracy of the model on the test set, that is, the proportion of correctly predicted samples; anti-interference ability R a It is used to test the stability of the model when the input data changes slightly or noise is added, reflecting the robustness of the model to uncertain data. The verification module comprehensively analyzes the performance of the model in different scenarios through these two indicators; (7-3) Finally, the verification module calculates the comprehensive performance score R based on the verification results, and combines the verification accuracy R v And anti-interference ability R a The calculation formula of the comprehensive score is: R=ω v R v +oh a R a (3) where ω v and ω a They represent the weights of the two indicators respectively. The verification module dynamically optimizes the enhancement strategy and model parameters according to the scoring results, adjusts the adjacency matrix weight, feature enhancement strength and channel weight ratio to further improve the robustness and generalization ability of the model.