Bidirectional GCN-BERT scientific data classification method based on rotary coding and dynamic gating
Patent Information
- Application Number
- CN202510294506.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The prior art is difficult to effectively capture the word order information and context semantic information of text in scientific data classification, resulting in insufficient performance.
A bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating is proposed. By combining BERT's high-quality semantic embedding and GCN's topological information capture capabilities, a multi-stage feature enhancement pipeline is built to realize text semantic depth mining and topological position perception modeling.
It effectively improves the accuracy, robustness and generalization performance of scientific data classification, and enhances the model's adaptability to complex data by optimizing context information capture and feature enhancement design.
Smart Images

Figure CN120216685A_ABST
Abstract
Description
Technical Field
[0001] This patent relates to the technical field of text classification, and particularly to a bidirectional GCN-BERT scientific data text classification method based on rotary encoding and dynamic gating. Background Art
[0002] In the technical field of text classification, the classification of scientific data has always been a challenging task. Traditional text classification methods, such as rule-based methods and shallow learning models, although they can achieve certain results in some cases, are often inadequate when dealing with large-scale, complex and semantically rich scientific data. In recent years, with the rapid development of deep learning technology, text classification methods based on deep neural networks have gradually emerged, providing new solutions to the scientific data classification problem.
[0003] As a powerful tool for processing graph-structured data, the Graph Convolutional Network (GCN) can effectively capture the topological relationships between nodes, thereby learning node representations with rich structural information. GCN generates graph-level representations by aggregating the features of neighboring nodes and performing linear transformations on each node. However, when dealing with text data, GCN often ignores the word order information and context semantic information of the text, which to some extent limits its performance in text classification tasks.
[0004] The emergence of the BERT pre-trained model has brought revolutionary changes to the field of natural language processing. BERT learns rich semantic information and language patterns through pre-training on a large-scale text corpus, and can generate high-quality semantic embeddings for each word in the text. In text classification tasks, using the semantic features extracted by BERT can significantly improve the classification accuracy. The RoBERTa model is an optimized version of the BERT model. It significantly improves the performance of the model through more refined pre-training strategies and larger datasets. Its structure is similar to BERT, but through improvements such as dynamic masking, larger training batches, and removing the NSP task, the model can learn richer semantic information. However, the RoBERTa model mainly focuses on the global semantic information of the text and lacks consideration of the local structure and context relevance of the text. Summary of the Invention
[0005] To overcome the limitations of existing methods, the present invention proposes a bidirectional GCN-BERT scientific data classification method based on rotation encoding and dynamic gating. The present invention fully combines the high-quality semantic embedding of BERT and the strong ability of GCN to capture topological information to achieve complementary advantages of the two. This method realizes in-depth text semantic mining and topological position-aware modeling by constructing a multi-stage feature enhancement pipeline. The multi-layer feature interaction network based on rotation position encoding realizes the position alignment of text global semantics and graph structure, and the bidirectional gated graph convolutional unit constructs a spatio-temporal joint feature propagation mechanism. In addition, the introduction of the CorNet network and the BiGRU module further enhances the model's ability to capture and understand semantic and context information. With the weight dynamic allocation strategy of the attention mechanism, this method effectively improves the accuracy, robustness, and generalization performance of scientific data classification while ensuring linear computational complexity. The method includes the following steps:
[0006] Step S1, perform adaptive preprocessing on the original text data to obtain the input data for training.
[0007] Step S2, use the RoBERTa pre-trained model to extract the deep semantic features of the sample data in the input data, obtain the classification token (CLS token) containing global information, add it as a special symbol to the beginning of the input text, and its corresponding output vector represents the global semantic information of the entire text.
[0008] Step S3, optimize the features of the CLS token through the CorNet neural network to enhance the feature representation;
[0009] Step S4, further perform position enhancement and global interaction on the CLS token with optimized features through the Rotation Position Enhanced Multi-Layer Feature Interaction Network (RP-MLFIN). This network combines rotation position encoding and linear attention, and enhances the local position information modeling ability through depthwise separable convolution and convolutional position encoding. Its linear complexity and efficient feature interaction mechanism enable the CLS token to contain text semantics, topological positions, and global dependencies simultaneously.
[0010] Step S5, embed a bidirectional gated recurrent unit (BiGRU) in the GCN to construct a bidirectional graph convolutional neural network based on BiGRU. After the output of each layer of GCN, a Dynamic Gating Feature Enhancement Network (DGFEN) is added before the GRU input. This network is a lightweight feature enhancer that combines a gating mechanism and additive attention.
[0011] Step S6, transfer the CLS token optimized by RP-MLFIN to the bidirectional graph convolutional neural network for feature interaction and graph convolution processing to generate a graph-level representation.
[0012] Step S7: Respectively pass the CLS token with optimized features and the graph-level representation after graph convolution processing to the classifier for classification prediction, generating BERT prediction and GCN prediction. Dynamically allocate the weights of the two through the attention mechanism, and fuse the prediction results to generate the final classification decision.
[0013] Preferably, the sample data adaptive preprocessing operation in step S1 includes: data cleaning, removing stop words in the sample data, generating embedding vectors of words and documents, calculating the adjacency matrix, etc. And convert the normalized adjacency matrix into a graph object DGL (Deep Graph Library), which is specifically designed to efficiently store graph structure data containing nodes, edges and their features and support complex graph operations, thus supporting efficient graph convolution operations. The normalization calculation formula for the adjacency matrix A with self-loops added is as follows:
[0014]
[0015] Where: I is the identity matrix, and D is the degree matrix after adding self-loops. is the inverse square root of the degree matrix, that is:
[0016] D ii =∑ j A ij +1
[0017] Where: D ii represents the degree of nodes i and, that is, the number of neighbor nodes of the node. A ij represents whether there is an edge between nodes i and j. If there is, it is 1, otherwise it is 0.
[0018] Preferably, the generation of embedding vectors of words and documents includes: loading pre-trained GloVe word vectors as the embedding vectors of words, and calculating the average word vector of each document as the embedding vector of the document.
[0019] Preferably, the construction of the DGL graph object includes: constructing a heterogeneous graph that contains both word nodes and document nodes, calculating the pointwise mutual information (PMI) and term frequency-inverse document frequency (TF-IDF) respectively. Define the edges between words and words, and between words and documents based on PMI and TF-IDF respectively. The weight of the edge between two nodes i and is defined as:
[0020]
[0021] Preferably, for obtaining the CLS token containing global information in step S2, it includes: using BERT to learn rich language structures and semantic information through pre-training. After the input text is encoded, a multi-layer bidirectional Transformer is used to capture context information. At the beginning of each input sequence, BERT inserts a CLS token, and the corresponding output vector synthesizes the semantic information of the entire sentence.
[0022] Preferably, for the CorNet residual network in step S3, it enhances the expressiveness of features through residual connections and non-linear transformations, including: applying an activation function to the output distribution and applying the ELU activation function to the context vector to introduce non-linearity. Adding the transformed output distribution to the original output distribution to form a residual connection.
[0023] Preferably, for the Rotation Position Enhanced Multi-Layer Feature Interaction Network (RP-MLFIN) in step S4, it is improved based on Transformer and achieves efficient long sequence modeling through a triple design of convolutional enhancement, hierarchical feature modulation, and rotation position encoding. The absolute position information is encoded into the Q / K vectors of the attention calculation through complex space rotation operations, which can more naturally maintain the relative position relationship compared with traditional position encoding. Using the QK projection activated by ELU and rotation position encoding to achieve low-complexity attention, while retaining the double-residual structure of attention residuals and MLP residuals to enhance gradient flow.
[0024] Step S4-1, convolutional enhancement captures local patterns in the sequence such as phrases, local dependencies, etc. by applying depthwise separable convolutions to the input features. The specific implementation includes: performing depthwise convolution operations independently for each channel to extract local features; using 1x1 pointwise convolutions to fuse channel information and enhance the feature expression ability; adding the convolutional output to the original input to retain the original feature information.
[0025] Step S4-2, hierarchical feature modulation gradually optimizes the feature representation through multi-layer feature interaction and a dynamic gating mechanism. The specific implementation includes: stacking multiple Transformer layers to gradually extract higher-level features; using gating weights to dynamically adjust the contribution degree of each layer of features and focus on key information; fusing features at different levels through weighted summation or concatenation to generate the final representation.
[0026] Step S4-3, rotation position encoding embeds position information into the feature vector through complex rotation operations. The specific implementation includes: regarding the feature vector as a complex number and encoding position information through a certain rotation angle; dividing the feature dimension into multiple two-dimensional subspaces and performing independent rotations for each subspace; applying rotation position encoding to the query Q and key K respectively to calculate the attention scores carrying relative position information.
[0027] Preferably, for the dynamic gating feature enhancement network (DGFEN) described in step S5, this network is a lightweight feature enhancer that combines a gating mechanism and additive attention. It achieves efficient feature enhancement through dynamic feature screening and global-local feature interaction, enabling the GRU to receive more expressive inputs, effectively improving the model performance without significantly increasing the model's computational burden.
[0028] The gating mechanism dynamically adjusts the importance of features through learnable weights. The specific implementation includes: generating gating weights by linearly transforming the input features; dynamically weighting the input features with the gating weights. Additive attention models the global dependencies between features through non-linear transformations. The specific implementation includes: projecting the input features into query Q and key K respectively; calculating attention scores through an additive function; aggregating global information according to the attention weights. DGFEN designs the gating mechanism and additive attention in series to form a two-stage process of "selection-enhancement".
[0029] Preferably, the bidirectional graph convolutional neural network includes multiple graph convolutional layers and bidirectional gated recurrent units. It extracts node features through the graph convolutional layers, and then captures the forward and backward context information of the features through the BiGRU, finally generating a more comprehensive feature representation. For each layer of graph convolution operation, the calculation formula is as follows:
[0030]
[0031] Where: H (l) is the input feature matrix of the l-th layer. W (l) is the weight matrix of the l-th layer. (A + I) is the adjacency matrix after adding self-loops. D is the degree matrix after adding self-loops. σ is the activation function.
[0032] The bidirectional graph convolutional neural network includes multiple graph convolutional layers and bidirectional gated recurrent units. It extracts node features through the graph convolutional layers, performs feature enhancement through DGFEN, and then inputs them into the BiGRU to capture the forward and backward context information of the features, finally generating a more comprehensive feature representation.
[0033] Preferably, for the dynamic weight allocation of the attention mechanism described in step S7, a fully connected layer and a Softmax function are used to calculate the attention weights, dynamically adjusting the contribution ratio of BERT and GCN, and fusing the output predictions of BERT and GCN through the attention weights to generate the final classification decision.
[0034] The substantial effects of the present invention are as follows:
[0035] Optimize the capture and expression of context information. In the present invention, a bidirectional gated recurrent unit (BiGRU) is incorporated into the traditional GCN. After the output of each layer of GCN and before the input of GRU, a dynamic gated feature enhancement network is added. Through dynamic feature screening and global-local feature interaction, efficient feature enhancement is achieved, enabling GRU to receive more expressive inputs, enabling the model to capture both forward and reverse context information simultaneously, and making up for the feature loss problem that may be caused by the unidirectional propagation mechanism. This bidirectional feature capture method enhances the comprehensiveness of feature representation, helps improve the processing ability of text data with strong context dependence, and does not significantly increase the computational burden of the model.
[0036] Innovative design of feature enhancement and fusion. By introducing a neural network (CorNet) to optimize the features of the CLS token, and further enhancing the position and global interaction through a rotated position-enhanced multi-layer feature interaction network (RP-MLFIN). This network combines rotated position encoding and linear attention, and enhances the local position information modeling ability through depthwise separable convolution and convolutional position encoding, enabling the CLS token to contain text semantics, topological position, and global dependencies simultaneously, strengthening the diversity and representational ability of the BERT output features. At the same time, the attention mechanism is used to assign weights and fuse the prediction results of BERT and GCN, achieving a dynamic balance of prediction information in different feature spaces, and improving the stability and robustness of the classification results. This design of feature optimization and fusion reflects the technical advantages of this method in complex feature processing and information integration.
[0037] Robustness adapted to scientific data classification tasks. Scientific data often has the characteristics of diversity and complexity. In the present invention, through DGL (Deep Graph Library) modeling, the adjacency relationship and feature distribution of scientific data are effectively characterized, and at the same time, combined with a bidirectional graph convolutional neural network, the implicit topological structure and semantic association between data are deeply mined. This design significantly enhances the adaptability of the model to the characteristics of scientific data, and improves the generalization performance and robustness to complex data.
[0038] In summary, through the collaborative design of multiple network modules and innovative feature optimization strategies, the present invention fully realizes the efficient integration of text semantic information and data structure information, and belongs to an innovative technology with substantial improvements in the field of text classification. Brief Description of the Drawings
[0039] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0040] Figure 2 It is a schematic diagram of the structure of the CorNet neural network;
[0041] Figure 3 Schematic diagram of optimized enhanced process for CLS marker
[0042] Figure 4 Schematic diagram of the structure of a bidirectional gated recurrent unit
[0043] Figure 5 Schematic diagram of the process of calculating graph convolution for a bidirectional graph convolutional neural network Specific implementation manners
[0044] The technical solution of the present invention will be further specifically described below through specific embodiments in combination with the accompanying drawings.
[0045] The present invention proposes a bidirectional GCN - BERT scientific data classification method based on rotation encoding and dynamic gating, as Figure 1 shown, including the following steps:
[0046] Step S1: Perform adaptive preprocessing on the original text data to obtain the input data for training. Data cleaning, removing stop words in the sample data, using word vectors to generate embedding vectors of words and documents, calculating the adjacency matrix and normalizing it, and constructing a heterogeneous graph representing the document - word relationship using PMI and TF - IDF weights. This heterogeneous graph is used as the input of the text classification model based on the graph convolutional network GCN. It includes the following sub - steps:
[0047] Step S1 - 1: Data loading and preprocessing. Among them,
[0048] Word vector loading, loading the pre - trained GloVe word vectors from the file glove.6B.300d.txt as the embedding vectors of words.
[0049] Data splitting, reading the document names and contents from the prepared files {dataset}.txt and {dataset}.clean.txt. Then, according to the labels in the document name file, such as (index, train / test, label), the data is split into a training set and a test set, and the training set and the dataset are shuffled. The training set is divided into a true training set and a validation set according to a ratio of 9:1. The document names and contents will be shuffled according to the shuffled training and test indexes. These shuffled datasets will be saved to the new files {dataset}_shuffle.txt and {dataset}_shuffle.txt.
[0050] Vocabulary construction: construct a vocabulary (vocab) from the shuffled document content. Use nltk.corpus.stopwords to remove stop words, such as common words like "the", "a", "is", etc. Perform text cleaning: lowercase, remove punctuation, etc. Calculate the occurrence times of words, map words to the list of documents in which they appear. Calculate how many documents each word appears in. Create a mapping word_id_map from words to IDs. Save the vocabulary to {dataset}_vocab.txt.
[0051] Label extraction: extract unique labels from the document metadata and save them to {dataset}_labels.txt.
[0052] Step S1-2: Feature vector construction. Among them,
[0053] Document feature vectors (x, tx): for each document in the training set and the test set, create a feature vector by averaging the word vectors of the words in the document. The resulting vector is stored as a sparse matrix using scipy.sparse.csr_matrix, where x is for training and tx is for testing. The formula for the document vector for each dimension j is:
[0054]
[0055] where: world_vector i [j] is the j-th dimension of the embedding vector of the i-th word in the document. doc_len is the number of words in the document. And checks are added to handle the cases where doc_len is 0, NaN, or Inf to prevent errors.
[0056] Label vectors (y, ty): create one-hot encoded label vectors y and ty for each document. y is for training and ty is for testing.
[0057] Combined feature vectors (allx): create a combined feature matrix allx that contains the document feature vector x and the word embedding vectors.
[0058] Step S1-3: Construct a graph object. Among them,
[0059] Window co-occurrence: create context windows of size window_size from the tokenized sentences to capture the co-occurrence relationships of words.
[0060] Word pair counting: used to calculate the co-occurrence times of word pairs within the window.
[0061] Adjacency matrix (adj): Construct an adjacency matrix adj to represent the graph, and the edges are weighted using point mutual information PMI. The normalization calculation formula for the adjacency matrix A with self-loops added is as follows:
[0062]
[0063] where: I is the identity matrix, and D is the degree matrix after adding self-loops, which is a diagonal matrix. is the inverse square root of the degree matrix, that is:
[0064] D ii = ∑ j A ij + 1
[0065]
[0066] where: n is the number of nodes. A ij indicates whether there is an edge between nodes i and j. If there is, it is 1; otherwise, it is 0. D ii represents the degree of node i, that is, the number of neighbor nodes of this node.
[0067] The PMI calculation formula is as follows:
[0068]
[0069] where: P(w1, w2) is the frequency of the simultaneous occurrence of words w1 and w2. P(w1) and P(w2) are the frequencies of the individual occurrences of words w1 and w2 respectively.
[0070] Document-word edges: The edges between documents and words are added to the adjacency matrix, and the weight is the product of the term frequency TF and the inverse document frequency IDF. The calculation formula is as follows:
[0071]
[0072] TF-IDF(t, d) = TF(t, d) × IDF(t)
[0073] where: t represents a specific term, that is, the word whose importance is to be evaluated. d represents the document, that is, the text segment containing the term t. D represents the document set, and N represents the total number of the entire document set. |d ∈ D: t ∈ d| represents the number of documents containing the term t, that is, the number of documents in the document set that contain this word. TF(t, d) represents the number of times the term t appears in the document d.
[0074] Define the edges between words and words, and between words and documents respectively based on PMI and TF-IDF. The weight of the edge between two nodes i and is defined as:
[0075]
[0076] The process of constructing a DGL graph object mainly includes defining the graph structure and setting the features of nodes and edges. First, extract the indices of source nodes and target nodes from the adjacency matrix, and use these indices to create a DGL graph instance. Then, if there is a node feature matrix, add it as node features to the graph; specifically, convert the normalized adjacency matrix into a list of source nodes and target nodes, use the graph function of DGL to create a graph object, and then add node features and edge features by setting ndata and edata respectively, thus completing the construction of the graph object.
[0077] Step S2: Use the RoBERTa pre-trained model to extract the deep semantic features of the sample data and obtain the CLS token containing global information; RoBERTa is a powerful pre-trained language model that can learn rich context information and semantic representations in the text. The architecture of the RoBERTa model is a multi-layer Transformer network that captures the relationships between words through the self-attention mechanism. In the input of RoBERTa, in addition to the text sequence itself, a special CLS token (Classification token) is added to the beginning of the sequence. Each layer of RoBERTa encodes the entire input sequence, and the final output of the CLS token is designed to aggregate the information of the entire input sequence, thus serving as the representation vector of the entire sentence. This design makes the output vector of the CLS token contain the global semantic information of the input text.
[0078] Load the pre-trained RoBERTa model and its corresponding tokenizer from the Hugging Face model repository. The tokenizer is responsible for converting the input text into a sequence of token IDs that the RoBERTa model can understand. Obtain the output feature dimension of the penultimate layer of the RoBERTa model. This is because the RoBERTa model usually contains a final linear layer for fine-tuning specific tasks, and the output of the penultimate layer is usually more suitable as a general semantic feature representation. Define a linear classifier to map the output features of the BERT model to the number of classes.
[0079] During the forward propagation process, pass the input token ID sequence (input_ids) and attention mask (attention_mask) to the RoBERTa model. Obtain the hidden state matrix output by the RoBERTa model, and select the output vector of the CLS token as the semantic representation of the text. These features contain global context information and can represent the semantics of the text better than traditional bag-of-words models or TF-IDF methods.
[0080] Step S3: Optimize the features of the CLS token through the CorNet residual network to enhance its feature representation; enhance the expressiveness of the features through residual connection and non-linear transformation, as Figure 3 shown, including: applying an activation function to the output distribution, and applying the ELU activation function to the context vector to introduce non-linearity. Add the transformed output distribution to the original output distribution to form a residual connection. The structure of the CorNet neural network is as Figure 2 shown.
[0081] Perform non-linear transformation on the input CLS features through the Sigmoid activation function:
[0082]
[0083] Convert the input output distribution into a context vector through a linear layer:
[0084] c = W1x + b1
[0085] Apply the ELU activation function to the context vector for non-linearization:
[0086] c = ELU(c)
[0087] Convert the processed context vector back into an output distribution through another linear layer:
[0088] x' = W2c + b2
[0089] Add the transformed output distribution to the output distribution of the original input to achieve a residual connection:
[0090] x out = x' + x
[0091] where: x is the output distribution of the input. W1, W2, b1, and b2 are the weights and biases of the linear layer.
[0092] Step S4: Further enhance the position and global interaction of the CLS token with optimized features through the Rotation Position-enhanced Multi-Layer Feature Interaction Network (RP-MLFIN). The Rotation Position-enhanced Multi-Layer Feature Interaction Network is improved based on the Transformer, and through the triple design of convolution enhancement, hierarchical feature modulation, and rotation position encoding, it realizes efficient long-sequence modeling. Encode the absolute position information into the Q / K vectors of the attention calculation through complex space rotation operations, which can more naturally maintain the relative position relationship compared with traditional position encoding. Use the QK projection activated by ELU and rotation position encoding to achieve low-complexity attention, while retaining the double-residual structure of attention residual and MLP residual to enhance gradient flow.
[0093] Given position p and frequency θk , and its complex rotation encoding is:
[0094]
[0095] where \(i\) is the imaginary unit. \(\theta\) k is a frequency-based scale parameter.
[0096] The complex rotation operation is achieved by multiplying with the original feature:
[0097]
[0098] The \(Q / K\) projection is linearized by the ELU function:
[0099] \(Q' = \text{ELU}(Q)+1\), \(K' = \text{ELU}(K)+1\)
[0100] where \(Q\) and \(K\) are the query and key vectors respectively. By using the ELU function, both \(Q'\) and \(K'\) are positive values, enabling cumulative calculations.
[0101] Step S5, embed a Bidirectional Gated Recurrent Unit (BiGRU) in the GCN, as Figure 4 shown, to construct a bidirectional graph convolutional neural network based on BiGRU. After the output of each layer of GCN and before the input to GRU, add a Dynamic Gated Feature Enhancement Network (DGFEN). DGFEN is a lightweight feature enhancer that fuses gating mechanisms with additive attention. Through dynamic feature screening and global-local feature interaction, efficient feature enhancement is achieved, enabling GRU to receive more expressive inputs, effectively improving model performance without significantly increasing the computational burden of the model.
[0102] The dynamic gating weight \(A\) is generated by the learnable parameter \(W\) g and the calculation formula is as follows:
[0103]
[0104] where: \(Q\) norm is the normalized position feature. \(W\) g \(\in \mathbb{R}\) d×1 is the learnable gating parameter vector. is the scaling factor to prevent gradient explosion.
[0105] The formula for global feature aggregation is as follows:
[0106]
[0107] where: \(A\) i is the attention weight at the \(i\)-th position, \(Q\) norm,i is the normalized \(i\)-th position feature. The calculation formula for feature modulation and residual fusion is as follows:
[0108] G expand = Expand(G, N)
[0109] Y = W2 · GELU(W1 · (G expand ⊙ K norm )) + Q
[0110] where: G expand represents replicating the global feature G N times along the sequence dimension. W1 and W2 are learnable fully connected parameters. K norm is the normalized local feature. Q is the original feature. GELU is the Gaussian error linear unit activation function.
[0111] Step S6, passing the optimized CLS token to a bidirectional graph convolutional neural network for feature interaction and graph convolutional processing to generate a graph-level representation; extracting node features through the graph convolutional layer, and then capturing the forward and backward context information of the features through BiGRU, finally generating a more comprehensive feature representation, as Figure 5 shown.
[0112] Loop through all graph convolutional layers, and apply dropout regularization before each layer except the first layer. Finally, perform the graph convolutional operation, store the output h of each layer of GCN in the gru_input list, and prepare to pass it to BiGRU. The graph convolutional formula is as follows:
[0113]
[0114] where: H (l) is the input feature matrix of the l-th layer. W (l) is the weight matrix of the l-th layer. (A + I) is the adjacency matrix after adding self-loops. D is the degree matrix after adding self-loops. σ is the activation function.
[0115] If gru_input is not empty, it needs to be reversed and stacked, and the output of GCN is input to GRU in the order from deep to shallow. Then stack the outputs of all GCN layers on dim = 1 to make it meet the input form of GRU. Pass the stacked gru_input to BiGRU, extract the hidden states of the last time step h and the first time step h_reverse of the BiGRU output respectively, which represent the forward and backward context information of the sequence, and concatenate them. For the forward GRU, the calculation formula is as follows:
[0116] z_t = sigmoid(W_z · x_t + U_z · h(t_1))
[0117] r_t = sigmoid(W_r · x_t + U_r · h(t_1))
[0118] h′_t = tanh(W·x_t + U(r_t ⊙ h_(t_1)))
[0119] h_t = (1 - z_t) ⊙ h_(t - 1) + z_t ⊙ h′_t
[0120] Where: x_t is the input of the GRU at time t, that is, the sequence formed by concatenating the output features of each layer of the GCN. h_(t) is the hidden state of the GRU at time t. z_t and r_t are the activation values of the update gate and reset gate of the GRU. W_z, U_z, W_r, U_r, W, and U are the weight matrices of the GRU. ⊙ is the Hadamard product (element-wise multiplication).
[0121] For the reverse GRU: The calculation is similar to that of the forward GRU, except that the time series is calculated in reverse.
[0122] Step S7, respectively pass the optimized CLS token and the features processed by graph convolution to the classifier to generate BERT predictions and GCN predictions. Dynamically allocate the weights of the two through the attention mechanism, fuse the prediction results, and generate the final classification decision. Use a fully connected layer and the Softmax function to calculate the attention weights, dynamically adjust the contribution ratios of BERT and GCN, and fuse the output predictions of BERT and GCN through the attention weights to generate the final classification decision. Concatenate the prediction results of BERT and GCN, and calculate the attention weights through a linear layer. The formula is as follows:
[0123]
[0124] Use the Softmax function to convert the output of the linear layer into attention weights:
[0125]
[0126] w att = Softmax(a)
[0127] Use the attention weights to perform weighted averaging on the prediction results of BERT and GCN to obtain the final prediction result:
[0128] P combined = w att [0]·P cls + w att [1]·P gcn
[0129] Use the Log_Softmax function to normalize the fused prediction result to obtain the final prediction probability distribution:
[0130]
[0131] P final = log_Softmax(P combined )
[0132] where: P cls is the Softmax probability output of the BERT classifier. P gcn is the Softmax probability output of the GCN. a is the output of the linear layer. z i is the i-th element of the input vector. n is the dimension of the vector z, i.e., the number of classes. is the sum of the exponents of all input elements, used for normalization. w att is the attention weight. P combined is the fused prediction result. P final is the final predicted probability distribution.
[0133] Next, the effectiveness of the bidirectional BERT-GCN scientific data classification method based on rotational position encoding and dynamic gating is demonstrated through experiments. Training is carried out on four public datasets, with accuracy used as the classification result metric. The details of the datasets are shown in Table 1. The experimental results of the classification accuracies of different models on different text classification datasets are shown in Table 2.
[0134] Table 1
[0135]
[0136] Table 2
[0137]
[0138] The model of the present invention is compared with the current state-of-the-art pre-trained models and GCN models: TextGCN, TextING, BERT, BERTGCN, RoBERTaGCN, and BERTGAT. The experimental results show that the method proposed in the present invention achieves better performance than the baseline models such as TextGCN, TextING, BERT, BERTGCN, RoBERTaGCN, and BERTGAT on the MR, R8, and R52 datasets. Although it is slightly inferior to BERTGCN on the ohsumed dataset, it is still highly competitive. These experimental results strongly verify the effectiveness of the method of the present invention.
[0139] To evaluate the model performance and verify the roles of the CorNet and RP-MLFIN networks as well as the importance of the DGFEN and BiGRU modules, this study designed two ablation experiments: The first experiment evaluated the impact of removing the CorNet and RP-MLFIN networks on the model performance. Specifically, the accuracy rates of the model of the present invention and the model with the CorNet and RP-MLFIN networks removed were compared. The second experiment aimed to explore the impact of removing the DGFEN and BiGRU modules on the model effect. Specifically, the accuracy rates of the model of the present invention and the model with the DGFEN and BiGRU modules removed were compared. In this way, by comparing with the performance of the original model, the roles of the CorNet and RP-MLFIN networks and the DGFEN and BiGRU modules in the model were accurately evaluated. The results of the accuracy rate comparison test are shown in Table 3.
[0140] Table 3
[0141]
[0142] The present invention does not directly use the CLS token extracted by the RoBERTa model, but inputs it into the CorNet and RP-MLFIN networks for feature enhancement. The CorNet and RP-MLFIN networks are designed with multi-layer non-linear transformation, context modeling, residual connection, convolution enhancement, hierarchical feature modulation, and rotational position encoding, effectively enhancing the deep semantic features contained in the CLS token and making it more expressive. This optimized CLS feature provides high-quality input for the subsequent graph convolution operation of the GCN.
[0143] The present invention embeds the DGFEN and BiGRU modules in the GCN to construct a bidirectional graph convolutional neural network. The output sequence of the multi-layer GCN is input into the BiGRU after passing through the DGFEN, and the BiGRU is used to capture the inter-layer sequence dependence of the GCN output, thereby capturing the information interaction at different levels in the graph convolution process. The bidirectional graph convolutional neural network can utilize both the graph structure information of the GCN, the dynamic feature screening and global-local feature interaction of the DGFEN, and the context modeling ability of the BiGRU for sequences, achieving complementary advantages.
[0144] The present invention dynamically fuses the BERT prediction from the optimized and enhanced CLS token and the graph convolution prediction from the GCN through the attention mechanism. The attention mechanism dynamically assigns weights to the BERT prediction and the GCN prediction according to the input samples, thereby achieving more effective prediction fusion and making full use of the advantages of both. This dynamic fusion strategy enables the model to adaptively adjust the prediction weights according to the data characteristics, thereby improving the robustness and generalization ability.
Claims
1. A bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating, characterized by: The following steps are involved: Step S1, performing adaptive preprocessing on the original text data to obtain input data for training; Step S2, using the RoBERTa pre-trained model to extract deep semantic features of sample data in the input data, and obtaining the classification tag CLS token containing global information as a symbol to be added to the beginning of the input text; Step S3, the CLS token is optimized through the CorNet neural network to enhance the feature representation; Step S4, the feature-optimized CLS token is subjected to position enhancement and global interaction through the rotation position enhanced multi-layer feature interaction network RP-MLFIN; Step S5, passing the CLS token optimized by RP-MLFIN to the bidirectional graph convolutional neural network for feature interaction and graph convolution processing to generate a graph-level representation; In step S6, the feature-optimized CLS token and graph-level representation are passed to the classifier respectively to generate BERT prediction and GCN prediction. The weights of the two are dynamically allocated through the attention mechanism, the prediction results are fused, and the final classification decision is generated.
2. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 1 is characterized in that: The adaptive preprocessing in step S1 includes: data cleaning, removing stop words in sample data, generating embedding vectors of words and documents, calculating the adjacency matrix A; and converting the normalized adjacency matrix into a graph object DGL.
3. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 2 is characterized in that: The generating of embedding vectors of words and documents includes: loading pre-trained GloVe word vectors as embedding vectors of words, and calculating the average word vector of each document as the embedding vector of the document.
4. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 3 is characterized in that: In the implementation process of the graph object DGL: a heterogeneous graph containing both word nodes and document nodes is constructed, and the point-by-point mutual information PMI and term frequency-inverse document frequency TF-IDF are calculated respectively; the edges between words and between words and documents are defined based on PMI and TF-IDF respectively; the weight of the edge between two nodes i, j and is defined as:
5. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 4 is characterized in that: The step S2 of obtaining the CLS token containing global information includes: using the language structure and semantic information learned by BERT through pre-training, encoding the input text, and then using a multi-layer bidirectional Transformer to capture contextual information; at the beginning of each input sequence, BERT inserts a CLS token, and its corresponding output vector integrates the semantic information of the entire sentence text.
6. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 5 is characterized in that: The rotation position enhanced multi-layer feature interaction network RP-MLFIN is based on the improvement of Transformer. It realizes long sequence modeling through the triple design of sequentially connected convolution enhancement, hierarchical feature modulation and rotation position encoding. The specific implementation process is as follows: Step S4-1, apply depthwise separable convolution on the input features, perform depthwise convolution operation on each channel independently to extract local features; use 1x1 point-by-point convolution to fuse channel information, and add the convolution output to the original input; Step S4-2: hierarchical features are stacked through multiple Transformer layers to gradually extract higher-level features; the contribution of each layer of features is dynamically adjusted using gating weights, and features at different levels are fused through weighted summation or concatenation to generate the final representation; Step S4-3, rotational position encoding embeds the position information into the feature vector through a complex rotation operation, regards the feature vector as a complex number, and encodes the position information through a preset rotation angle; The feature dimension is divided into multiple two-dimensional subspaces, and each subspace is rotated independently; rotation position encoding is applied to the query Q and key K respectively, and the attention score carrying relative position information is calculated.
7. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 6 is characterized in that: The bidirectional graph convolutional neural network includes multiple graph convolutional layers and bidirectional gated recurrent units. The node features are extracted through the graph convolutional layer, and the features are enhanced through DGFEN. Then, the features are input into BiGRU to capture the forward and reverse context information of the features, and finally a comprehensive feature representation is generated. The specific implementation process is as follows: A bidirectional gated recurrent unit BiGRU is embedded in GCN to build a bidirectional graph convolutional neural network based on BiGRU; a dynamic gated feature enhancement network DGFEN is added after the output of each layer of GCN and before the input of GRU; The dynamic gated feature enhancement network DGFEN is a lightweight feature enhancer that integrates a gating mechanism and additive attention; the gating mechanism dynamically adjusts the importance of features through learnable weights, and generates gating weights by linearly transforming input features; the input features are dynamically weighted by the gating weights; additive attention models the global dependencies between features through nonlinear transformations, and projects the input features into queries Q and keys K respectively; the attention scores are calculated through additive functions; global information is aggregated according to the attention weights, and DGFEN designs the gating mechanism and additive attention in series to form a two-stage process.
8. The bidirectional GCN-BERT scientific data classification method based on rotational coding and dynamic gating according to claim 7 is characterized in that: The dynamic weight allocation of the attention mechanism described in step S6 uses a fully connected layer and a Softmax function to calculate the attention weight, dynamically adjusts the contribution ratio of BERT and GCN, and fuses the output predictions of BERT and GCN through the attention weight to generate the final classification decision.
Citation Information
Patent Citations
Chinese short text classification method fusing contextual information graph convolution
CN114048754A
Chinese sentiment analysis method fusing syntactic dependency and part-of-speech based on graph convolutional network
CN114881042A
Column semantic recognition method and system based on context awareness of GCN and RoBERTa
CN117312989A
Network public opinion text classification method based on social network and public opinion theme
CN118377907A
Chinese text classification method and system based on structured multi-granularity feature fusion
CN119377740A
Cited By
Group behavior identification method and system based on cross-feature interaction Transform
CN120388335A
Academic paper title grading device and method based on Bert
CN120996032A