Short text classification processing method using graph attention network

By constructing text graphs and label graphs in short text classification processing, and applying graph attention mechanism for feature weighting and interactive processing, the problem of low classification accuracy caused by short text sparsity is solved, and higher classification accuracy and generalization ability are achieved.

CN120162435APending Publication Date: 2025-06-17JIANGSU UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510333554.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing short text classification processing methods are difficult to accurately predict labels due to the sparseness and characteristics of short texts, and the graph attention network is difficult to build rich and accurate graph structures when processing short texts, resulting in low model accuracy.

Method used

A short text classification processing method using graph attention network is proposed. Through text preprocessing, graph construction, graph attention mechanism learning, text label interaction and model training and prediction steps, text graph diagram and label diagram are constructed, and feature weighted summation and text label interaction processing are applied to improve the accuracy and generalization ability of the model.

Benefits of technology

By alleviating the sparseness of short texts, building rich graph structures, automatically paying attention to important node information, effectively capturing key information in short texts, improving model accuracy and generalization capabilities, and improving the accuracy of short text classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162435A_ABST
    Figure CN120162435A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and discloses a short text classification processing method utilizing a graph attention network, which comprises the following steps of: cleaning short text data, performing word segmentation and part-of-speech tagging, removing redundant words, converting each word into a vector representation with a fixed length by utilizing a word embedding technology, and extracting the vector representation by utilizing a graph attention network; obtaining a text vector matrix; constructing a text graph based on the text vector matrix, and constructing a tag graph according to a known tag relationship; a graph attention mechanism is applied to nodes in the text graph and the label graph, attention coefficients between the nodes are calculated, feature weighted summation is carried out, and updated node embedding representation is obtained. The short text classification processing method using the graph attention network aims at solving the problem that in an existing short text classification processing method, due to the fact that sparsity and characteristics of short texts are insufficient, labels are difficult to predict accurately through a traditional classification method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and specifically to a short text classification processing method using a graph attention network. Background Art

[0002] With the rapid advancement of social media and the Internet, short text data has shown an explosive growth trend. Short text forms such as tweets and comments, with their concise and immediate characteristics, have quickly become important carriers of information dissemination. Such data is not only huge in quantity but also diverse in content, covering all aspects of social life. Facing this trend, existing technologies are constantly innovating to meet the needs of short text processing. Natural language processing technology is widely used in fields such as short text classification and sentiment analysis to help people quickly extract valuable information from massive data. At the same time, machine learning algorithms are also continuously optimized to improve the accuracy and efficiency of short text processing.

[0003] However, due to the inherent sparsity and insufficient features of short texts, it poses a huge challenge to traditional classification methods. Since the length of short texts is limited and the amount of information contained is small, it is extremely difficult for traditional classification methods to extract effective features and build accurate models, making it difficult to accurately predict their labels. In existing technologies, although there have been attempts to use graph attention networks for text classification, expecting to improve the classification effect by capturing the correlation information between texts, when dealing with short texts, due to the limited features of short texts themselves, it is difficult to construct a rich and accurate graph structure, resulting in low model accuracy and inability to effectively capture the key information in short texts. Summary of the Invention

[0004] The purpose of the present invention is to solve the problem that in the existing short text classification processing methods, due to the sparsity and insufficient features of short texts, it is difficult for traditional classification methods to accurately predict labels, and to propose a short text classification processing method using a graph attention network.

[0005] The technical solution of the present invention to solve the above technical problems is as follows:

[0006] A short text classification processing method using a graph attention network includes the following steps:

[0007] S10. Text preprocessing step: Clean, tokenize, and perform part-of-speech tagging on the short text data, remove redundant words, and use word embedding technology to convert each word into a fixed-length vector representation to obtain a text vector matrix;

[0008] S20. Graph construction step: Construct a text graph based on the text vector matrix and construct a label graph according to the known label relationships;

[0009] S30. Graph Attention Mechanism Learning Step: Apply the graph attention mechanism to the nodes in the text graph and the label graph, calculate the attention coefficients between the nodes, and perform weighted summation of the features to obtain the updated node embedding representation;

[0010] S40. Text-Label Interaction Step: Perform interaction processing on the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain the text-label interaction matrix;

[0011] S50. Model Training and Prediction Step: Train the model using the cross-entropy loss function, update the model parameters through the backpropagation algorithm until the loss value converges, preprocess the short text to be predicted and input it into the trained model to obtain the predicted label.

[0012] Based on the above technical solutions, the present invention can also be improved as follows.

[0013] Furthermore, the S10. Text Preprocessing Step specifically includes:

[0014] S11. Data Cleaning Sub-Step: Remove duplicate data, similar data, and stop words in the short text data;

[0015] S12. Word Segmentation and Part-of-Speech Tagging Sub-Step: Use natural language processing technology to segment the short text and perform part-of-speech tagging, remove redundant conjunctions, auxiliary words, and adverbs, and retain nouns, verbs, adjectives, and adverbs as graph nodes;

[0016] S13. Text Embedding Sub-Step: Adopt the WordEmbedding technology to convert each word into a fixed-length vector representation to form a text vector matrix.

[0017] Furthermore, the S20. Graph Construction Step specifically includes:

[0018] S21. Text Graph Construction Sub-Step: Use a sliding window to construct edges on the text vector matrix. If two words are in the same window, add an undirected edge to form a text graph;

[0019] S22. Label Graph Construction Sub-Step: Construct a label graph according to the known label relationships. If there is an association between labels, add an undirected edge.

[0020] Furthermore, the text graph attention layer in the graph attention mechanism learning step specifically uses the following formula to calculate the attention coefficients and update the node embedding representation:

[0021] e ij =LeakyReLU(a T [Wh i ∥Wh j )

[0022]

[0023] h i ′ = σ(∑ j∈N(i) α ij Wh j )

[0024] where a is a learnable attention vector, W is a trainable transformation matrix, and σ is an activation function.

[0025] Furthermore, the text-label interaction step in S40 specifically includes:

[0026] S41, Interaction matrix calculation sub-step: Perform a dot product on the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain a text-label interaction matrix;

[0027] S42, Dimensionality reduction processing sub-step: Use a fully connected layer to perform dimensionality reduction processing on the interaction matrix to obtain the final text-label vector.

[0028] Furthermore, the loss function in the model training and prediction steps uses a cross-entropy loss function, and the specific formula is:

[0029]

[0030] where N is the number of samples, C is the number of classes, y ic is the true label, and p ic is the predicted probability.

[0031] Furthermore, it also includes a multi-head attention mechanism expansion step, introducing a multi-head attention mechanism in the text graph attention layer and the label graph attention layer to enrich the extraction ability of the model and improve the classification accuracy.

[0032] Furthermore, the calculation formula of the attention coefficient in the multi-head attention mechanism expansion step is modified to:

[0033]

[0034] where h represents the index of the attention head, and the final node embedding is represented as the average of the outputs of all attention heads.

[0035] Furthermore, it also includes a model optimization step, using an optimization algorithm during model training to accelerate the convergence process and introducing regularization techniques to prevent overfitting. When dealing with short social media text data, this method constructs efficient text graphs and label graphs and combines graph attention mechanisms to improve the accuracy and generalization of short text classification.

[0036] Compared with the prior art, the technical solution of the present application has the following beneficial technical effects:

[0037] In the present invention, by cleaning, segmenting and part-of-speech tagging short text data, and removing redundant words, the problem of short text sparsity is effectively alleviated. At the same time, the word embedding technology is used to convert each word into a vector representation of a fixed length, obtaining a text vector matrix, which provides rich feature representations for subsequent processing. Secondly, a text graph is constructed based on the text vector matrix, and a label graph is constructed according to known label relationships. This way of constructing the graph structure can fully exploit the correlation information between texts and between texts and labels, making up for the deficiency in the prior art that it is difficult to construct a rich and accurate graph structure, and providing a good foundation for the application of the graph attention mechanism. Then, the graph attention mechanism is applied to the nodes in the text graph and the label graph, the attention coefficients between nodes are calculated, and feature weighted summation is performed to obtain the updated node embedding representation. This can automatically focus on the node information that is more important for the classification task, effectively capture the key information in short texts, and improve the accuracy of the model. Then, the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer are interactively processed to obtain a text-label interaction matrix, further enhancing the connection between texts and labels, helping the model better understand the correspondence between texts and labels, and improving the accuracy of classification. Finally, the model is trained using the cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm until the loss value converges. This training method can continuously optimize the model, improve the generalization ability, and thus better adapt to the short text classification requirements in different fields and styles. Description of the Drawings

[0038] Figure 1 It is a method flow chart of a short text classification processing method using a graph attention network according to the present invention. Detailed Embodiments

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] A short text classification processing method using a graph attention network according to the present invention includes the following steps:

[0041] S10. Text preprocessing step: Clean, segment and part-of-speech tag the short text data, remove redundant words, and use the word embedding technology to convert each word into a vector representation of a fixed length, obtaining a text vector matrix;

[0042] S20, Graph Construction Step: Construct a text graph based on the text vector matrix and construct a label graph according to the known label relationships;

[0043] S30, Graph Attention Mechanism Learning Step: Apply the graph attention mechanism to the nodes in the text graph and the label graph, calculate the attention coefficients between the nodes, and perform feature weighted summation to obtain the updated node embedding representations;

[0044] S40, Text-Label Interaction Step: Perform interaction processing on the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain a text-label interaction matrix;

[0045] S50, Model Training and Prediction Step: Train the model using the cross-entropy loss function, update the model parameters through the backpropagation algorithm until the loss value converges, preprocess the short text to be predicted and input it into the trained model to obtain the predicted labels.

[0046] In a preferred embodiment of the present invention, it can be further configured as: S10, the text preprocessing step specifically includes:

[0047] S11, Data Cleaning Sub-step: Remove duplicate data, similar data, and stop words in the short text data;

[0048] S12, Word Segmentation and Part-of-Speech Tagging Sub-step: Use natural language processing techniques to segment the short text and perform part-of-speech tagging, remove redundant conjunctions, auxiliary words, and adverbs, and retain nouns, verbs, adjectives, and adverbs as graph nodes;

[0049] S13. Text Embedding Sub-step: Using the WordEmbedding technique, each word is converted into a fixed-length vector representation to form a text vector matrix. Duplicate data, similar data, and stop words in the short text data are removed, reducing data redundancy and noise, enabling subsequent processing to focus on more valuable information, alleviating the interference caused by short text sparsity, and avoiding the problem of model overfitting caused by duplicate and similar data, improving the generalization ability of the model. Secondly, natural language processing techniques are used to tokenize the short text and perform part-of-speech tagging. Redundant conjunctions, auxiliary words, and adverbs are removed, and nouns, verbs, adjectives, and adverbs are retained as graph nodes. This refined processing method can accurately extract key information in the short text, providing high-quality nodes for constructing a more accurate graph structure and helping the graph attention mechanism better capture the correlation information between texts. Furthermore, the WordEmbedding technique is adopted to convert each word into a fixed-length vector representation to form a text vector matrix. This word embedding method can encode the semantic information of words into vectors, enabling the model to better understand the semantic relationships between words, further enhancing the model's ability to capture short text features, and improving the accuracy of classification.

[0050] In the data cleaning sub-step, in addition to removing duplicate data, similar data, and stop words, special characters, garbled characters, and meaningless symbols in the short text are also cleaned to ensure the purity of the data. At the same time, a text similarity algorithm is used to identify and merge short texts with a similarity exceeding a certain threshold to avoid data redundancy. In the tokenization and part-of-speech tagging sub-step, a deep learning-based tokenization model is adopted, combined with dictionary matching and statistical methods, to improve the accuracy of tokenization. Part-of-speech tagging uses a pre-trained part-of-speech tagging model to perform part-of-speech tagging on the tokenized results, and redundant conjunctions, auxiliary words, adverbs, etc. are removed according to preset rules, and only nouns, verbs, adjectives, and adverbs that are important for the classification task are retained as graph nodes. In the text embedding sub-step, a pre-trained WordEmbedding model, such as Word2Vec, GloVe, etc., is adopted to convert each word into a fixed-length vector representation. At the same time, considering the context information of the short text, the word vectors are fine-tuned in combination with the context to make the word vectors better reflect the semantics of words in the short text. In addition, to further enhance the expressive ability of the text vector matrix, the word vectors are normalized, and an attention mechanism is introduced to assign different weights to word vectors at different positions to form a more representative text vector matrix.

[0051] In a preferred embodiment of the present invention, it can be further configured as: S20. The graph construction step specifically includes:

[0052] S21. Text graph construction sub-step: Use a sliding window to construct edges on the text vector matrix. If two words are within the same window, add an undirected edge to form a text graph.

[0053] S22. Label graph construction sub-step: Construct a label graph based on the known label relationships. If there is an association between labels, add an undirected edge. In the text graph construction sub-step, use a sliding window to construct edges on the text vector matrix. If two words are within the same window, add an undirected edge to form a text graph. This method can fully exploit the local association information between words in short texts. Due to the limited features of short texts, traditional graph construction methods are difficult to capture rich enough association information. However, the sliding window mechanism can effectively connect adjacent words to construct a graph structure that reflects the internal structure of the text, making up for the deficiency in the prior art of being difficult to construct a rich and accurate graph structure, providing a richer information basis for the subsequent graph attention mechanism, and helping the model better understand the semantics of short texts. In the label graph construction sub-step, construct a label graph based on the known label relationships. If there is an association between labels, add an undirected edge. This operation can clarify the semantic connection between labels. By constructing the label graph, the model can learn the co-occurrence pattern and semantic relevance between labels, so as to more accurately judge the label to which the short text belongs during classification, improving the accuracy of classification.

[0054] In the text graph construction sub-step, the size of the sliding window is dynamically adjusted according to the average length of the short text and the word distribution. If the average length of the short text is short, the window size will be appropriately reduced to ensure that finer associations between words can be captured; conversely, if the average length is long, the window size will be appropriately enlarged to cover more association information. At the same time, when adding undirected edges, different weights will be assigned according to the semantic similarity between words. The higher the semantic similarity, the greater the weight of the edge, which can better reflect the association strength between words. In addition, to avoid the constructed graph being too sparse or dense, the number of edges will be restricted and optimized. In the label graph construction sub-step, the determination of the association between labels is not only based on simple co-occurrence relationships, but also combines domain knowledge and semantic analysis. For example, for some labels with hyponymy, synonymy, or antonymy relationships, corresponding edges will be explicitly added to the label graph. At the same time, the label graph will be further optimized and adjusted according to the occurrence frequency and association strength of the labels in the training data, so that the label graph can more accurately reflect the semantic structure and association relationship between labels.

[0055] In a preferred embodiment of the present invention, it can be further configured that: the text graph attention layer in the graph attention mechanism learning step specifically calculates the attention coefficient and updates the node embedding representation using the following formula:

[0056] eij = LeakyReLU(a T [Wh i ∥Wh j )

[0057]

[0058] h i ′ = σ(∑ j∈N(i) α ij Wh j )

[0059] Where a is a learnable attention vector, W is a trainable transformation matrix, σ is an activation function. When the graph attention network processes short texts, due to the limited characteristics of short texts themselves, it is difficult to construct a rich and accurate graph structure, resulting in low model accuracy and inability to effectively capture key information in short texts. By using a specific formula to calculate the attention coefficient and update the node embedding representation, the correlation importance between different nodes can be automatically learned. Among them, the learnable attention vector can dynamically adjust the attention degree to different nodes, the trainable transformation matrix can flexibly transform the node features, and the activation function introduces non-linear factors, enabling the model to learn more complex feature representations.

[0060] The initialization of the learnable attention vector will be adjusted according to the distribution characteristics of short text data and the parameters of the pre-trained model. For example, if the short text data is mainly concentrated in certain specific fields, the initialization of the attention vector will tend to focus on the features related to these fields. The trainable transformation matrix will adopt the structure of a multi-layer perceptron (MLP). Through multi-layer linear transformations and non-linear activation functions, the node features are gradually abstracted and extracted. The activation function usually selects the ReLU function, which can effectively alleviate the gradient vanishing problem and introduce non-linear characteristics at the same time, enabling the model to learn more rich feature representations. When calculating the attention coefficient, the node features will be linearly transformed first, then the similarity between nodes is obtained through dot product operation, and finally normalized by the softmax function to obtain the final attention coefficient. When updating the node embedding representation, the attention coefficient is weighted and summed with the features of adjacent nodes, and then added with the features of the node itself. After being processed by the transformation matrix and the activation function, the updated node embedding representation is obtained. In addition, in order to further improve the performance of the model, a multi-head attention mechanism is also adopted, and the outputs of multiple attention heads are concatenated and fused to capture more comprehensive node relationships.

[0061] In a preferred embodiment of the present invention, it can be further configured as: S40. The text label interaction step specifically includes:

[0062] S41. Interaction Matrix Calculation Sub-step: Multiply the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain a text-label interaction matrix;

[0063] S42. Dimensionality Reduction Processing Sub-step: Use a fully connected layer to perform dimensionality reduction processing on the interaction matrix to obtain the final text-label vector. By multiplying the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer, a text-label interaction matrix is obtained. This operation can directly establish the association between text and labels, enabling the model to learn the semantic correspondence between text and labels, making up for the deficiencies in this aspect of the prior art. The dimensionality reduction processing sub-step uses a fully connected layer to perform dimensionality reduction processing on the interaction matrix to obtain the final text-label vector. This not only reduces the dimension of the data, lowers the computational complexity, but also extracts more representative features, avoiding the noise interference caused by high-dimensional data, further enhancing the model's understanding of the relationship between text and labels, and improving the accuracy and efficiency of short text classification.

[0064] In the interaction matrix calculation sub-step, before the dot product operation, the text embedding matrix and the label embedding matrix are normalized so that the norm of each vector is 1. This can avoid the influence of vector norm differences on the dot product result and ensure that the interaction matrix can more accurately reflect the association strength between text and labels. At the same time, to improve the sufficiency of interaction, a multi-head interaction method is adopted, that is, the dot product results of the text embedding matrices and label embedding matrices of multiple different heads are calculated respectively, and then these results are concatenated to form the final interaction matrix. In the dimensionality reduction processing sub-step, the number of neurons in the fully connected layer is adjusted according to the actual task and data scale. If the data scale is large and the feature dimension is high, the number of neurons will be appropriately increased to retain more valid information; conversely, if the data scale is small, the number of neurons will be appropriately reduced to avoid overfitting. In addition, during the dimensionality reduction process, a regularization term, such as L2 regularization, is introduced to constrain the weights of the fully connected layer to prevent the model from overfitting. At the same time, to accelerate the convergence of the model and improve performance, the batch normalization (Batch Normalization) technique is used to normalize the input of the fully connected layer.

[0065] In a preferred embodiment of the present invention, it can be further configured that: the loss function in the model training and prediction steps uses the cross-entropy loss function, and the specific formula is:

[0066]

[0067] where N is the number of samples, C is the number of categories, y ic is the true label, p icIt is the predicted probability. Due to the limited and sparse features of short texts, it is often difficult for the model to accurately learn the complex mapping relationship between the text and the label, resulting in poor classification performance. The cross-entropy loss function can measure the difference between the probability distribution predicted by the model and the probability distribution of the true label. By minimizing this loss function, the model can continuously adjust its parameters to make the predicted probability as close as possible to the true label, thereby enhancing the model's learning ability for short text classification tasks. The choice of this loss function helps the model better capture the key information in short texts and improve the classification accuracy. At the same time, due to the good mathematical properties and optimization characteristics of the cross-entropy loss function, the model can converge faster during training, improving the training efficiency and effectively overcoming the problems of low model accuracy and low training efficiency in the prior art.

[0068] To prevent the model from overfitting, a regularization term is added to the cross-entropy loss function. Common regularization terms include L1 regularization and L2 regularization. L1 regularization can make the model parameters sparser, which helps with feature selection; L2 regularization can prevent the model parameters from being too large and improve the generalization ability of the model. The coefficient of the regularization term is adjusted according to the data scale and model complexity. If the data scale is small and the model complexity is high, the coefficient of the regularization term will be appropriately increased to enhance the regularization effect. In addition, to accelerate the convergence of the model, optimization algorithms such as Adam and Adagrad are used to optimize the loss function. These optimization algorithms can automatically adjust the learning rate based on the historical information of the gradient, enabling the model to find the optimal solution faster during training. At the same time, during the training process, an early stopping strategy is adopted, that is, when the performance of the model on the validation set no longer improves, the training is stopped in advance to avoid model overfitting. In the prediction stage, for the input short text to be predicted, the same preprocessing operations as in the training stage are first performed, including data cleaning, word segmentation, part-of-speech tagging, text embedding, etc. Then the processed text is input into the trained model to obtain the predicted probability. Finally, the category with the highest probability is selected as the predicted label according to the predicted probability.

[0069] In a preferred embodiment, the present invention can be further configured to: further include a multi-head attention mechanism expansion step, introducing a multi-head attention mechanism in the text graph attention layer and the label graph attention layer to enrich the extraction ability of the model and improve the classification accuracy. By using multiple attention heads in parallel, the multi-head attention mechanism can focus on information in the text and labels from different angles and levels, enriching the extraction ability of the model. Each attention head can learn different feature representations and association patterns, enabling the model to capture richer semantic information and more complex text structures, thereby improving the ability to capture key information in short texts. At the same time, the multi-head attention mechanism can also enhance the robustness and generalization ability of the model, enabling the model to better adapt to short text data in different fields and styles.

[0070] Each attention head in the multi-head attention mechanism has independent learnable parameters, including attention vectors and transformation matrices. These parameters will be adaptively adjusted according to different input data during the training process, enabling each attention head to focus on different features. For example, some attention heads may focus more on nouns and verbs in the text, while some may focus more on adjectives and adverbs. In the text graph attention layer, the multi-head attention mechanism will perform multiple different linear transformations and attention calculations on the text embedding matrix to obtain multiple different attention outputs, and then concatenate these outputs and perform a linear transformation to obtain the final text embedding representation. In the label graph attention layer, the multi-head attention calculation will also be performed on the label embedding matrix to obtain richer label feature representations. To further improve the performance of the model, regularization processing will be performed on the output of the multi-head attention mechanism to prevent overfitting. At the same time, during the training process, the dropout technique will be used to randomly discard the outputs of some attention heads to enhance the robustness of the model. In addition, the number of heads in the multi-head attention mechanism will be adjusted according to the data scale and model complexity. If the data scale is large and the model complexity is high, the number of heads will be appropriately increased to improve the expression ability of the model; conversely, if the data scale is small, the number of heads will be appropriately reduced to avoid overfitting of the model.

[0071] In a preferred embodiment, the present invention can be further configured to: modify the attention coefficient calculation formula in the multi-head attention mechanism expansion step to:

[0072]

[0073] Among them, h represents the index of the attention head. The final node embedding is represented as the average of the outputs of all attention heads. When using the graph attention network to process short text classification, due to the sparse features and complex semantics of short texts, a single attention mechanism often has difficulty comprehensively capturing the key information and semantic relationships in the text, resulting in insufficient understanding of the text by the model and limited classification accuracy. The modified attention coefficient calculation formula introduces the attention head index, enabling each attention head to independently calculate the attention coefficient and focus on the information in the text from different perspectives, enhancing the model's ability to extract text features. At the same time, representing the final node embedding as the average of the outputs of all attention heads can integrate the advantages of each attention head, making the node embedding representation more comprehensive and stable, avoiding the bias that may be brought by a single attention head, further improving the model's ability to understand the semantics of short texts, and thus effectively enhancing the accuracy of short text classification.

[0074] For each attention head, the calculation of its attention coefficient is based on different learnable parameters, including the attention vector and the transformation matrix. These parameters will be adaptively adjusted according to the input data during the training process, enabling each attention head to focus on different feature patterns in the text. For example, some attention heads may be better at capturing local semantic information in the text, while others may focus more on the global structure information of the text. When calculating the attention coefficient, in addition to considering the similarity between node features, some additional information, such as the position encoding of the node and the degree of the node, will be introduced to enrich the calculation basis of the attention coefficient. After obtaining the output of each attention head, it will be normalized to ensure the comparability of the outputs of different attention heads. Then, the outputs of all attention heads will be averaged to obtain the final node embedding representation. This averaging operation can balance the contributions of each attention head, making the final node embedding representation more robust. In addition, to further improve the performance of the model, the number of attention heads will also be optimized. The number of attention heads will be dynamically adjusted according to the scale, complexity of the dataset, and the training situation of the model. If the dataset scale is large and the complexity is high, the number of attention heads will be appropriately increased to improve the model's expressive ability; conversely, if the dataset scale is small, the number of attention heads will be appropriately reduced to avoid overfitting of the model. At the same time, during the training process, the gradient clipping technique will be used to prevent gradient explosion and ensure the stable training of the model.

[0075] In a preferred embodiment, the present invention can be further configured as follows: It further includes a model optimization step. During the model training process, an optimization algorithm (such as the Adam optimizer) is used to accelerate the convergence process, and a regularization technique (such as L2 regularization) is introduced to prevent overfitting. When the method processes short social media text data (such as tweets, comments, etc.), by constructing an efficient text graph and label graph, and combining the graph attention mechanism, the accuracy and generalization of short text classification are improved. By using optimization algorithms such as Adam and RMSprop, the learning rate can be dynamically adjusted according to the historical information of the gradient, enabling the model to find the optimal solution faster during the training process, greatly accelerating the convergence process of the model and improving the training efficiency. At the same time, by introducing regularization techniques such as L1 and L2 regularization, the model parameters are constrained to prevent the model from overfitting the training data and enhancing the generalization ability of the model. In addition, by constructing an efficient text graph and label graph, the semantic information in the short text and the correlation information between labels can be fully mined. Combining the graph attention mechanism enables the model to better capture the complex relationship between the text and the label, thereby improving the accuracy of short text classification.

[0076] The choice of the optimization algorithm will be adjusted according to the characteristics of the data and the complexity of the model. For example, for the case where the data scale is large and the gradient changes violently, the Adam optimization algorithm will be selected. It can adaptively adjust the learning rate and estimate the second moment of the gradient, making the model more stable during the training process. For the case where the data scale is small and the gradient changes gently, the RMSprop optimization algorithm may be selected. It can adjust the learning rate according to the weighted average of the square of the gradient, avoiding the learning rate being too large or too small. When introducing the regularization technique, the regularization coefficient is determined by the method of cross-validation. The data set is divided into a training set, a validation set, and a test set. The model is trained on the training set, the regularization coefficient is adjusted on the validation set, and the coefficient with the best performance on the validation set is selected as the final regularization coefficient. When constructing an efficient text graph and label graph, multiple strategies will be adopted. For the text graph, in addition to using a sliding window to construct edges, factors such as the semantic similarity and co-occurrence frequency of words will also be considered, so that the text graph can more accurately reflect the semantic structure inside the text. For the label graph, edges will be constructed according to the hierarchical relationship and co-occurrence relationship of the labels, so that the label graph can more comprehensively reflect the correlation information between the labels. When combining the graph attention mechanism, the attention mechanism will be improved, such as introducing multi-head attention, self-attention and other mechanisms to further enhance the model's ability to extract text and label information. At the same time, during the model training process, an early stopping strategy will also be adopted. When the performance of the model on the validation set no longer improves, the training is stopped in advance to avoid model overfitting.

[0077] In the text preprocessing step, short text data is comprehensively processed. Duplicate, similar data, and stop words are removed through the data cleaning sub-step to reduce noise interference. Then, in the word segmentation and part-of-speech tagging sub-step, natural language processing techniques are used for word segmentation and part-of-speech tagging, and nouns, verbs, adjectives, etc. that are significant for classification are selected and retained as graph nodes. Next, in the text embedding sub-step, the WordEmbedding technique is adopted to convert each word into a fixed-length vector representation, forming a text vector matrix, which provides a basis for subsequent graph construction.

[0078] Secondly, enter the graph construction step. On the one hand, in the text graph construction sub-step, a sliding window is used to construct edges on the text vector matrix. If two words are in the same window, an undirected edge is added to form a text graph, thereby capturing the local associations between words in the text. On the other hand, in the label graph construction sub-step, according to the known label relationships, if there is an association between labels, an undirected edge is added to construct a label graph to clarify the internal connections between labels.

[0079] Then, in the graph attention mechanism learning step, the graph attention mechanism is applied to the nodes in the text graph and the label graph. The text graph attention layer calculates the attention coefficients according to a specific formula. Through a learnable attention vector, a trainable transformation matrix, and an activation function, the node features are transformed and the attention weights are calculated, and then the feature weighted sum is performed to obtain the updated node embedding representation. This process can automatically focus on important word information in the text. At the same time, a similar operation is also carried out in the label graph attention layer to update the label node embedding representation. Moreover, the multi-head attention mechanism extension step is introduced. The multi-head attention mechanism is adopted in the text graph attention layer and the label graph attention layer, the attention coefficient calculation formula is modified, and the attention head index is introduced. Finally, the node embedding representation is the average of the outputs of all attention heads, so as to enrich the extraction ability of the model and capture text and label information from different perspectives.

[0080] Then, in the text-label interaction step, first, the text embedding matrix output by the text graph attention layer is dot-multiplied with the label embedding matrix output by the label graph attention layer to obtain a text-label interaction matrix, realizing the preliminary association between text and labels. Then, a fully connected layer is used to perform dimensionality reduction processing on the interaction matrix to obtain the final text-label vector and extract more representative features.

[0081] After that, enter the model training and prediction step. The cross-entropy loss function is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true labels. The model parameters are continuously updated through the backpropagation algorithm until the loss value converges, completing the model training. For the short text to be predicted, after the same preprocessing operations as in the training stage, it is input into the trained model to obtain the predicted label.

[0082] Finally, the model optimization step is also included in the whole process. During the model training process, optimization algorithms such as the Adam optimizer are adopted to accelerate the convergence process and improve the training efficiency. At the same time, regularization techniques such as L2 regularization are introduced to constrain the model parameters, prevent the model from overfitting, and enhance the generalization ability of the model. This method is especially suitable for processing short text data on social media. By constructing an efficient text graph and label graph and combining the graph attention mechanism, the accuracy and generalization of short text classification are effectively improved.

[0083] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0084] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A short text classification processing method using a graph attention network, characterized in that: The following steps are involved: S10, text preprocessing step: cleaning, word segmentation and part-of-speech tagging of short text data, removing redundant words, and using word embedding technology to convert each word into a vector representation of a fixed length to obtain a text vector matrix; S20, graph construction step: constructing a text graph based on the text vector matrix, and constructing a label graph according to known label relationships; S30, graph attention mechanism learning step: apply the graph attention mechanism to the nodes in the text graph and the label graph, calculate the attention coefficients between the nodes, and perform weighted summation of the features to obtain the updated node embedding representation; S40, text label interaction step: interactively processing the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain a text label interaction matrix; S50, model training and prediction steps: use the cross entropy loss function to train the model, update the model parameters through the back propagation algorithm until the loss value converges, pre-process the short text to be predicted and input it into the trained model to obtain the predicted label.

2. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The text preprocessing step S10 specifically includes: S11, data cleaning sub-step: removing duplicate data, similar data and stop words in short text data; S12, word segmentation and part-of-speech tagging sub-step: Use natural language processing technology to segment the short text and perform part-of-speech tagging, remove redundant conjunctions, auxiliary words, and adverbs, and retain nouns, verbs, adjectives, and adverbs as graph nodes; S13, text embedding sub-step: Use WordEmbedding technology to convert each word into a vector representation of a fixed length to form a text vector matrix.

3. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The step S20, graph construction, specifically includes: S21, text graph construction sub-step: use a sliding window to build edges on the text vector matrix. If two words are in the same window, an undirected edge is added to form a text graph; S22, label graph construction sub-step: construct a label graph based on the known label relationship, and add an undirected edge if there is a relationship between the labels.

4. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The text graph attention layer in the graph attention mechanism learning step specifically uses the following formula to calculate the attention coefficient and update the node embedding representation: e ij =LeakyReLU(a T [Wh i ||Wh j ]) h′ i =σ(∑ j∈N(i) a ij Wh j ) Among them, a is the learnable attention vector, W is the trainable transformation matrix, and σ is the activation function.

5. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The step of S40, text label interaction, specifically includes: S41, interaction matrix calculation sub-step: perform dot multiplication on the text embedding matrix output by the text graph attention layer and the label embedding matrix output by the label graph attention layer to obtain a text label interaction matrix; S42, dimensionality reduction processing sub-step: Use the fully connected layer to reduce the dimensionality of the interaction matrix to obtain the final text label vector.

6. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The loss function in the model training and prediction steps adopts the cross entropy loss function, and the specific formula is: Where N is the number of samples, C is the number of categories, and y ic is the true label, p ic is the predicted probability.

7. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: It also includes a multi-head attention mechanism expansion step, which introduces a multi-head attention mechanism in the text graph attention layer and the label graph attention layer to enrich the extraction capability of the model and improve the classification accuracy.

8. According to claim 7, a short text classification processing method using a graph attention network is characterized in that: The attention coefficient calculation formula in the multi-head attention mechanism expansion step is modified as follows: Here, h represents the index of the attention head, and the final node embedding is represented as the average of the outputs of all attention heads.

9. According to claim 1, a short text classification processing method using a graph attention network is characterized in that: The method also includes a model optimization step, in which an optimization algorithm is used to accelerate the convergence process during model training, and a regularization technique is introduced to prevent overfitting. When processing short text data from social media, the method constructs an efficient text graph and label graph and combines it with a graph attention mechanism to improve the accuracy and generalization of short text classification.

Citation Information

Cited By

  • Implementation method and device of disaster information early warning and storage medium

    CN121121956A