Short Text Classification Method and System Based on Multimodal Feature Fusion and Graph Convolution
Through the methods of multimodal feature fusion and graph convolution, text graphs are constructed and model parameters are optimized, which solves the problems of inaccurate classification and insufficient generalization ability in short text classification, and realizes accurate short text classification.
Patent Information
- Application Number
- CN202310007806.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-04
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-01-04
AI Technical Summary
The prior art has problems of classification inaccuracy and insufficient generalization ability of classification models in short text classification, especially the problems of semantic features and gradient vanishing/explosion caused by sparseness and ambiguity of short text data.
The multimodal feature fusion and graph convolution methods are adopted to construct text graphs, including document nodes, word nodes and edges, and text vector quantum model and graph convolution network sub-model, combined with cross entropy loss function to optimize model parameters, strengthen document node feature representation and jump fusion features to avoid gradient vanishing and explosion.
It effectively improves the accuracy of short text classification and generalization ability of classification models, solves the problems of semantic ambiguity and poor classification effect caused by scarce short text data features and insufficient utilization of label information, and realizes accurate category prediction.
Smart Images

Figure CN115982361B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer natural language processing, and more specifically, to a short text classification method and system based on multi-modal feature fusion and graph convolution. Background Art
[0002] With the booming development of numerous social media platforms, the short text data on many platforms has grown exponentially, and it is very necessary to correctly classify the massive data; classifying short texts can help various platforms and users on the network efficiently process text data, such as news classification, user intention analysis, etc. Since this information is generally relatively short, resulting in sparse semantic features, accurate classification cannot be achieved during the classification process. In addition, most current classification methods are more suitable for long texts, and when applied to short texts, the classification effect is poor.
[0003] Currently, there are two classification methods for short texts. One is to input the short text to be classified into a pre-trained deep neural network model and use the output as the final classification result; the other is to use an external knowledge base to expand the text content to improve the classification accuracy. However, due to the sparsity and ambiguity of short text data, traditional deep neural networks cannot fully capture the importance of the keywords in short text data, and the class label information of the text is often ignored; moreover, in model training, if the network is too deep and the network weight update is unstable, problems such as gradient disappearance and gradient explosion will occur; for the technology of using an external knowledge base to expand short texts, although it expands the short text data, it also increases redundant features, and the classification accuracy is easily affected by the quality of the external knowledge base.
[0004] The prior art discloses a Chinese short text classification method based on a graph attention network, including the following steps: preprocessing text data to obtain a set of word lists corresponding to the text; text feature extraction: performing word embedding processing on the set of word lists corresponding to the text using a feature embedding tool to obtain corresponding word vectors; constructing a graph using a graph structure, taking the text and the words in the text as graph nodes to construct a heterogeneous graph; establishing a graph attention network text classification model; using an open-source Chinese short text dataset with category annotations as the training language dataset, and training the graph attention network text classification model using the heterogeneous graph; outputting the category to which the text belongs: obtaining the final classified category through a softmax classification layer for the node features; this application ignores the importance of words and the label information of the text, and there is a problem of weak expression of document node features, resulting in inaccurate final classification categories. Summary of the Invention
[0005] To overcome the defect of inaccurate short text classification in the above-mentioned prior art, the present invention provides a short text classification method and system based on multi-modal feature fusion and graph convolution, which effectively improves the accuracy of short text classification and the generalization ability of the classification model, and avoids gradient explosion and disappearance.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] The present invention provides a short text classification method based on multi-modal feature fusion and graph convolution, including:
[0008] S1: Obtain short text data and corresponding labels, and preprocess the short text data to obtain preprocessed short text data;
[0009] S2: Convert the preprocessed short text data into a text graph, where the text graph includes document nodes, word nodes, document-word edges, and word-word edges;
[0010] S3: Build a short text classification model, including a text vector sub-model and a graph convolution network sub-model;
[0011] S4: Input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial features of the document nodes and text features; and embed the initial features of the document nodes into the text graph to obtain an optimized text graph;
[0012] S5: Input the optimized text graph into the graph convolution network sub-model to obtain the final features of the text graph;
[0013] S6: Fuse the text features and the final features of the text graph to obtain the final classification features, and calculate the predicted classification probability of the short text data;
[0014] S7: Set a cross-entropy loss function to optimize the short text classification model, adjust the model parameters of the short text classification model, and obtain an optimized short text classification model;
[0015] S8: Obtain the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
[0016] Preferably, in the step S1, the preprocessing performed on the short text data includes denoising, duplicate removal, special symbol removal, and stop word removal operations.
[0017] Preferably, in the step S2, the specific method for converting the preprocessed short text data into a text graph is:
[0018] Take all the documents in the preprocessed short text data as document nodes, and all the words in the preprocessed short text data as word nodes; use the PMI algorithm to calculate the edge weights between words, realize the connection between words, and obtain the word-word edges; use the improved word frequency statistics algorithm to calculate the edge weights between words and documents, realize the connection between words and documents, and obtain the document-word edges.
[0019] The word-word edges are calculated by the traditional PMI algorithm. If two words always appear in the same text, then these two words are considered relevant, and the edge weight value between the two words can be set according to the relationship between the frequencies of the two words and the total frequency of all word pairs; the traditional word frequency method cannot effectively reflect the importance of words and the distribution of feature words, so the improved word frequency statistics algorithm is used to weight the weight of words in the document by using the word frequency relationship between words and the global, reducing the influence of the same type of text in the corpus on the word weight and the problem of too small weight value.
[0020] Preferably, the specific method for using the PMI algorithm to calculate the edge weights between words, realize the connection between words, and obtain the word-word edges is as follows:
[0021]
[0022] In the formula, PMI(t i , t j ) represents the edge weight between word t i and word t j , p(t i , t j ) represents the probability that word t i and word t j appear in the same text, p(t i ) represents the probability that word t i appears, and p(t i ) represents the probability that word t j appears.
[0023] Preferably, the specific method for using the improved word frequency statistics algorithm to calculate the edge weights between words and documents, realize the connection between words and documents, and obtain the document-word edges is as follows:
[0024]
[0025] In the formula, TF-IDF-Pro(t i , d j ) represents the edge weight between word t i and document d j , represents word t i in document d jThe number of occurrences in represents document d j the frequency of all words in; |D| represents the number of documents in the preprocessed short text data, |{j: t i ∈d j}| represents the number of documents containing document d j in which the word t i ; M represents the total number of words in the preprocessed short text data, represents the word t i the total number of occurrences in the preprocessed short text data.
[0026] Preferably, in the step S4, the specific method for obtaining the initial features of the document node is:
[0027] Represent the text contained in the document node as a vector set W i ={w1, w2,..., w p}), where W i represents the vector set of the i-th text in the document, w p represents the word vector at the p-th position of the i-th text; the feature vector set of the short text data label is represented as Y={L1,..., L i ,..., L Q}, where L i represents the label feature vector of the i-th text in the document, and the feature vector of the short text data label corresponds one-to-one with the vector of the text;
[0028] Calculate the text category features:
[0029]
[0030] In the formula, LE[i] represents the text category feature of the i-th text in the document, WE[t j , v i represents the feature vector of the word at the j-th position of the i-th text, similarity(*) represents the vector similarity calculation function, and k represents the preset similarity threshold;
[0031] Calculate the initial features of the document node:
[0032]
[0033] In the formula, H R represents the initial features of the document node, and N R represents the number of words participating in the aggregation.
[0034] To strengthen the expression of document nodes, first, the feature vectors of each piece of text and text tags are collected, and then they are introduced into a common feature space for similarity calculation; through the feature selection operation by setting a similarity threshold, the word features with near-synonym properties in the text that exceed the similarity threshold condition with the tags are retained, reducing the influence of redundant features and achieving the effect of strengthening the expression of features; finally, the document nodes are assigned by accumulating the feature vectors and taking the mean, and embedded into the corresponding text graph; when the tag consists of more than one word, the word pair can be split at this time, and they are respectively used as the features of the tag to participate in the calculation in the feature space.
[0035] Preferably, in the step S5, the graph convolutional network sub-model includes a first graph convolutional layer and a second graph convolutional layer;
[0036] The optimized text graph is input into the graph convolutional network sub-model, and the original feature matrix of the optimized text graph is input into the first graph convolutional layer for information propagation, adaptively learning the feature information with an edge distance of 1, and obtaining the first text graph feature:
[0037]
[0038] In the formula, L () represents the first text graph feature, L () represents the original feature matrix of the optimized text graph, represents the normalized Laplacian matrix of the optimized text graph, W () represents the weight matrix of the first graph convolutional layer, and ReLU(*) represents the activation function; where:
[0039]
[0040] In the formula, D represents the degree matrix of the optimized text graph, A represents the adjacency matrix of the optimized text graph, and I represents the identity matrix;
[0041] To solve the problems of gradient disappearance, explosion, and feature loss during network training, using the method of fusing features in a skip manner, after combining the first text graph feature and the original feature matrix of the optimized text graph, it is input into the second graph convolutional layer for information propagation, adaptively learning the feature information with an edge distance of 2, and obtaining the final text graph feature:
[0042]
[0043] In the formula, Z G represents the final text graph feature, W () represents the weight matrix of the second graph convolutional layer.
[0044] Preferably, in the step S6, the specific method for fusing the text features and the final text graph feature to obtain the final classification feature is:
[0045] R = Z G *+(1-)* R
[0046] Wherein, R represents the final classification feature, ε represents the balance parameter, and Z R represents the text feature.
[0047] Preferably, in the step S6, the specific method for calculating the predicted classification probability of the short text data is:
[0048]
[0049] Wherein, softmax(R i ) represents the predicted classification probability of the i-th document node in the short text data, softmax(*) represents the activation function, exp(*) represents the exponential function, j represents the number of classification categories of the short text data, and R i represents the final classification feature of the i-th document node in the output data, and R j represents the final classification feature of the j-th document node in the output data.
[0050] Preferably, in the step S7, the cross-entropy loss function is specifically:
[0051]
[0052] Wherein, Loss represents the cross-entropy loss value, y N represents the preprocessed short text data set, m represents the feature dimension, M represents the total number of feature dimensions, and ln(*) represents the logarithmic function.
[0053] After calculating the loss value through the cross-entropy loss function, by calculating the gap between the forward calculation result of each iteration and the true value, the weight parameters in the training process of the short text classification model are adjusted, so that the next training proceeds in the correct direction.
[0054] The present invention also provides a short text classification system based on multi-modal feature fusion and graph convolution for implementing the above-mentioned short text classification method based on multi-modal feature fusion and graph convolution, including:
[0055] A data acquisition and preprocessing module, configured to acquire short text data and corresponding labels, and preprocess the short text data to obtain preprocessed short text data;
[0056] A text graph conversion module, configured to convert the preprocessed short text data into a text graph, where the text graph includes document nodes, word nodes, document-word edges, and word-word edges;
[0057] A classification model construction module for constructing a short text classification model, including a text vector sub-model and a graph convolutional network sub-model;
[0058] An initial feature acquisition module for inputting word nodes and their corresponding labels into the text vector sub-model to obtain initial document node features and text features; and embedding the initial document node features into a text graph to obtain an optimized text graph;
[0059] A text graph feature acquisition module for inputting the optimized text graph into the graph convolutional network sub-model to obtain the final text graph features;
[0060] A feature fusion module for fusing the text features and the final text graph features to obtain final classification features and calculating the predicted classification probability of the short text data;
[0061] A classification model optimization module for setting a cross-entropy loss function to optimize the short text classification model, adjusting the model parameters of the short text classification model, and obtaining an optimized short text classification model;
[0062] A short text classification module for obtaining short text data to be classified, inputting the optimized short text classification model, obtaining the predicted classification probability of the short text data to be classified, and determining the category of the short text data to be classified.
[0063] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0064] After preprocessing the short text data, this application converts it into a text graph, including document nodes, word nodes, document-word edges, and word-word edges; constructs a short text classification model containing a text vector sub-model and a graph convolutional network sub-model, inputs the word nodes and corresponding labels into the text vector sub-model, combines the near-synonym features existing in the text features and label information, strengthens the feature representation of the document nodes in the text graph by using the label information, effectively reflects the importance of words, solves the problem of weak feature expression of document nodes, reduces the impact of feature smoothing, obtains the initial features of document nodes and text features, and embeds the initial features of document nodes into the text graph; then inputs the optimized text graph into the graph convolutional network sub-model, uses the method of jump-style fusion features to obtain the final features of the text graph, further reduces the loss of initial features, and can also avoid the problems of gradient disappearance and explosion existing in the training process of the graph convolutional network; finally, fuses the text features and the final features of the text graph to obtain the final classification features, calculates the predicted classification probability of the short text data, sets the cross-entropy loss function, and optimizes the short text classification model. It combines the diverse features learned by two different models, namely the text vector sub-model and the graph convolutional network sub-model, effectively improves the classification quality of short text data and the generalization ability of the text vector sub-model; uses the optimized short text classification model to predict the category of short text data, effectively solves the problems of semantic ambiguity and poor classification effect caused by scarce features of short text data and insufficient utilization of label information, and obtains accurate categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 FIG. is a flowchart of the short text classification method based on multi-modal feature fusion and graph convolution described in Embodiment 1.
[0066] Figure 2 FIG. is a schematic diagram of obtaining the initial features of document nodes described in Embodiment 2.
[0067] Figure 3 FIG. is a schematic diagram of the structure of the short text classification system based on multi-modal feature fusion and graph convolution described in Embodiment 3. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0069] For better illustration of this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the actual size of the product;
[0070] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0071] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0072] Example 1
[0073] This example provides a short text classification method based on multi-modal feature fusion and graph convolution, as Figure 1 shown, including:
[0074] S1: Obtain short text data and corresponding labels, preprocess the short text data to obtain preprocessed short text data;
[0075] S2: Convert the preprocessed short text data into a text graph, where the text graph includes document nodes, word nodes, document-word edges, and word-word edges;
[0076] S3: Build a short text classification model, including a text vector sub-model and a graph convolution network sub-model;
[0077] S4: Input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial document node features and text features; and embed the initial document node features into the text graph to obtain an optimized text graph;
[0078] S5: Input the optimized text graph into the graph convolution network sub-model to obtain the final text graph features;
[0079] S6: Fuse the text features and the final text graph features to obtain the final classification features, and calculate the predicted classification probability of the short text data;
[0080] S7: Set the cross-entropy loss function, optimize the short text classification model, adjust the model parameters of the short text classification model to obtain an optimized short text classification model;
[0081] S8: Obtain the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
[0082] In the specific implementation process, after preprocessing the short text data, this embodiment converts it into a text graph, including document nodes, word nodes, document-word edges, and word-word edges; constructs a short text classification model containing a text vector sub-model and a graph convolutional network sub-model, inputs the word nodes and corresponding labels into the text vector sub-model, combines the near-synonym features existing in the text features and label information, uses the label information to strengthen the feature representation of the document nodes in the text graph, effectively reflects the importance of words, solves the problem of weak feature expression of document nodes, reduces the influence of feature smoothing, obtains the initial features of document nodes and text features, and embeds the initial features of document nodes into the text graph; then inputs the optimized text graph into the graph convolutional network sub-model, uses the method of jump fusion features to obtain the final features of the text graph, further reduces the loss of initial features, and can also avoid the problems of gradient disappearance and explosion during the training of the graph convolutional network; finally, fuses the text features and the final features of the text graph to obtain the final classification features, calculates the predicted classification probability of the short text data, sets the cross-entropy loss function, and optimizes the short text classification model. It combines the diverse features learned by two different models, namely the text vector sub-model and the graph convolutional network sub-model, effectively improves the classification quality of short text data and the generalization ability of the text vector sub-model; uses the optimized short text classification model to predict the category of short text data, effectively solves the problems of semantic ambiguity and poor classification effect caused by scarce features of short text data and insufficient utilization of label information, and obtains accurate categories.
[0083] Embodiment 2
[0084] This embodiment provides a short text classification method based on multi-modal feature fusion and graph convolution, including:
[0085] S1: Obtain short text data and corresponding labels, preprocess the short text data, and obtain the preprocessed short text data;
[0086] The preprocessing includes operations of denoising, duplicate removal, removing special symbols, and removing stop words;
[0087] S2: Convert the preprocessed short text data into a text graph, where the text graph includes document nodes, word nodes, document-word edges, and word-word edges;
[0088] Take all documents in the preprocessed short text data as document nodes, and take all words in the preprocessed short text data as word nodes; use the PMI algorithm to calculate the edge weights between words to realize the connection between words and obtain word-word edges; specifically:
[0089]
[0090] where PMI(t i , t j ) represents the edge weight between word t i and word t j , p(t i , t j ) represents the probability that word t i and word t j appear in the same text, p(t i ) represents the probability that word t i appears, and p(t i ) represents the probability that word t j appears.
[0091] Calculate the edge weight between words and documents using an improved word frequency statistical algorithm, realize the connection between words and documents, and obtain the document-word edge; specifically:
[0092]
[0093] where TF-IDF-Pro(t i , d j ) represents the edge weight between word t i and document d j , represents the number of times word t i appears in document d j , represents the frequency of all words in document d j ; |D| represents the number of documents in the preprocessed short text data, |{j: t i ∈ d j}| represents the number of documents containing word t j in document d i ; M represents the total number of words in the preprocessed short text data, represents the total number of times word t i appears in the preprocessed short text data.
[0094] The word-word edge is calculated by the traditional PMI algorithm. If two words always appear in the same text, then these two words are considered related, and the edge weight value between words can be set according to the relationship between the frequencies of these two words and the total frequency of all word pairs; the traditional word frequency method cannot effectively reflect the importance of words and the distribution of feature words. Therefore, an improved word frequency statistical algorithm is used to weight the weight of words in the document using the word frequency relationship between words and the global, reducing the influence of the same type of text in the corpus on the word weight and the problem of too small weight values.
[0095] S3: Construct a short text classification model, including a text vector sub-model and a graph convolutional network sub-model;
[0096] S4: Input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial features of the document nodes and text features; and embed the initial features of the document nodes into the text graph to obtain an optimized text graph;
[0097] The specific method for obtaining the initial features of the document nodes is as follows:
[0098] Represent the text contained in the document node as a vector set W i ={w1, w2, …, w p}, where W i represents the vector set of the i-th text in the document, and w p represents the word vector at the p-th position of the i-th text; the feature vector set of the short text data labels is represented as Y = {L1, …, L i , …, L Q}, where L i represents the label feature vector of the i-th text in the document, and the feature vector of the short text data label corresponds one-to-one with the vector of the text;
[0099] Calculate the text category features:
[0100] "
[0101] In the formula, LE[i] represents the text category feature of the i-th text in the document, WE[t j , v i represents the feature vector of the word at the j-th position of the i-th text, similarity(*) represents the vector similarity calculation function, and k represents the preset similarity threshold;
[0102] Calculate the initial features of the document nodes:
[0103]
[0104] In the formula, H R represents the initial features of the document nodes, and N R represents the number of words participating in the aggregation.
[0105] To strengthen the expression of document nodes, first, the feature vectors of each piece of text and text tags are collected, and then they are introduced into a common feature space for similarity calculation; through the feature selection operation set by the similarity threshold, the word features with near-synonym properties in the text that exceed the similarity threshold condition with the tag are retained, reducing the influence of redundant features and achieving the effect of strengthening the expression of features; finally, the document nodes are assigned by accumulating the feature vectors and taking the average value, and embedded into the corresponding text graph; when the tag consists of more than one word, the word pair can be split at this time, and they are respectively used as the features of the tag to participate in the calculation in the feature space. As Figure 2 shown, x represents all word vectors in the text, y represents the corresponding tag, and m represents the dimension; both are introduced into the feature space for similarity screening, and the word features with near-synonym properties that exceed the similarity threshold are retained for feature fusion as the initial features of the document nodes.
[0106] S5: Input the optimized text graph into the graph convolutional network sub-model to obtain the final features of the text graph;
[0107] The graph convolutional network sub-model includes a first graph convolutional layer and a second graph convolutional layer;
[0108] Input the optimized text graph into the graph convolutional network sub-model, and the original feature matrix of the optimized text graph is input into the first graph convolutional layer for information propagation, adaptively learning the feature information with an edge distance of 1 to obtain the first text graph feature:
[0109]
[0110] In the formula, L () represents the first text graph feature, L () represents the original feature matrix of the optimized text graph, represents the normalized Laplacian matrix of the optimized text graph, W () represents the weight matrix of the first graph convolutional layer, and ReLU(*) represents the activation function; where:
[0111]
[0112] In the formula, D represents the degree matrix of the optimized text graph, A represents the adjacency matrix of the optimized text graph, and I represents the identity matrix;
[0113] To solve the problems of gradient disappearance, explosion, and feature loss during network training, the method of fusing features in a skip manner is used. After combining the first text graph feature and the original feature matrix of the optimized text graph, it is input into the second graph convolutional layer for information propagation, adaptively learning the feature information with an edge distance of 2 to obtain the final features of the text graph:
[0114]
[0115] In the formula, Z G represents the final feature of the text graph, and W () represents the weight matrix of the second graph convolutional layer.
[0116] S6: Fuse the text feature and the final feature of the text graph to obtain the final classification feature, and calculate the predicted classification probability of the short text data;
[0117] The final classification feature is:
[0118] R = Z G *+(1 - )* R
[0119] In the formula, R represents the final classification feature, ε represents the balance parameter, and Z R represents the text feature;
[0120] The predicted classification probability of the short text data is:
[0121]
[0122] In the formula, softmax(R i ) represents the predicted classification probability of the i-th document node in the short text data, softmax(*) represents the activation function, exp(*) represents the exponential function, j represents the number of classification categories of the short text data, and R i represents the final classification feature of the i-th document node in the output data, and R j represents the final classification feature of the j-th document node in the output data.
[0123] S7: Set the cross-entropy loss function, optimize the short text classification model, adjust the model parameters of the short text classification model, and obtain the optimized short text classification model;
[0124] The specific cross-entropy loss function is:
[0125]
[0126] In the formula, Loss represents the cross-entropy loss value, y N represents the preprocessed short text data set, m represents the feature dimension, M represents the total number of feature dimensions, and ln(*) represents the logarithmic function;
[0127] After the cross-entropy loss function calculates the loss value, by calculating the gap between the forward calculation result of each iteration and the true value, the weight parameters in the training process of the short text classification model are adjusted, so that the next training proceeds in the correct direction.
[0128] S8: Obtain the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
[0129] Embodiment 3
[0130] This embodiment provides a short text classification system based on multi-modal feature fusion and graph convolution, which is used to implement the short text classification method based on multi-modal feature fusion and graph convolution described in Embodiment 1 or 2. As Figure 3 shown, it includes:
[0131] A data acquisition and preprocessing module, which is used to obtain short text data and corresponding labels, preprocess the short text data, and obtain the preprocessed short text data;
[0132] A text graph conversion module, which is used to convert the preprocessed short text data into a text graph. The text graph includes document nodes, word nodes, document-word edges, and word-word edges;
[0133] A classification model construction module, which is used to construct a short text classification model, including a text vector sub-model and a graph convolution network sub-model;
[0134] An initial feature acquisition module, which is used to input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial document node features and text features; and embed the initial document node features into the text graph to obtain an optimized text graph;
[0135] A text graph feature acquisition module, which is used to input the optimized text graph into the graph convolution network sub-model to obtain the final text graph features;
[0136] A feature fusion module, which is used to fuse the text features and the final text graph features to obtain the final classification features, and calculate the predicted classification probability of the short text data;
[0137] A classification model optimization module, which is used to set the cross-entropy loss function, optimize the short text classification model, adjust the model parameters of the short text classification model, and obtain an optimized short text classification model;
[0138] A short text classification module, which is used to obtain the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
[0139] The same or similar reference numerals correspond to the same or similar components;
[0140] The terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0141] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A short text classification method based on multi-modal feature fusion and graph convolution, characterized in that Including: S1: Obtain short text data and corresponding labels, preprocess the short text data, and obtain the preprocessed short text data; S2: Convert the preprocessed short text data into a text graph, where the text graph includes document nodes, word nodes, document-word edges, and word-word edges; The specific method for converting the preprocessed short text data into a text graph is: Take all documents in the preprocessed short text data as document nodes, and take all words in the preprocessed short text data as word nodes; Use the PMI algorithm to calculate the edge weights between words, realize the connection between words, and obtain word-word edges; Use an improved word frequency statistics algorithm to calculate the edge weights between words and documents, realize the connection between words and documents, and obtain document-word edges; The specific method for using an improved word frequency statistics algorithm to calculate the edge weights between words and documents, realize the connection between words and documents, and obtain document-word edges is: Where, TF_IDF_Pro(t i ,d j ) represents the edge weight of word t i and document d j ; represents the number of times word t i appears in document d j ; represents the frequency of all words in document d j ; |D| represents the number of documents in the preprocessed short text data, |{j:t i ∈d j}| represents the number of documents containing word t j in document d i ; N represents the total number of words in the preprocessed short text data, represents the total number of times word t i appears in the preprocessed short text data; S3: Construct a short text classification model, including a text vector sub-model and a graph convolutional network sub-model; S4: Input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial document node features and text features; And embed the initial document node features into the text graph to obtain an optimized text graph; S5: Input the optimized text graph into the graph convolutional network sub-model to obtain the final text graph features; S6: Fuse the text features and the final text graph features to obtain the final classification features, and calculate the predicted classification probability of the short text data; S7: Set the cross-entropy loss function, optimize the short text classification model, adjust the model parameters of the short text classification model, and obtain an optimized short text classification model; S8: Obtain the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
2. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 1, wherein The specific method for using the PMI algorithm to calculate the edge weights between words, realize the connection between words, and obtain word-word edges is: where PMI(t i ,t j ) represents the edge weight between word t i and word t j , p(t i ,t j ) represents the probability that word t i and word t j appear in the same text, p(t i ) represents the probability that word t i appears, and p(t i ) represents the probability that word t j appears.
3. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 1 or 2, characterized in that, In step S4, the specific method for obtaining the initial document node features is: Represent the text included in the document node as a vector set W i ={w1, w2, …, w p}, where W i represents the vector set of the i-th text in the document, and w p represents the word vector at the p-th position of the i-th text; the feature vector set of the short text data label is represented as Y = {L1, …, L i , …, L Q}, where L i represents the label feature vector of the i-th text in the document, and the feature vector of the short text data label corresponds one-to-one with the vector of the text; Calculate the text category features: where LE[i] represents the text category feature of the i-th text in the document, and WE[t j ,v i represents the feature vector of the word at the j-th position of the i-th text, similarity(*) represents the vector similarity calculation function, and k represents the preset similarity threshold; Calculate the initial document node features: Where H R represents the initial feature of the document node, and N R represents the number of words participating in the aggregation.
4. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 3, wherein In step S5, the graph convolutional network sub-model includes a first graph convolutional layer and a second graph convolutional layer; Input the optimized text graph into the graph convolutional network sub-model, and input the original feature matrix of the optimized text graph into the first graph convolutional layer for information propagation to learn the feature information with an edge distance of 1, and obtain the first text graph features: Wherein, L (1) represents the first text graph feature, and L (0) represents the original feature matrix of the optimized text graph, represents the normalized Laplacian matrix of the optimized text graph, W (1) represents the weight matrix of the first graph convolutional layer, and ReLU(*) represents the activation function; wherein: Where D represents the degree matrix of the optimized text graph, A represents the adjacency matrix of the optimized text graph, and I represents the identity matrix; After combining the first text graph features and the original feature matrix of the optimized text graph, input them into the second graph convolutional layer for information propagation to learn the feature information with an edge distance of 2, and obtain the final text graph features: where Z G represents the final feature of the text graph, and W (2) represents the weight matrix of the second graph convolutional layer.
5. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 4, characterized in that In step S6, the specific method for fusing the text features and the final text graph features to obtain the final classification features is: R = Z G *ε+(1 - ε)*Z R wherein, R represents the final classification feature, ε represents the balance parameter, and Z R represents the text feature.
6. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 5, wherein In step S6, the specific method for calculating the predicted classification probability of the short text data is: where softmax(R i ) represents the predicted classification probability of the i-th document node in the short text data, softmax(*) represents the activation function, exp(*) represents the exponential function, j represents the number of classification categories of the short text data, R i represents the final classification feature of the i-th document node in the output data, and R j represents the final classification feature of the j-th document node in the output data.
7. The short text classification method based on multi-modal feature fusion and graph convolution according to claim 6, wherein In step S7, the cross-entropy loss function is specifically: where Loss represents the cross-entropy loss value, y N represents the set of short text data after preprocessing, m represents the feature dimension, M represents the total number of feature dimensions, and ln(*) represents the logarithmic function.
8. A short text classification system based on multi-modal feature fusion and graph convolution, which is used to implement the short text classification method based on multi-modal feature fusion and graph convolution described in any one of claims 1-7, characterized in that, Including: A data acquisition and preprocessing module, which is used to acquire short text data and corresponding labels, preprocess the short text data, and obtain the preprocessed short text data; A text graph conversion module, which is used to convert the preprocessed short text data into a text graph, and the text graph includes document nodes, word nodes, document-word edges, and word-word edges; A classification model construction module, which is used to construct a short text classification model, including a text vector sub-model and a graph convolutional network sub-model; An initial feature acquisition module, which is used to input the word nodes and their corresponding labels into the text vector sub-model to obtain the initial features of the document nodes and text features; and embed the initial features of the document nodes into the text graph to obtain an optimized text graph; A text graph feature acquisition module, which is used to input the optimized text graph into the graph convolutional network sub-model to obtain the final features of the text graph; A feature fusion module, which is used to fuse the text features and the final features of the text graph to obtain the final classification features, and calculate the predicted classification probability of the short text data; A classification model optimization module, which is used to set a cross-entropy loss function, optimize the short text classification model, adjust the model parameters of the short text classification model, and obtain an optimized short text classification model; A short text classification module, which is used to acquire the short text data to be classified, input it into the optimized short text classification model, obtain the predicted classification probability of the short text data to be classified, and determine the category of the short text data to be classified.
Citation Information
Patent Citations
Short text classification method, device and equipment based on feature extension
CN109960730A
Chinese short text classification method fusing contextual information graph convolution
CN114048754A