A news classification method based on multimodal feature fusion

Through the multimodal feature fusion method, LSTM, attention network and cross-attention network are used to pre-train news text and image information on a large-scale multimodal dataset, which solves the problem of insufficient utilization of image information in the existing technology and improves the accuracy and credibility of news classification.

CN115588122BActive Publication Date: 2025-09-12CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211383002.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-09-12
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

Existing news classification models fail to fully utilize the image information in news, resulting in a decrease in classification accuracy. In particular, visual extractors that have not been pre-trained on large-scale multimodal datasets are prone to introduce information bias in downstream interactions.

Method used

A multimodal feature fusion method is adopted to pre-train news text vectors and image sequence vectors on a large-scale multimodal dataset through LSTM, attention network, gated memory network and cross-attention network. The text and image information are fused using the cross-attention network and classified using the softmax function.

Benefits of technology

It improves the accuracy and credibility of news classification and ensures the accuracy and consistency of information in downstream interaction processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115588122B_ABST
    Figure CN115588122B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of text classification, and specifically relates to a news classification method based on multimodal feature fusion, comprising: obtaining original news sample data; performing feature extraction on the original news text to obtain an original news text vector, and performing feature extraction on each original news illustration to obtain an image sequence vector of each original news illustration; inputting the original news text vector and the image sequence vectors of all original news illustrations into a news classification model for training; obtaining target news sample data, obtaining a target news text vector and image sequence vectors of multiple target news illustrations, and inputting the target news text vector and the image sequence vectors of the multiple target news illustrations into the news classification model to obtain classification results of the target news sample data. The present invention classifies news uploaded by users to social platforms by performing feature extraction on news text and illustrations in the news, so that the classification results have higher accuracy and credibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of text classification, and in particular to a news classification method based on multimodal feature fusion. Background Art

[0002] With the development of Internet technology, more and more social platforms are widely used. Users often browse current real-time news through these platforms. Due to user preferences, social platforms often divide news into multiple categories. According to different user preferences, it is convenient for social platforms to push news of corresponding categories to users. How to classify the news uploaded by users on social platforms has become a current research hotspot.

[0003] Most existing models for news text classification focus on a single text modality, such as the convolutional neural network-based TextCNN and TextGCN models, the recurrent neural network-based Bi-LSTM and Bi-LSTM-Attention models, and, in recent years, various pre-trained and fine-tuned BERT models. These primarily extract features from news text and classify news based on these features. However, these models ignore the information contained in the visual modality of news. When the text lacks obvious keyword information, models based solely on the text modality struggle to classify. In recent years, some multimodal news classification models have emerged. These primarily use an attention mechanism to fuse the title, paragraph, and image information and then concatenate them as the fused result. However, this approach fails to fully utilize the full range of image information in the news, and the visual extractor that extracts image information is not pre-trained on large-scale multimodal datasets. This can lead to information bias in downstream interactions, thus affecting the accuracy of news classification. Summary of the Invention

[0004] In order to solve the problems in the prior art of failing to fully utilize all the image information in news and the fact that the visual extractor for extracting image information has not been pre-trained on a large-scale multimodal dataset, which easily leads to information bias in the downstream interaction process and thus affects the accuracy of news classification, the present invention provides a news classification method based on multimodal feature fusion, comprising:

[0005] S1: Obtain original news sample data; the original news sample data includes: original news text and multiple original news pictures; label the original news sample data; and divide each original news picture into p image blocks of the same scale to obtain an image set of each original news picture;

[0006] S2: Extract features from the original news text to obtain the original news text vector, and extract features from the image set of each original news picture to obtain the image sequence vector of each original news picture;

[0007] S3: Build a news classification model, which includes: LSTM, attention network, gated memory network, cross-attention network, and softmax function;

[0008] S4: Use the original news text vectors and all the image sequence vectors of the original news illustrations as training samples to train the news classification model;

[0009] S5: Obtain target news sample data, perform feature extraction on the target news sample data to obtain a target news text vector and a plurality of target news picture sequence vectors, and input the target news text vector and the plurality of target news picture sequence vectors into a news classification model to obtain a classification result of the target news sample data.

[0010] The present invention has at least the following beneficial effects

[0011] The present invention performs feature extraction on the text in the news, and the extracted news text vector has the text information of the news. By dividing each picture of the news into P image blocks, the image block of each original news picture is characterized to obtain the picture sequence vector of each news picture, and the news picture sequence vector and the news text vector are used as training samples to train the news classification model, so that the extracted picture information and text information are both pre-trained on a large-scale multimodal data set, avoiding the information deviation that is easily caused in the downstream interaction process, and improving the accuracy of news classification. The present invention adopts the LSTM algorithm in the news classification model to calculate the hidden information in the pictures, and adds weight information to each picture through the attention network, which can reflect the importance of each picture. By combining the text information and picture information of the news across the attention network and inputting them into the softmax function for classification, the obtained news classification result has higher credibility and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flow chart of the method of the present invention;

[0013] Figure 2 This is a system block diagram of the news classification model of the present invention. DETAILED DESCRIPTION

[0014] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0015] See also Figure 1 and Figure 2 The present invention provides a news classification method based on multimodal feature fusion, comprising:

[0016] S1: Get the original news sample data (T, {a1, ...a m ,...,a M The original news sample data includes: original news text T and multiple original news pictures {a1, ...a m ,...,a M}; label the original news sample data; and divide each original news picture into p image blocks with the same R degree to obtain the image set B of each original news picture m ={b m1 ,...,b mn ,...,b mp}, where a m represents the mth original news picture in the original news sample data; M represents the number of original news pictures in the original news sample data; b mn Indicates original news picture a m The nth image block;

[0017] The label information is the category of the original news sample data. The original news text T and each original news picture carry label information. For example, the categories of the original news sample data include: military, entertainment, education, current affairs, weather, epidemic prevention, etc. In the present invention, the original news sample data of each category is obtained from the social platform through the API interface provided by the social APP; the number of original news sample data of each category is the same, and in the present invention, the number of original news sample data of each category is 1000.

[0018] S2: Extract features from the original news text T to obtain the original news text vector T′∈R s×d ; For each original news picture image set B m Perform feature extraction to obtain the image sequence vector e of each original news picture m ∈Rp×d ; s represents the number of tokens in the original news text; p represents the number of patches in the mth original news picture; e m The image sequence vector representing the mth original news picture.

[0019] The original news text T is extracted by the Deberta model to obtain the original news text vector T′∈R s ×d ;

[0020] The Chinese CLIP model is used to extract features from the image set B of the original news illustrations to obtain the image sequence vector e of the original news illustrations. m ∈R p×d ;

[0021] The Deberta model proposes two major improvements based on the BERT model, namely DisentangledAttention decoding attention and Decoding-enhanced decoding enhancement, which further enhance the encoding effect.

[0022] The Chinese CLIP model uses the Dual Encoder method of the English CLIP model to encode images and text separately, and performs large-scale unsupervised pre-training through similarity calculations. The difference is that the text data is replaced by Chinese.

[0023] S3: Construct a news classification model, which includes: a long short-term memory network (LSTM), an attention network, a gated memory network, a cross-attention network and a softmax function; the present invention performs feature extraction on the text in the news, and the extracted news text vector has the text information of the news. By dividing each picture of the news into P image blocks, the image block of each original news picture is feature-processed to obtain the image sequence vector of each news picture, and the news picture sequence vector and the news text vector are used as training samples to train the news classification model, so that the extracted image information and text information are pre-trained on a large-scale multimodal data set, avoiding the information deviation that is easily caused in the downstream interaction process, and improving the accuracy of news classification.

[0024] S4: The original news text vector T∈R s×d The image sequence vectors of all original news pictures are used as training samples to train the news classification model;

[0025] S41: Input the image sequence vectors of all original news pictures into LSTM to calculate the hidden state vector of each image sequence vector at the current moment

[0026] The structure of the neuron cell in the LSTM includes a cell state, a forget gate, an input gate, and an output gate. An internal memory unit state, namely the cell state, is defined and maintained throughout the entire cycle. The cell state is updated through the forget gate, input gate, and output gate. The calculation formula for different gates in each time step in the neuron is as follows:

[0027] The forget gate is used to implement information filtering between the current input and the hidden layer output of the previous time step. The forget gate is as follows:

[0028]

[0029] in, Refers to the hidden layer output, σ refers to the gate activation function, usually the Sigmoid function, to ensure that the output value of the forget gate is within a certain range; W f Is the current input value The associated forget gate weight value; represents the hidden state vector output at the previous time step, b f is the bias weight vector of the forget gate;

[0030] The input gate mainly needs to do two things: one is to decide which information needs to be stored and perform a second round of information update; the other is to generate a new candidate memory unit by the tanh function. Add to the state as follows:

[0031]

[0032]

[0033]

[0034] in, is the input gate output; W i 、W c are the weight values ​​of the input gate respectively; Represents the hidden state vector output at the previous time step; is the new cell state candidate value vector; b c is the weight of the new cell state candidate value vector, It is a new cell state; is the cell state at the previous time step;

[0035] The output gate determines the final output and controls the influence of the current memory cell state on the output. The sigmoid function determines the output portion of the state, while the tanh function maintains the output between -1 and 1. Multiplying the two yields the final result, as shown below:

[0036]

[0037]

[0038] in, is the output gate output of the current time step; Wo is the input value of the current time step The associated output gate weight value; b o is the bias weight vector of the output gate; is the hidden state vector of the image sequence vector of the mth original news picture at the current time step, represents 0; e m Represents the image sequence vector of the mth original news picture. When t=0,

[0039] S42: Concatenate the hidden state vectors of all the image sequence vectors at the current moment to obtain the first combination vector c of the current time step t ;

[0040] The first combination vector includes:

[0041]

[0042] Among them, c t represents the first combination vector of the current time step, is the hidden state vector of the image sequence vector of the mth original news picture at the current time step; Represents the hidden state vector of the image sequence vector of the mth original news picture at the previous time step.

[0043] S43: Concatenate the first combination vector of the current time step and the first combination vector of the previous time step to obtain the second combination vector c of the current time step [t-1;t] ;

[0044] c [t-1;t] =[c t-1 , c t ]

[0045] Among them, c [t-1;t] Represents the second combined vector of the current time step, c t-1 Represents the first combination vector of the previous time step, c t Represents the first combination vector at the current time step.

[0046] S44: Input the second combination vector of the current time step into the self-attention network to obtain the importance score matrix of the second combination vector of the current time step through the attention mechanism;

[0047]

[0048] Among them, a [t-1;t] Represents the importance score matrix of the second combination vector at the current time step, c [t-1;t] represents the second combined vector of the current time step, is the function for calculating the attention score;

[0049] The self-attention network includes: a first fully connected layer and a second fully connected layer;

[0050] The second combination vector of the current time step is input into the first fully connected layer and the second fully connected layer respectively to calculate the first fully connected feature and the second fully connected feature. Finally, the first fully connected feature and the second fully connected feature are input into the softmax function to calculate the importance score matrix of the second combination vector of the current time step;

[0051] The self-attention network is a conventional network in this field. Represents the calculation process of the self-attention network.

[0052] S45: performing a Hadamard product calculation on the second combination vector of the current time step and the importance score matrix of the second combination vector of the current time step to obtain a third combination vector of the current time step;

[0053]

[0054] Among them, a [t-1;t] Represents the importance score matrix of the second combination vector at the current time step, c [t-1;t] represents the second combined vector of the current time step, Represents the third combined vector at the current time step.

[0055] S46: Input the third combination vector of the current time step into the gated memory network to calculate the fourth combination vector of the current time step. The first combination vector, the second combination vector, the third combination vector, and the fourth combination vector are updated at each time step. When the LSTM converges, that is, at the last time step, the fourth combination vector calculated by the gated memory network is the final combination vector.

[0056] The gated memory network includes: a first MLP multi-layer perceptron, a second MLP multi-layer perceptron, a σ (sigmod) activation function, and a τ (tanh) activation function;

[0057] The step of inputting the third combination vector of the current time step into the gated memory network to calculate the fourth combination vector of the current time step includes:

[0058] S461: Input the third combined vector of the current time step into the first MLP multi-layer perceptron and then input the σ (sigmod) activation function to obtain the first weight value γ of the current time stept ;

[0059]

[0060] Among them, g1 represents the first MLP multi-layer perceptron, σ represents the (sigmod) activation function, represents the weight and bias parameters of the first MLP multilayer perceptron; γ t Represents the first weight value of the current time step. The MLP multi-layer perceptron is an existing conventional technology, and the present invention does not further introduce the MLP multi-layer perceptron.

[0061] S462: Input the third combination vector of the current time step into the second MLP multi-layer perceptron and then input the τ (tanh) activation function to obtain the updated suggestion vector of the current time step Preferably, the MLP multilayer perceptron includes a fully connected layer with 768 input and 768 output neurons respectively, a Dropout layer, and a batch normalization layer;

[0062]

[0063] Among them, g2 represents the second MLP multi-layer perceptron, represents the third combined vector of the current time step, Represents the weight and bias parameters of the second MLP multilayer perceptron; MLP multilayer perceptron is an existing conventional technology, and the present invention does not further introduce MLP multilayer perceptron, τ represents the (tanh) activation function,, Represents the updated proposal vector for the current time step.

[0064] S463: Update suggestion vector based on current time step and the first weight value γ t Calculate the fourth combination vector for the current time step:

[0065]

[0066] Among them, u t Represents the fourth combination vector of the current time step t, γ t represents the first weight value of the current time step, represents the updated proposal vector for the current time step, u t-1 Represents the fourth combination vector of the previous time step t-1.

[0067] S47: The final combined vector and the original news text vector are input into the cross-attention network to calculate the multimodal fusion vector;

[0068] The cross-attention network includes: a first sub-MLP layer (multi-layer perceptron), a second sub-multi-layer perceptron MLP layer (multi-layer perceptron), and a third sub-multi-layer perceptron MLP layer (multi-layer perceptron);

[0069] The specific method of extracting feature information in the MLP network is as follows: in the MLP input layer, the previous output is expanded into x1, x2, x3... and converted into vector X[1]. The weights between the MLP input layer and the next layer are w1, w2, w3... and converted into vector W[1], where 1 represents the weight of the first layer of MLP, and the bias b[1] is similar. The calculation of the first layer is Z[1]=W[1]X+b[1], and then A[1]=Sigmoid(Z[1]), where Z[1] is a linear combination of the inputs, and A[1] is the value obtained by the activation function Sigmoid for Z[1]. For the first layer of MLP, the input is X[1] and the output is A[1], which is the input value of the next layer, that is, X[2]=A[1], and so on to the next layer. The MLP network of the present invention is composed of five layers of fully connected networks stacked together. In order to avoid overfitting and enhance the generalization performance of the model, a random dropout layer is added between every two fully connected layers. At the same time, L2 regularization is added to each fully connected layer, and he_normal is used as the weight initialization method.

[0070] The step of inputting the final combined vector and the original news text vector into the cross-attention network to calculate the multimodal fusion vector includes:

[0071] S471: Input the final state vector into the first sub-MLP layer and the second sub-MLP layer respectively to calculate the key K and value V;

[0072] S472: Input the original news text vector into the third sub-MLP layer to calculate the query Q;

[0073] S473: Perform matrix multiplication on the key K and the query Q to obtain an attention score;

[0074] S474: Perform matrix multiplication on the attention score and the value V to obtain a multimodal fusion vector.

[0075] S48: Input the multimodal fusion vector into the softmax function to calculate the category prediction result of the original news sample data;

[0076] The specific classification of fault categories using the Softmax classifier is as follows: Softmax is selected to calculate the probability that a sample belongs to different categories. For a given input x, the function is used to estimate the probability value p for each category j, and finally the diagnosis result is output.

[0077] S49: Based on the category prediction results of the original news sample data and the label information of the original news sample data, the parameters of the news classification model are updated through the back propagation mechanism using the cross entropy loss function;

[0078] Then the cross entropy calculation is performed between the Softmax output vector [y1, y2, y3...] and the actual label of the sample. The formula is as follows:

[0079]

[0080] where y' i is the label of the original news sample data, y i is the diagnosis result of the news classification model on the original news sample data, H y‘ (y) represents the loss function. Then the output vector is averaged to get the desired loss function value. After defining the loss function, the back propagation algorithm is used to minimize the loss function.

[0081] S5: Obtain target news sample data, perform feature extraction on the news text and news pictures in the target news sample data to obtain a target news text vector and a target news picture vector set, and input the target news text vector and the target news picture vector set into a news classification model to obtain a classification result for the target news sample data.

[0082] The present invention adopts the LSTM algorithm in the news classification model to calculate the hidden information in the illustrations, and adds weight information to each illustration through the attention network to reflect the importance of each illustration. By combining the text information and illustration information of the news across the attention network and inputting them into the softmax function for classification, the obtained news classification results have higher credibility and accuracy.

[0083] According to the classification results of the target news sample data, the corresponding users are matched and the target news data is recommended to the users through the social platform.

[0084] The present invention can design the news classification method described in the present invention into a computer program, store it in the storage of smart devices such as mobile phones, computers, calculators, and counter-insurgency devices, obtain news data through mobile phones, computers, calculators, and counter-insurgency smart devices, and run the computer program to implement it.

[0085] The above preferred embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the present invention. Those skilled in the art should understand that various changes can be made thereto in form and details without departing from the scope defined by the claims of the present invention.

Claims

1. A news classification method based on multimodal feature fusion, characterized in that: The following steps are involved: S1: Obtain original news sample data; the original news sample data includes: original news text and multiple original news pictures; label the original news sample data; and divide each original news picture into p image blocks of the same scale to obtain an image set of each original news picture; S2: Extract features from the original news text to obtain the original news text vector, and extract features from the image set of each original news picture to obtain the image sequence vector of each original news picture; S3: Build a news classification model, which includes: LSTM, attention network, gated memory network, cross-attention network, and softmax function; S4: Use the original news text vectors and all the image sequence vectors of the original news illustrations as training samples to train the news classification model; The training of the news classification model using the original news text vector and the original news picture vector sequence as training samples includes: S41: Input the image sequence vectors of all original news pictures into the LSTM to calculate the hidden state vector of each image sequence vector at the current time step; S42: Concatenate the hidden state vectors of all image sequence vectors at the current time step to obtain a first combined vector at the current time step; S43: performing feature concatenation on the first combined vector of the current time step and the first combined vector of the previous time step to obtain the second combined vector of the current time step; S44: Input the second combination vector of the current time step into the self-attention network to obtain the importance score matrix of the second combination vector of the current time step through the attention mechanism; S45: performing a Hadamard product calculation on the second combination vector of the current time step and the importance score matrix of the second combination vector of the current time step to obtain a third combination vector of the current time step; S46: Input the third combination vector of the current time step into the gated memory network to calculate the fourth combination vector of the current time step. The first combination vector, the second combination vector, the third combination vector, and the fourth combination vector are updated at each time step. When the LSTM converges, that is, at the last time step, the fourth combination vector calculated by the gated memory network is the final combination vector. S47: The final combined vector and the original news text vector are input into the cross-attention network to calculate the multimodal fusion vector; S48: Input the multimodal fusion vector into the softmax function to calculate the category prediction result of the original news sample data; S49: Based on the category prediction results of the original news sample data and the label information of the original news sample data, the parameters of the news classification model are updated through the back propagation mechanism using the cross entropy loss function; S5: Obtain target news sample data, perform feature extraction on the target news sample data to obtain a target news text vector and a plurality of target news picture sequence vectors, and input the target news text vector and the plurality of target news picture sequence vectors into a news classification model to obtain a classification result of the target news sample data.

2. A news classification method based on multimodal feature fusion according to claim 1, characterized in that: The gated memory network includes: a first MLP multi-layer perceptron, a second MLP multi-layer perceptron, a σ activation function, and a tanh activation function; The step of inputting the third combination vector of the current time step into the gated memory network to calculate the fourth combination vector of the current time step includes: S461: Input the third combined vector of the current time step into the first MLP multi-layer perceptron and then into the σ activation function to obtain the first weight value of the current time step; S462: Input the third combination vector of the current time step into the second MLP multi-layer perceptron and then into the τ activation function to obtain the updated suggestion vector of the current time step; S463: Calculate a fourth combination vector at the current time step according to the updated suggestion vector of the current time step and the first weight value.

3. A news classification method based on multimodal feature fusion according to claim 2, characterized in that: The fourth combination vector of the current time step includes: Among them, u t Represents the fourth combination vector of the current time step t, γ t represents the first weight value of the current time step, represents the updated proposal vector for the current time step, u t-1 Represents the fourth combination vector of the previous time step t-1.

4. The news classification method based on multimodal feature fusion according to claim 1 is characterized in that: The cross-attention network includes: a first sub-MLP layer, a second sub-MLP layer, and a third sub-MLP layer; The step of inputting the final combined vector and the original news text vector into the cross-attention network to calculate the multimodal fusion vector includes: S471: Input the final state vector into the first sub-MLP layer and the second sub-MLP layer respectively to calculate the key K and value V; S472: Input the original news text vector into the third sub-MLP layer to calculate the query Q; S473: Perform matrix multiplication on the key K and the query Q to obtain an attention score; S474: Perform matrix multiplication on the attention score and the value V to obtain a multimodal fusion vector.

5. The news classification method based on multimodal feature fusion according to claim 1 is characterized in that: The cross entropy loss function includes: where y i ’ is the label of the original news sample data, y i is the diagnosis result of the news classification model on the original news sample data, H y‘ (y) represents the loss function.