Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement

By adopting multi-source bilingual and cross-modal feature fusion and enhancement methods in the emotion classification method, the lack of performance problems in the existing technology in handling mixed language text and image sentiment analysis is solved, and a more efficient and accurate emotion classification effect is achieved.

CN119961759AActive Publication Date: 2025-05-09UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510322558.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-09
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Existing emotion classification methods are difficult to make full use of emotional information when processing mixed language texts, and fail to effectively combine text information and visual elements in the image, resulting in insufficient performance in Chinese social media text data processing.

Method used

The Chinese emotion classification method based on multi-source bilingual and cross-modal feature fusion and enhancement is adopted. Through the bilingual feature extraction model and computer vision model, the bilingual features in Chinese text and images are extracted, and the cross-modal attention fusion is performed to generate cross-modal emotional representations.

Benefits of technology

The accuracy of Chinese emotion classification task is significantly improved, especially in the emotion analysis task of processing Chinese-English mixed text and images, and the performance and robustness of emotion classification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961759A_ABST
    Figure CN119961759A_ABST
Patent Text Reader

Abstract

The invention discloses a Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement, which belongs to the technical field of sentiment calculation, and specifically comprises the following steps: acquiring and translating a Chinese text to be classified; obtaining an image Chinese text of the to-be-classified image and translating the image Chinese text; constructing a bilingual feature extraction model, and respectively extracting bilingual fusion features of the Chinese text to be classified and the image Chinese text by using the bilingual feature extraction model; extracting global visual emotion semantic representation of the to-be-classified image; constructing an emotion knowledge base, and performing search processing on the bilingual fusion features of the Chinese texts to be classified to obtain external emotion representations; and through cross-modal attention fusion of the to-be-classified image, enhancing bilingual fusion features of the to-be-classified Chinese text, and predicting an emotion classification result based on the obtained cross-modal emotion representation. By fully mining the characteristics of the Chinese-English mixed text and illustration published by the user, strong sentiment expression vocabularies from different sources are effectively captured, the accuracy of a Chinese sentiment classification task is remarkably improved, and the model calculation efficiency is considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of natural language processing, computer vision and sentiment computing, and specifically relates to a Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement. Background Art

[0002] With the booming development of various social media platforms, the massive amount of text and image data generated by users provides a valuable resource for insight into public sentiment. However, current sentiment classification methods face many challenges when dealing with mixed-language texts, especially when English words and mixed expressions are frequently interspersed in Chinese texts. In addition, current sentiment classification methods do not consider image data with obvious emotional information. These images with emotional information often also contain text with concise and clear emotional characteristics. Nowadays, the text in images often appears in a mixture of Chinese and English, expressing emotions strongly. For example, mixed-language text expressing sadness, "If your hair is not to your liking, you will be in a low mood all day", and such Figure 1 As shown in the accompanying pictures, these mixed language words usually contain rich emotional information. For example, the English word "low" has a clear negative emotional tendency. In addition, the accompanying pictures brought by the user also contain the obvious emotional word "emo" and visual elements with corresponding emotional characteristics.

[0003] Sentiment classification technology has gradually developed from early machine learning methods, such as support vector machines (SVM), naive Bayes, and classification and regression trees (CRT), to deep learning methods. The emergence of Transformer networks has greatly promoted the progress in the field of natural language processing. Transformer-based pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and its variants, such as RoBERTa, XLNet, Nezha, ELECTRA, and ERNIE, have become the mainstream models for sentiment analysis tasks. However, when faced with mixed-language texts, existing models often fail to fully utilize the sentiment information therein.

[0004] Current multimodal sentiment classification methods attempt to integrate multiple signals such as text, speech, and vision. However, in the text channel, they do not take into account the increasingly common phenomenon of mixed language expressions, fail to effectively utilize the multilingual modeling capabilities of pre-trained models, and the analysis of sentiment in mixed language environments needs to be further improved. In addition, these multimodal methods only consider the emotional visual elements expressed by the image itself, and often ignore the text information contained in the image. The text information in these images usually contains a lot of clear and concise emotional information. Moreover, with the development of the Internet, scenes where Chinese and English are mixed are gradually becoming popular in images.

[0005] Therefore, there is an urgent need for a sentiment classification method that can effectively process multi-source mixed-language texts. This method should make full use of the multilingual representation capabilities of the pre-trained language model, extract and fuse bilingual features, and thus significantly improve the sentiment analysis performance for mixed Chinese and English texts. In addition, this method can also extract the corresponding significant emotional features from the image information carried by the user, as well as the mixed-language text information in the image, combine the emotional elements in the image with the emotional elements in the text, and fuse the significant emotional characteristics from different languages ​​with the significant emotional characteristics of different modalities, so as to better meet the actual application needs in the current Chinese social media context. Summary of the invention

[0006] In response to the problems existing in existing multimodal sentiment classification methods, the present invention provides a Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement, which can significantly improve the accuracy of Chinese sentiment classification tasks while taking into account the model calculation efficiency, and has strong practical value.

[0007] In order to achieve the above purpose, the technical method adopted by the present invention is as follows:

[0008] The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement includes the following steps:

[0009] Step 1: obtaining Chinese text to be classified and a corresponding image to be classified, and translating the Chinese text to be classified into semantically aligned English text to be classified; the Chinese text to be classified is a mixed Chinese and English text containing a small amount of English;

[0010] Step 2: Build and train a bilingual feature extraction model, including a pre-trained language model, a Chinese lightweight network module, an English lightweight network module, and a bilingual cross-attention module;

[0011] Step 3: Input the Chinese text to be classified and the English text to be classified into the trained bilingual feature extraction model to extract the corresponding Chinese features. and English features , and bilingual integration features ;

[0012] Step 4: Use the computer vision model to perform text detection on the image to be classified and extract the text area contained in the image to be classified; then use the text recognition module (OCR) to recognize the text area as a readable string as the Chinese text of the image, and translate it into semantically aligned English text of the image;

[0013] Input the image Chinese text and image English text into the trained bilingual feature extraction model to extract the corresponding image Chinese features and image English features , and image bilingual fusion features ;

[0014] Step 5: Based on a multilingual corpus containing multiple text segments and known sentiment classification, the Chinese text segments are translated into semantically aligned English text segments, and the trained bilingual feature extraction model is used to extract Chinese features and English features corresponding to the text segments to construct a sentiment knowledge base;

[0015] Step 6: In the sentiment knowledge base Perform nearest neighbor retrieval to obtain multiple most relevant semantic representations, and obtain external sentiment representation by weighting, weighted summing and averaging them. ;

[0016] Step 7: Use the pre-trained convolutional neural network to extract global features of the classified image to obtain a global image representation; construct and train the emotion feature extraction module, and use the trained emotion feature extraction module to extract emotion-related visual features in the global image representation as the global visual emotion semantic representation. ;

[0017] Step 8: Build and train a cross-modal attention fusion module, and use the trained cross-modal attention fusion module to , and Perform cross-modal attention fusion to obtain cross-modal fusion representation, and use the residual mechanism to combine the cross-modal fusion representation with Combined to obtain cross-modal sentiment representation ;

[0018] Step 9: Use sentiment classifier to Perform sentiment category prediction and obtain the sentiment classification results of the Chinese text to be classified.

[0019] Furthermore, the bilingual feature extraction model is specifically:

[0020] Taking Chinese text and English text as the input of the bilingual feature extraction model, the pre-trained language model extracts the hidden layer representation of the Chinese text and the hidden layer representation of the English text; the Chinese lightweight network module and the English lightweight network module are added to each linear layer of the pre-trained language model; after the Chinese text is processed by the Chinese lightweight network module, it is combined with the hidden layer representation of the Chinese text to obtain the Chinese features; after the English text is processed by the English lightweight network module, it is combined with the hidden layer representation of the English text to obtain the English features; the Chinese features and the English features are fused through the bilingual cross-attention module to output the bilingual fused features.

[0021] Furthermore, the Chinese lightweight network module and the English lightweight network module are designed based on low rank adaptation (LoRA) technology.

[0022] Furthermore, the bilingual feature extraction model is trained based on a sentiment analysis evaluation database, and in the process of training a pre-trained language model in the bilingual feature extraction model, only parameters in the embedding layer are trained, and other parameters of the pre-trained language model are frozen.

[0023] Furthermore, the pre-trained language model is BERT, RoBERTa or XLNet.

[0024] Furthermore, in step 3, bilingual fusion features are extracted The specific process is:

[0025] Chinese Features and English features Perform linear transformations respectively to obtain the corresponding Chinese query matrix (Query Matrix, Q) 、Chinese Key Matrix (Key Matrix, K) , Chinese value matrix (Value Matrix, V) , English query matrix , English key matrix and English value matrix ;

[0026] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get feature enhancement from English to Chinese ;

[0027] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement from Chinese to English ;

[0028] Using Feed-Forward Network (FFN) and Further information interaction and semantic integration are performed to obtain the Chinese hidden layer representation after crossover and English hidden layer representation after cross :

[0029] ;

[0030] ;

[0031] In the formula, Indicates normalization processing; represents feed-forward neural network processing;

[0032] Furthermore, and Sum and get the bilingual fusion feature .

[0033] Furthermore, in step 4, the image bilingual fusion features are extracted The specific process is:

[0034] Chinese features of images and image English features Perform linear transformations respectively to obtain the corresponding image Chinese query matrix , Image Chinese Key Matrix , image Chinese value matrix , image English query matrix , Image English Key Matrix and image English value matrix ;

[0035] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement of image English to Chinese ;

[0036] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, the image Chinese to English feature enhancement is obtained ;

[0037] Using feed-forward neural network and Further information interaction and semantic integration are performed to obtain the Chinese hidden layer representation of the cross-image And the English hidden layer representation of the image after cross :

[0038] ;

[0039] ;

[0040] right Average all word vectors contained in The average of all word vectors contained in is calculated, and the two average values ​​are summed to obtain the image bilingual fusion feature. .

[0041] Furthermore, in step 6, a scoring mechanism is used to assign weights to the most relevant semantic representations.

[0042] Furthermore, in step 7, the sentiment feature extraction module is implemented based on a fully connected layer and trained based on a sentiment analysis evaluation database. Specifically, the global image representation output by a pre-trained convolutional neural network is used as input to train the parameters of the sentiment feature extraction module.

[0043] Further, in step 8, the training process of the cross-modal attention fusion module is: training based on the sentiment analysis evaluation database, specifically using the output of the trained bilingual feature extraction model , the output of the trained sentiment feature extraction module And the result of step 6 The weight of the cross-modal attention fusion module is trained as the input of the cross-modal attention fusion module, and the parameters of the trained bilingual feature extraction model and the trained sentiment feature extraction module are kept unchanged during the training process.

[0044] The emotion feature extraction module is implemented based on a fully connected layer.

[0045] Compared with the prior art, the present invention has the following beneficial effects: 1. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement proposed in the present invention fully mines the features of the Chinese-English mixed texts published by users and the corresponding images (i.e., pictures), especially the bilingual significant emotional text information contained in the images, effectively captures strong emotional expression words from different sources (including text content and text information in images), thereby significantly improving the accuracy of Chinese sentiment classification tasks; 2. The present invention constructs an emotional knowledge base to find historical information for texts with unclear emotional expression, so as to enhance features and thus improve the effectiveness of emotional information extraction in bilingual mixed scenarios; 3. The Chinese sentiment classification method proposed in the present invention performs particularly well in processing social media text data containing mixed languages ​​and a mixture of Chinese and English. Compared with existing methods, the present invention not only integrates bilingual features, but also adopts a lightweight model fine-tuning strategy. While significantly improving the sentiment classification performance, it also takes into account the model calculation efficiency. It has strong practical value and is expected to be widely used in the fields of intelligent customer service, personalized recommendations, etc., providing strong support for the implementation of artificial intelligence technology in actual scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0050] Figure 1 It is a picture with the emotional word "emo";

[0051] Figure 2 It is a flowchart of the Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement proposed in Example 1;

[0052] Figure 3 is a schematic diagram of the principle of the bilingual feature extraction model proposed in Example 1;

[0053] Figure 4 It is a flow chart of the lightweight model fine-tuning strategy proposed in Example 1. DETAILED DESCRIPTION

[0054] In order to further understand the present invention, preferred embodiments of the present invention are described below in conjunction with examples, but it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, rather than limiting the claims of the invention.

[0055] Example 1

[0056] This embodiment proposes a Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement. The process is as follows: Figure 2 As shown, the following steps are included:

[0057] Step 1: obtain the Chinese text to be classified and the corresponding image to be classified, and use the machine translation model to translate the Chinese text to be classified into semantically aligned English text to be classified;

[0058] The Chinese text to be classified is a mixed Chinese-English text containing a small amount of English; the machine translation model can adopt existing online translation tools, such as Google Translate, Bing Microsoft Translator, etc., or a self-trained high-quality neural machine translation model, such as Transformer, BART, etc.;

[0059] Example: The Chinese text to be classified is "If your hair does not follow your heart, your mood will be very low all day long", and the translated English text to be classified is "If your hair does not follow your heart, your emotion will be very low".

[0060] Step 2: Build a bilingual feature extraction model, such as Figure 3 As shown, it includes a pre-trained language model, a Chinese lightweight network module, an English lightweight network module, and a bilingual cross-attention module, and is trained based on a sentiment analysis evaluation database. Specifically:

[0061] The pre-trained language model is specifically BERT;

[0062] The Chinese lightweight network module and the English lightweight network module are designed based on low-rank adaptation technology and added to each linear layer of the pre-trained language model. Specifically, a low-rank matrix is ​​introduced into each linear layer. , and compare it with the original weight matrix Add together to get the adapted weight matrix ;

[0063] Taking Chinese text and English text as the input of the bilingual feature extraction model, the pre-trained language model extracts the hidden layer representation of the Chinese text and the hidden layer representation of the English text; after the Chinese text is processed by the Chinese lightweight network module, it is combined with the hidden layer representation of the Chinese text to obtain the Chinese features; after the English text is processed by the English lightweight network module, it is combined with the hidden layer representation of the English text to obtain the English features; the Chinese features and the English features are fused through the bilingual cross attention module to output the bilingual fusion features; wherein the Chinese features and the English features both include the corresponding global sentiment representation and the local word-level sentiment representation; for example, in BERT, the [CLS]th or 0th word vector can be regarded as the global sentiment representation, and the subsequent word vectors are used as the local word-level sentiment representation; in the training process of the bilingual feature extraction model, a lightweight model fine-tuning strategy is adopted, such as Figure 4As shown in the figure, the original parameters of each layer in the fixed pre-trained language model remain unchanged, and only the various embedding layers of the pre-trained language model (including word type embedding coding, position embedding coding and word embedding coding), the low-rank matrix of the Chinese lightweight network module and the English lightweight network module are trained. And the bilingual cross attention module; among them, only the low-rank matrix needs to be learned , while keeping the original weight matrix The pre-trained language model remains unchanged, which greatly reduces the number of parameters that need to be updated. In this way, the model can be efficiently fine-tuned on downstream tasks without significantly increasing the computational overhead. Only training the embedding matrix in the embedding layer can align the word representations of different languages ​​in the semantic space to make them closer to the target task, and freezing other parameters of the pre-trained language model can greatly reduce the number of fine-tuning parameters while still being able to effectively adapt to downstream tasks.

[0064] A lightweight model fine-tuning strategy is adopted to minimize the computational overhead while ensuring the fine-tuning effect, so that the bilingual feature extraction model can adapt to the sentiment classification task more efficiently, significantly reducing the training time and resource consumption, and has strong practical value.

[0065] Step 3: Input the Chinese text to be classified and the English text to be classified into the trained bilingual feature extraction model to extract the corresponding Chinese features. and English features , and bilingual integration features ;

[0066] Among them, extracting bilingual fusion features The specific process is:

[0067] Chinese Features and English features Perform linear transformations respectively to obtain the corresponding Chinese query matrix , Chinese key matrix , Chinese value matrix , English query matrix , English key matrix and English value matrix :

[0068] ;

[0069] ;

[0070] ;

[0071] ;

[0072] ;

[0073] ;

[0074] In the formula, , and All represent the linear transformation matrix of Chinese; , and Both represent linear transformation matrices in English. Through linear transformation, the original features are mapped to the same feature space to prepare for subsequent attention calculations.

[0075] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get feature enhancement from English to Chinese :

[0076] ;

[0077] In the formula, the superscript "T" represents transposition; d represents the length of the word vector; Represents the Softmax function; Indicates normalization processing;

[0078] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement from Chinese to English :

[0079] ;

[0080] Using Feed-Forward Network (FFN) and Further information interaction and semantic integration are performed to obtain the Chinese hidden layer representation after crossover and English hidden layer representation after cross :

[0081] ;

[0082] ;

[0083] In the formula, represents feed-forward neural network processing;

[0084] Furthermore, and Sum and get the bilingual feature representation ; Including the corresponding global sentiment representation, as well as the local word-level sentiment representation.

[0085] Step 4: Use a computer vision model to perform text detection on the image to be classified and extract the text area contained in the image to be classified; then use a text recognition module to recognize the text area as a readable string as the image Chinese text, and translate it into semantically aligned image English text; wherein, this embodiment specifically uses DBNet (Differentiable Binarization Network) as the computer vision model and CRNN (Convolutional Recurrent Neural Network) as the text recognition module;

[0086] Input the image Chinese text and image English text into the trained bilingual feature extraction model to extract the corresponding image Chinese features and image English features , and image bilingual fusion features ;

[0087] Among them, extracting image bilingual fusion features The specific process is:

[0088] Chinese features of images and image English features Perform linear transformations respectively to obtain the corresponding image Chinese query matrix , Image Chinese Key Matrix , image Chinese value matrix , image English query matrix , Image English Key Matrix and image English value matrix ;

[0089] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement of image English to Chinese ;

[0090] calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, the image Chinese to English feature enhancement is obtained ;

[0091] Using feed-forward neural network and Further information interaction and semantic integration are performed to obtain the Chinese hidden layer representation of the cross-image And the English hidden layer representation of the image after cross :

[0092] ;

[0093] ;

[0094] right Average all word vectors contained in The average of all word vectors contained in is calculated, and the two average values ​​are summed to obtain the image bilingual fusion feature. :

[0095] ;

[0096] In the formula, express The number of word vectors contained in ; express The number of word vectors contained in ; express The word vectors; express The word vectors; and is just a vector;

[0097] The bilingual cross-attention module adopted in this embodiment can fully explore the relationship between Chinese and English features, complement and enhance each other, make the emotional tendency clearer, and provide high-quality feature representation for subsequent emotional classification tasks.

[0098] Step 5: Based on a multilingual corpus containing multiple text segments and known sentiment classification, the Chinese text segments are translated into semantically aligned English text segments, and the trained bilingual feature extraction model is used to extract Chinese features and English features corresponding to the text segments to construct a sentiment knowledge base;

[0099] Step 6: In the sentiment knowledge base Perform nearest neighbor search, specifically by calculating The cosine similarity with all features in the sentiment knowledge base is used to obtain multiple most relevant semantic representations. The most relevant semantic representations are weighted using a scoring mechanism. Specifically, the cosine similarity score is normalized by Softmax. After weighted summation and averaging of multiple most relevant semantic representations, the external sentiment representation is obtained. , specifically a vector;

[0100] Step 7: Use the pre-trained convolutional neural network to extract global features of the image to be classified and obtain a global image representation;

[0101] Construct and train an emotion feature extraction module. In this embodiment, a fully connected layer is used as the emotion feature extraction module, and the global image representation output by the pre-trained convolutional neural network is used as input to train the parameters of the emotion feature extraction module.

[0102] The trained emotion feature extraction module is used to extract emotion-related visual features in the global image representation as the global visual emotion semantic representation. , specifically a vector;

[0103] Step 8: Construct a cross-modal attention fusion module based on the cross-attention mechanism, using the output of the trained bilingual feature extraction model , the output of the trained sentiment feature extraction module And the result of step 6 As the input of the cross-modal attention fusion module, the weight of the cross-modal attention fusion module is trained, and the parameters of the trained bilingual feature extraction model and the trained sentiment feature extraction module are kept unchanged during the training process;

[0104] Using the trained cross-modal attention fusion module , and Perform cross-modal attention fusion. , and Concatenate to get the matrix H f =[ h' text , h img , h ̂ text ] , using learnable query weights , key weight Sum weight ,Will Mapping to query matrix , key matrix Sum Matrix :

[0105] ;

[0106] ;

[0107] ;

[0108] By attention operation, update :

[0109] ;

[0110] Updated Three new vectors , and Spliced ​​together, that is H f =[ h 1 f , h 2 f , h 3 f ] ;

[0111] Then the cross-modal fusion representation is calculated :

[0112] ;

[0113] By adaptively allocating the importance of each modality’s features, full complementation and enhancement can be achieved between vision and text, bilingual and external emotional knowledge;

[0114] Then the residual mechanism is used to represent the cross-modal fusion and Combined to obtain a unified and enhanced cross-modal sentiment representation ;

[0115] Step 9: Use sentiment classifier to Perform sentiment category prediction to obtain the sentiment classification results of the Chinese text to be classified, specifically:

[0116] A linear classification layer is used to represent the cross-modal sentiment Probability distribution mapped to emotion categories :

[0117] ;

[0118] In the formula, Represents the weight matrix of the linear classification layer; Represents the bias vector of the linear classification layer; the output of the linear transformation is normalized into a probabilistic form using the Softmax function for multi-classification tasks;

[0119] Contains global sentiment category probabilities And the local word-level sentiment category probability, the final probability vector is obtained by weighted average , and determine the final sentiment category based on the element with the highest score;

[0120] =0.8× +0.2× ;

[0121] In the formula, express The number of local word-level sentiment category probabilities contained in ; Indicates Local word-level sentiment category probabilities.

[0122] In order to comprehensively evaluate the performance of the Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement, this embodiment conducted a large number of experiments on the SMP2020-EWECT dataset. The SMP2020-EWECT dataset contains social media text data and general field text data from a special period, covering different emotional polarities and themes. It is currently one of the largest and most comprehensive benchmark datasets for Chinese sentiment analysis tasks. In order to further enrich the dataset, this embodiment uses the Stable Diffusion model to automatically generate corresponding illustrations based on the text content of SMP2020-EWECT and designs specific instructions to guide Stable Diffusion to generate images with obvious emotional tendencies to supplement the missing visual information in the original dataset. In addition, this embodiment also uses the trained bilingual feature extraction model to deeply process the SMP2020-EWECT dataset and construct a comprehensive emotional knowledge base.

[0123] In the experiment, this embodiment is compared with multiple existing pre-trained language models, including BERT, RoBERTa, ERNIE, Nezha and BSF, and the data results shown in Table 1 are obtained. It can be seen that the Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement proposed in this embodiment significantly exceeds the existing pre-trained language model in terms of accuracy and macro-average F1 value, proving its effectiveness and robustness in processing mixed language text sentiment classification tasks. Compared with ERNIE (Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation), the method of this embodiment has improved the accuracy by 4.53% and the macro-average F1 value by 10.21%, which fully reflects the significant advantages of the method of this embodiment in processing mixed language text sentiment classification tasks. By integrating bilingual features, especially using the richer and more explicit emotional expressions in the English part, the method of this embodiment can more accurately capture the emotional information contained in the text, thereby greatly improving the performance of sentiment classification. Compared with BSF (Improving Chinese Emotion Classification based on Bilingual Feature Fusion), the method of this embodiment improves the macro-average F1 value by 3.22% on the basis of improving the accuracy by 3.1%, indicating that the method of this embodiment is effective in constructing and applying multi-source bilingual information, image emotion features, and emotion databases to emotion recognition, thereby improving the overall emotion classification performance.

[0124] Table 1

[0125]

[0126] In addition to the SMP2020-EWECT dataset, the Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement proposed in this embodiment also performs well on general datasets. General datasets cover a wider range of topics and fields, and put forward higher requirements on the generalization ability of the model. Table 2 shows the experimental results of the method of this embodiment on general datasets and compares them with other existing pre-trained language models. It can be seen that the method of this embodiment also achieves the best performance on general datasets, ranking first in both accuracy and macro-average F1 value. Compared with ERNIE, the method of this embodiment has improved by 1.63% in accuracy and 1.96% in macro-average F1 value, which once again proves the significant effect of integrating bilingual features, multi-source bilingual features, image emotional elements and emotional feature library support on improving sentiment classification performance. It is worth noting that compared with the SMP2020-EWECT dataset, the fields of general datasets are more diverse, including different types of texts such as e-commerce reviews, social media, and news reviews. Nevertheless, the method of this embodiment is still able to stably surpass various existing models, demonstrating strong generalization ability and robustness.

[0127] Table 2

[0128]

[0129] In summary, the Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement proposed in this embodiment has the following significant advantages over the existing pre-trained language model:

[0130] 1. The multilingual modeling capability of the pre-trained language model is fully utilized. By simultaneously inputting Chinese text and its English translation, bilingual information is effectively captured, making up for the shortcomings of the monolingual model. In addition, this embodiment also extracts the bilingual text information contained in the image, further enriching the input of sentiment analysis;

[0131] 2. The bilingual cross-attention mechanism is introduced. By calculating the similarity and information transfer between Chinese and English features, the deep fusion and interaction of text bilingual features and image bilingual features are realized, enabling the model to understand text sentiment more comprehensively from the perspectives of different languages ​​and modalities;

[0132] 3. We built a sentiment knowledge base that stores sentiment words and their semantic features in different languages. We used this knowledge base for retrieval and fusion, providing external knowledge enhancement for texts with unclear sentiment expression, and strengthening the effective extraction of sentiment information in mixed language scenarios.

[0133] 4. A cross-modal attention fusion mechanism is designed to adaptively adjust the importance of text bilingual features, image visual sentiment features, and knowledge base retrieval features, so that they can enhance each other and obtain a unified cross-modal sentiment representation. The cross-modal features are then fused into the original bilingual features through residual connections, further enhancing the sentiment representation capability.

[0134] 5. A lightweight model fine-tuning strategy is adopted, combined with low-rank adaptation and embedding layer fine-tuning technology, which achieves efficient adaptation of the model to downstream tasks without significantly increasing computational overhead, and has strong practical value;

[0135] 6. Comprehensive experimental evaluations were conducted on multiple benchmark datasets, demonstrating the significant advantages of the method of this embodiment in mixed-language text sentiment classification tasks, providing new ideas and references for the development of Chinese sentiment analysis technology.

[0136] The above embodiments are intended to provide a better understanding of the present invention, but are not limited to the optimal implementation mode and do not constitute a limitation on the content and protection scope of the present invention. Any product identical or similar to the present invention obtained by anyone under the inspiration of the present invention or by combining the present invention with the features of other prior arts shall be within the protection scope of the present invention.

Claims

1. A Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement, characterized in that: The following steps are involved: Step 1: obtaining Chinese text to be classified and a corresponding image to be classified, and translating the Chinese text to be classified into semantically aligned English text to be classified; the Chinese text to be classified is a mixed Chinese and English text containing a small amount of English; Step 2: Build and train a bilingual feature extraction model, including a pre-trained language model, a Chinese lightweight network module, an English lightweight network module, and a bilingual cross-attention module; Step 3: Input the Chinese text to be classified and the English text to be classified into the trained bilingual feature extraction model to extract the corresponding Chinese features. and English features , and bilingual integration features ; Step 4: Use the computer vision model to perform text detection on the image to be classified and extract the text area contained in the image to be classified; then use the text recognition module to recognize the text area as a string as the Chinese text of the image, and translate it into semantically aligned English text of the image; Input the image Chinese text and image English text into the trained bilingual feature extraction model to extract the corresponding image Chinese features and image English features , and image bilingual fusion features ; Step 5: Based on a multilingual corpus containing multiple text segments and known sentiment classification, the Chinese text segments are translated into semantically aligned English text segments, and the trained bilingual feature extraction model is used to extract Chinese features and English features corresponding to the text segments to construct a sentiment knowledge base; Step 6: In the sentiment knowledge base Perform nearest neighbor retrieval to obtain multiple most relevant semantic representations, and obtain external sentiment representation by weighting, weighted summing and averaging them. ; Step 7: Use the pre-trained convolutional neural network to extract global features of the classified image to obtain a global image representation; construct and train the emotion feature extraction module, and use the trained emotion feature extraction module to extract emotion-related visual features in the global image representation as the global visual emotion semantic representation. ; Step 8: Build and train a cross-modal attention fusion module, and use the trained cross-modal attention fusion module to , and Perform cross-modal attention fusion to obtain cross-modal fusion representation, and use the residual mechanism to combine the cross-modal fusion representation with Combined to obtain cross-modal sentiment representation ; Step 9: Use sentiment classifier to Perform sentiment category prediction and obtain the sentiment classification results of the Chinese text to be classified.

2. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 1 is characterized in that: The bilingual feature extraction model is specifically: Taking Chinese text and English text as the input of the bilingual feature extraction model, the pre-trained language model extracts the hidden layer representation of the Chinese text and the hidden layer representation of the English text; the Chinese lightweight network module and the English lightweight network module are added to each linear layer of the pre-trained language model; after the Chinese text is processed by the Chinese lightweight network module, it is combined with the hidden layer representation of the Chinese text to obtain the Chinese features; after the English text is processed by the English lightweight network module, it is combined with the hidden layer representation of the English text to obtain the English features; the Chinese features and the English features are fused through the bilingual cross-attention module to output the bilingual fused features.

3. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: Extract bilingual fusion features in step 3 The specific process is: Chinese Features and English features Perform linear transformations respectively to obtain the corresponding Chinese query matrix , Chinese key matrix , Chinese value matrix , English query matrix , English key matrix and English value matrix ; calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get feature enhancement from English to Chinese ; calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement from Chinese to English ; Then we get the Chinese hidden layer representation after crossover and English hidden layer representation after cross : ; ; In the formula, Indicates normalization processing; represents feed-forward neural network processing; Furthermore, and Sum and get the bilingual fusion feature .

4. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 3 is characterized in that: Step 4: Extract image bilingual fusion features The specific process is: Chinese features of images and image English features Perform linear transformations respectively to obtain the corresponding image Chinese query matrix , Image Chinese Key Matrix , image Chinese value matrix , image English query matrix , Image English Key Matrix and image English value matrix ; calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, we get the feature enhancement of image English to Chinese ; calculate and The similarity of , using the Softmax function to obtain the attention weight matrix of the similarity, and After weighting, Combined and normalized, the image Chinese to English feature enhancement is obtained ; Then we get the Chinese hidden layer representation of the cross-image And the English hidden layer representation of the image after cross : ; ; right Average all word vectors contained in The average of all word vectors contained in is calculated, and the two average values ​​are summed to obtain the image bilingual fusion feature. .

5. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: The Chinese lightweight network module and the English lightweight network module are designed based on low-rank adaptation technology.

6. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: In step 2, the bilingual feature extraction model is trained based on the sentiment analysis evaluation database, and in the process of training the pre-trained language model in the bilingual feature extraction model, only the parameters in the embedding layer are trained, and other parameters of the pre-trained language model are frozen.

7. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: The pre-trained language model is BERT, RoBERTa or XLNet.

8. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: In step 6, a scoring mechanism is used to assign weights to the most relevant semantic representations.

9. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 2 is characterized in that: In step 7, the sentiment feature extraction module is implemented based on a fully connected layer and trained based on a sentiment analysis evaluation database. Specifically, the global image representation output by a pre-trained convolutional neural network is used as input to train the parameters of the sentiment feature extraction module.

10. The Chinese sentiment classification method based on multi-source bilingual and cross-modal feature fusion and enhancement according to claim 9 is characterized in that: In step 8, the training process of the cross-modal attention fusion module is: training based on the sentiment analysis evaluation database, specifically using the trained bilingual feature extraction model output , the output of the trained sentiment feature extraction module And the result of step 6 The weight of the cross-modal attention fusion module is trained as the input of the cross-modal attention fusion module, and the parameters of the trained bilingual feature extraction model and the trained sentiment feature extraction module are kept unchanged during the training process.

Citation Information

Patent Citations

  • Multi-modal Mongolian sentiment analysis method based on irony recognition and fine-grained feature fusion

    CN113657115A

  • Emotion analysis method based on multi-modal cross attention mechanism image-text fusion

    CN116844179A

  • AU2021104773A4