Multi-modal image-text emotion recognition method and system based on dynamic routing hybrid expert model
Through the dynamic routing hybrid expert model, the image and text data is encoded and feature fusion is solved, and the problem of static modal weight allocation in the prior art is achieved, achieving more efficient and interpretable multimodal graphic and text emotion recognition.
Patent Information
- Application Number
- CN202510478044.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-08
AI Technical Summary
Existing multimodal graphic and text emotion recognition methods usually adopt static fusion mechanisms, and the weight allocation between modes cannot be adaptively adjusted according to the semantic features of specific samples, resulting in poor model recognition effect.
A multimodal graphic and text emotion recognition method based on the dynamic routing hybrid expert model is adopted, and the image and text data are encoded through the target encoder to generate global features. The dynamic routing hybrid expert model is used to perform dynamic expert calculation and weighted feature fusion, and finally the recognition results are generated through the emotion classifier.
Adaptively selecting experts to perform expert calculations based on the input content is realized, which improves the recognition effect of the model when dealing with complex graphics and text relationships, and improves the calculation efficiency and interpretability of the model.
Smart Images

Figure CN120277613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a multi-modal graphic-text sentiment recognition method and system based on a dynamic routing mixture of experts model. Background Technique
[0002] As a key technology in natural language processing, sentiment recognition technology aims to infer the emotional polarity that a user wants to express through various types of information published by the user, and is mainly used to identify and extract the emotional tendencies, attitudes, and emotions in audio, images, or texts.
[0003] With the booming development of the Internet, people increasingly use multi-modal information such as text, pictures, voice, and video on social media to share their views. By performing sentiment analysis on the multi-modal data published by people, not only the emotional tendencies of pictures and texts are considered simultaneously, but also the semantic associations between graphics and texts are captured, providing new ideas and perspectives for multi-modal sentiment analysis on social media. Currently, multi-modal sentiment recognition for social media has become an important research hotspot in the field of artificial intelligence.
[0004] Existing multi-modal graphic-text sentiment recognition methods usually adopt a static fusion mechanism to integrate modal features through preset fixed attention weights or expert assignment rules. This design has significant limitations: since the fusion strategy lacks the ability to dynamically respond to the input content, when faced with graphic-text semantic conflicts or modal quality differences, the model cannot adaptively adjust the weight distribution between modalities according to the semantic features of specific samples, resulting in poor model recognition effects. Summary of the Invention
[0005] The present invention provides a multi-modal graphic-text sentiment recognition method and system based on a dynamic routing mixture of experts model for the technical problem that existing multi-modal graphic-text sentiment recognition methods usually adopt a static fusion mechanism, resulting in poor model recognition effects.
[0006] A multi-modal graphic-text sentiment recognition method based on a dynamic routing mixture of experts model provided by the first aspect of the present invention includes:
[0007] Obtain image data and text data, and input the image data and the text data into a preset multi-modal graphic-text sentiment recognition network, where the preset multi-modal graphic-text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier;
[0008] Encode the image data and the text data through the target encoder, and output an image global feature and a text global feature;
[0009] Perform multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature;
[0010] Perform dynamic expert calculation on the image-text fusion features using the dynamic routing mixture-of-experts model, and output weighted features;
[0011] Input the weighted features into the sentiment classifier for recognition to generate the target multi-modal image-text sentiment recognition result.
[0012] Optionally, the target encoder includes an image encoder and a text encoder; encoding the image data and the text data through the target encoder to output an image global feature and a text global feature, including:
[0013] Perform image encoding on the image data through the image encoder to output an image global feature;
[0014] Use the text encoder to perform text encoding on the text data to generate a text global feature.
[0015] Optionally, performing image encoding on the image data through the image encoder to output an image global feature includes:
[0016] Segment the image data to generate a plurality of image patches;
[0017] Perform linear embedding on each of the image patches to generate a plurality of image embeddings;
[0018] Use a plurality of the image embeddings to form an image embedding sequence, and insert a preset first classification token into the image embedding sequence to output a marked image embedding sequence;
[0019] Perform Transformer encoding on the marked image embedding sequence to generate an image global feature.
[0020] Optionally, using the text encoder to perform text encoding on the text data to generate a text global feature includes:
[0021] Perform word segmentation on the text data to generate a plurality of text words;
[0022] Perform linear embedding on each of the text words to generate a plurality of text word embeddings;
[0023] Use a plurality of the text word embeddings to form a text word embedding sequence, and insert a preset second classification token into the text word embedding sequence to output a marked text word embedding sequence;
[0024] Perform Transformer encoding on the marked text word embedding sequence to output a text global feature.
[0025] Optionally, the multi-modal feature fusion of the image global feature and the text global feature to generate an image-text fusion feature includes:
[0026] Perform cross-attention calculation on the image global feature and the text global feature to generate a cross-attention alignment feature;
[0027] Concatenate the cross-attention alignment feature, the image global feature, and the text global feature, and output a concatenated feature;
[0028] Reduce the dimension of the concatenated feature to generate an image-text fusion feature.
[0029] Optionally, the dynamic routing mixture-of-experts model includes two serially connected fully connected layers, a Softmax activation function layer, and a mixture-of-experts layer; performing dynamic expert calculation on the image-text fusion feature using the dynamic routing mixture-of-experts model to output a weighted feature includes:
[0030] Perform feature transformation on the image-text fusion feature through the two serially connected fully connected layers to generate a reduced-dimensional feature;
[0031] Use the Softmax activation function layer to calculate weights based on the reduced-dimensional feature to determine the weights corresponding to multiple experts in the mixture-of-experts layer;
[0032] Sort the weights corresponding to each expert in descending order, select the experts corresponding to the top pre-set number of weights to perform expert calculation on the multi-modal fusion feature, and output multiple expert results;
[0033] Perform weighted summation on each expert result according to the top pre-set number of weights to generate a weighted feature.
[0034] Optionally, the sentiment classifier includes two serially connected fully connected layers and an activation function layer; inputting the weighted feature into the sentiment classifier for recognition to generate a target multi-modal image-text sentiment recognition result includes:
[0035] Perform feature transformation on the weighted feature using the two serially connected fully connected layers and output a transformed feature;
[0036] Perform non-linear mapping on the transformed feature through the activation function layer to generate a target multi-modal image-text sentiment recognition result.
[0037] Optionally, the model training process of the pre-set multi-modal image-text sentiment recognition network includes:
[0038] Obtain a training image data set and a training text data set;
[0039] Using a preset loss function and an initial multi-modal image-text sentiment recognition network, determine the model gradient according to the training image data set and the training text data set;
[0040] Using the model gradient, update the model parameters of the initial multi-modal image-text sentiment recognition network, determine an intermediate multi-modal image-text sentiment recognition network, and real-time statistically count the number of iterative model updates;
[0041] Judge whether the number of model updates reaches a preset number of training times;
[0042] If it reaches, use the intermediate multi-modal image-text sentiment recognition network as the trained preset multi-modal image-text sentiment recognition network.
[0043] A multi-modal image-text sentiment recognition system based on a dynamic routing mixture of experts model provided by the second aspect of the present invention includes:
[0044] An acquisition module, configured to acquire image data and text data, and input the image data and the text data into a preset multi-modal image-text sentiment recognition network, where the preset multi-modal image-text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier;
[0045] An encoding module, configured to encode the image data and the text data through the target encoder, and output an image global feature and a text global feature;
[0046] A fusion module, configured to perform multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature;
[0047] An expert calculation module, configured to perform dynamic expert calculation on the image-text fusion feature by using the dynamic routing mixture of experts model, and output a weighted feature;
[0048] A recognition module, configured to input the weighted feature into the sentiment classifier for recognition, and generate a target multi-modal image-text sentiment recognition result.
[0049] A computer device provided by the third aspect of the present invention includes a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the multi-modal image-text sentiment recognition method according to any one of the above.
[0050] From the above technical solutions, it can be seen that the present invention has the following advantages:
[0051] The above technical solution of the present invention provides a multi-modal graphic and text sentiment recognition method based on a dynamic routing mixture of experts model. First, image data and text data are obtained, and the image data and text data are input into a preset multi-modal graphic and text sentiment recognition network. The preset multi-modal graphic and text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier. Then, the image data and text data are encoded by the target encoder to output an image global feature and a text global feature. The image global feature and the text global feature are subjected to multi-modal feature fusion to generate an image-text fusion feature. The dynamic routing mixture of experts model is used to perform dynamic expert calculation on the image-text fusion feature to output a weighted feature. Finally, the weighted feature is input into the sentiment classifier for recognition to generate a target multi-modal graphic and text sentiment recognition result. Based on the above solution, in the process of processing the obtained image data and text data through the target encoder, the dynamic routing mixture of experts model, and the sentiment classifier to output the target multi-modal graphic and text sentiment recognition result, the present invention adaptively selects an expert according to the input features through the dynamic routing mixture of experts model to perform expert calculation, so as to process the complex relationship between graphics and texts, and further improve the model recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the steps of a multi-modal graphic and text sentiment recognition method based on a dynamic routing mixture of experts model provided by Embodiment 1 of the present invention;
[0054] Figure 2 It is an overall framework diagram of the multi-modal graphic and text sentiment recognition method based on the dynamic routing mixture of experts model provided by Embodiment 1 of the present invention;
[0055] Figure 3 It is a schematic diagram of the dynamic routing mechanism provided by Embodiment 1 of the present invention;
[0056] Figure 4 It is a flowchart of the steps of model training of a preset multi-modal graphic and text sentiment recognition network provided by Embodiment 2 of the present invention;
[0057] Figure 5 It is a structural block diagram of a multi-modal graphic and text sentiment recognition system based on a dynamic routing mixture of experts model provided by Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] An embodiment of the present invention provides a multi-modal graphic-text sentiment recognition method and system based on a dynamic routing mixture of experts model, which is used to solve the technical problem that existing multi-modal graphic-text sentiment recognition methods usually adopt a static fusion mechanism, resulting in poor model recognition effects.
[0059] In order to make the invention purpose, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0060] Please refer to Figure 1 , Figure 1 which is a flowchart of the steps of a multi-modal graphic-text sentiment recognition method provided in Embodiment 1 of the present invention.
[0061] A multi-modal graphic-text sentiment recognition method based on a dynamic routing mixture of experts model provided by the present invention includes:
[0062] Step 101: Obtain image data and text data, and input the image data and text data into a pre-set multi-modal graphic-text sentiment recognition network. The pre-set multi-modal graphic-text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier.
[0063] It should be noted that, please refer to Figure 2 , the multi-modal graphic-text sentiment recognition proposed by the present invention is composed of a target encoder, a dynamic routing mixture of experts model (MoE, Mixture of Experts), and a sentiment classifier. The target encoder includes a Vision Transformer (ViT, image encoder) and a BERT (Bidirectional Encoder Representations from Transformers, text encoder). The present invention takes the Vision Transformer (ViT), BERT, and mixture of experts model (MoE) as the core, aiming to adaptively fuse image and text features through a dynamic routing mechanism to improve the accuracy and efficiency of sentiment analysis.
[0064] Step 102: Encode the image data and text data through the target encoder, and output the image global feature and the text global feature.
[0065] The target encoder includes an image encoder and a text encoder.
[0066] Specifically, step 102 may include the following sub-steps S21 - S22:
[0067] Step S21: Perform image encoding on the image data through an image encoder to output image global features;
[0068] Furthermore, step S21 may include the following sub-steps S211 - S214:
[0069] Step S211: Segment the image data to generate multiple image patches;
[0070] Step S212: Perform linear embedding on each image patch to generate multiple image embeddings;
[0071] Step S213: Use multiple image embeddings to form an image embedding sequence, and insert a preset first classification token into the image embedding sequence to output a marked image embedding sequence;
[0072] Step S214: Perform Transformer encoding on the marked image embedding sequence to generate image global features.
[0073] It should be noted that for the processing process of the image encoder (ViT) for the input image: Input: 224×224 RGB image, segmented into 16×16 patches (a total of 196). Processing flow: (1) Patch linear embedding: Convert each patch into a 768-dimensional vector. (2) Add [CLS] Token (i.e., the preset first classification token): Used to aggregate global features. (3) Transformer encoding: 12-layer ViT blocks extract multi-scale features. Output: [CLS] Token feature (768-dimensional): As the image global representation (i.e., the image global feature). Patch features (196×768): Retain local detail information (i.e., multiple image patch features).
[0074] Step S22: Perform text encoding on the text data through a text encoder to generate text global features.
[0075] Furthermore, step S22 may include the following sub-steps S221 - S224:
[0076] Step S221: Segment the text data to generate multiple text words;
[0077] Step S222: Perform linear embedding on each text word to generate multiple text word embeddings;
[0078] Step S223: Use multiple text word embeddings to form a text word embedding sequence, and insert a preset second classification token into the text word embedding sequence to output a marked text word embedding sequence;
[0079] Step S224: Perform Transformer encoding on the marked text word embedding sequence and output the text global features.
[0080] It should be noted that for the processing process of the text encoder (BERT) for the input text: Input: After text tokenization, the maximum length is 128. Processing flow: Token embedding: Convert words into 768-dimensional vectors (i.e., perform linear embedding on each text word to generate multiple text word embeddings; use multiple text word embeddings to form a text word embedding sequence, and insert a preset second classification token into the text word embedding sequence to output the marked text word embedding sequence). Transformer encoding: 12 layers of BERT extract context semantics. Output: [CLS] Token feature (768-dimensional): As the text global representation (i.e., the text global feature). Token features (128×768): Retain word-level semantic information (i.e., multiple text word features).
[0081] Step 103: Perform multi-modal feature fusion on the image global feature and the text global feature to generate the image-text fusion feature.
[0082] Specifically, step 103 may include the following sub-steps S31 - S33:
[0083] Step S31: Perform cross-attention calculation on the image global feature and the text global feature to generate the cross-attention alignment feature;
[0084] Step S32: Concatenate the cross-attention alignment feature, the image global feature, and the text global feature and output the concatenated feature;
[0085] Step S33: Reduce the dimension of the concatenated feature to generate the image-text fusion feature.
[0086] It should be noted that for the multi-modal feature fusion process: Input: [CLS] feature of ViT + [CLS] feature of BERT (i.e., the image global feature and the text global feature). Fusion strategy: (1) Cross-attention mechanism: Use the text feature as the Query, and the image feature as the Key / Value to calculate cross-modal alignment (i.e., the cross-attention alignment feature). The corresponding pseudo-code example for this fusion strategy is:
[0087]
[0088] Furthermore, concatenate the aligned feature with the original feature and reduce the dimension to 512 dimensions, that is, concatenate the cross-attention alignment feature, the image global feature, and the text global feature and output the concatenated feature; reduce the dimension of the concatenated feature and output the fused multi-modal feature (512-dimensional), that is, the image-text fusion feature.
[0089] In this embodiment, for the multi-modal feature fusion strategy: cross-attention and concatenation: In the multi-modal feature fusion stage, the cross-attention mechanism is used as Figure 3 the text features are used as Queries, and the image features are used as Keys / Values for cross-modal alignment, and then the aligned features are concatenated with the original features and dimension-reduced. This fusion strategy fully considers the semantic correlation and complementarity between different modal features, and can effectively mine the deep emotional information between text and images, which is different from the traditional methods of simple feature concatenation or static attention mechanism.
[0090] Step 104: Use the dynamic routing mixture-of-experts model to perform dynamic expert calculations on the image-text fusion features and output weighted features.
[0091] The dynamic routing mixture-of-experts model includes two cascaded fully-connected layers, a Softmax activation function layer, and a mixture-of-experts layer; the mixture-of-experts layer includes multiple experts. The two cascaded fully-connected layers and the Softmax activation function layer constitute a gating network. The multiple experts include a vision-dominant expert, a text-dominant expert, a cross-modal contradiction detection expert, and a shared basic expert. Among them, the vision-dominant expert: a 3-layer MLP, specializing in image feature enhancement (such as facial expression recognition). The text-dominant expert: a 3-layer MLP, specializing in text semantic analysis (such as sentiment word weight calculation). The cross-modal contradiction detection expert: cross-attention + residual network, detecting text-image conflicts. The shared basic expert: general feature processing to prevent routing failure.
[0092] Specifically, step 104 may include the following sub-steps S41-S44:
[0093] Step S41: Perform feature transformation on the image-text fusion features through two cascaded fully-connected layers to generate dimension-reduced features;
[0094] Step S42: Use the Softmax activation function layer to calculate weights according to the dimension-reduced features to determine the weights corresponding to multiple experts in the mixture-of-experts layer;
[0095] Step S43: Sort the weights corresponding to each expert in descending order, and select the experts corresponding to the top pre-set number of weights to perform expert calculations on the multi-modal fusion features and output multiple expert results;
[0096] Step S44: Perform weighted summation on each expert result according to the top pre-set number of weights to generate weighted features.
[0097] It should be noted that, please refer to Figure 3, For the feature processing process of the Gating Network: Input: Multimodal fusion features (512 - dimensional), i.e., image - text fusion features. Structure: Two - layer fully - connected (512→256→4), outputting expert weights. Activation function: Softmax normalizes to a probability distribution. Routing strategy: Select the Top - 2 experts and perform weighted summation. Specifically, through two cascaded fully - connected layers and a Softmax activation function layer, the weights corresponding to each expert are output. The weights corresponding to each expert are sorted in descending order, and the experts corresponding to the top pre - set number of weights (taking the value of 2) are selected for the multimodal fusion features. The expert results corresponding to the top 2 experts are output. The expert results are weighted and summed according to the weights corresponding to the top 2 experts to generate weighted features.
[0098] It is worth mentioning that for the Top - k expert selection: The MoE layer adopts a sparse activation mechanism. The gating network in the MoE layer selects Top - k (such as Top - 2) experts to participate in the calculation, only activating the expert parameters related to the input, which greatly reduces the computational amount (such as a significant reduction in FLOPs). Compared with the traditional fully - parameter - activated model, while ensuring the performance of sentiment recognition, it greatly improves the inference speed, enabling it to meet the requirements of real - time application scenarios (such as real - time review of social media, real - time interaction of intelligent customer service, etc.). This is an important means to achieve efficient multimodal sentiment recognition.
[0099] In this embodiment, for the dynamic - routing hybrid - expert model architecture: Innovative fusion: VisionTransformer (ViT) is used for image feature extraction, BERT is used for text feature extraction, and they are innovatively combined with the hybrid - expert model (MoE) to construct a unified architecture that can effectively process multimodal data (image - text). Dynamic - routing mechanism: The gating network in the MoE layer dynamically selects expert combinations according to the input multimodal fusion features, realizing adaptive modal interaction. For example, it can flexibly allocate experts with different specializations (such as vision - dominant experts, text - dominant experts, cross - modal conflict - detection experts, etc.) to participate in the calculation according to the characteristics of the image - text content, such as whether the text sentiment is fuzzy, whether there are image - text conflicts, etc., rather than using a fixed modal - fusion method. This is the key to improving the accuracy and flexibility of multimodal sentiment recognition.
[0100] Step 105: Input the weighted features into the sentiment classifier for recognition to generate the target multimodal image - text sentiment recognition result.
[0101] The sentiment classifier includes two cascaded fully - connected layers and an activation - function layer; the activation - function layer includes a cascaded ReLU activation - function layer and a Softmax activation - function layer.
[0102] Specifically, step 105 may include the following sub - steps S51 - S52:
[0103] Step S51: Perform feature transformation on the weighted features using two cascaded fully-connected layers, and output the transformed features.
[0104] Step S52: Perform a non-linear mapping on the transformed features through an activation function layer to generate the target multi-modal graphic-text sentiment recognition result.
[0105] It should be noted that for the recognition process of the sentiment classifier: Input: The weighted features (512 dimensions) output by the MoE layer. Structure: Two fully-connected layers (512→256→number of sentiment categories). Activation function: ReLU + Softmax. Output: The sentiment probability distribution (such as positive / neutral / negative), that is, the target multi-modal graphic-text sentiment recognition result.
[0106] As a comparison of technical effects, it can be referenced in combination with the existing technology. With the popularization of applications such as social media, intelligent customer service, and human-computer interaction, user-generated content (UGC, User-Generated Content) presents multi-modal characteristics (such as mixed data of graphics, text, and video). Traditional single-modal sentiment analysis (such as text-based sentiment classification) is difficult to comprehensively capture the user's intention, and there are obvious deficiencies especially in the following scenarios: Graphic-text contradiction: For example, the caption "I'm so happy" with a "smiling emoji" may be ironic. Modal complementarity: For example, the intonation of speech (high-pitched / low-pitched) can assist in judging the true emotion of the text (neutral words).
[0107] For the current technical challenges: ① Modal alignment: The feature spaces of different modalities are heterogeneous (such as image pixels vs text word vectors), and effective fusion is required. ② Computational efficiency: The number of parameters of multi-modal models is large (such as the number of parameters of ViLBERT (Vision-and-Language BERT) reaches 320 million), making it difficult to deploy in real time. ③ Model capacity: Complex sentiment patterns (such as implicit sentiment, cultural difference expressions) require higher-capacity models to capture.
[0108] In the evolution of existing technologies, (1) Early methods (2010): Feature concatenation: Concatenate text TF-IDF (Term Frequency-Inverse Document Frequency) and image HOG (Histogram of Oriented Gradients) features and input them into an SVM (Support Vector Machine) classifier. Limitations: Ignore cross-modal interactions and cannot handle non-linear relationships. (2) Deep learning era (2017 - 2020): Attention mechanism: Align features through cross-modal attention (such as bilinear attention). Representative models: ViLBERT (joint image-text training), LXMERT (Language-Xtreme Multimodal Encoding and Representation Transformer). Limitations: Static attention weights and cannot adapt to dynamic scenarios. (3) Transformer unified architecture (2021 - present): Multimodal Transformer: CLIP (Contrastive Language-Image Pre-training, text-image contrastive learning), FLAVA (Foundation Language-And-Vision Alignment, unified multimodal pre-training). Limitations: High computational cost due to full parameter activation and difficult to scale.
[0109] Furthermore, regarding the existing ViLBERT-based multimodal sentiment analysis system: (1) Architecture: Image encoder: Faster R-CNN extracts regional features. Text encoder: BERT processes text. Cross-modal interaction: Fuse image and text features through a co-attention layer. (2) Training objective: Multi-task learning: Sentiment classification + image-text matching task. (3) Performance: On the CMU-MOSEI dataset, the sentiment classification accuracy is 81.2%.
[0110] Based on the above foundation, the shortcomings of existing technologies are as follows: (1) Insufficient modal alignment and interaction: Existing methods (such as ViLBERT and traditional multimodal MoE) rely on static fusion mechanisms (such as fixed attention weights or rule-based expert allocation) and cannot dynamically adjust modal weights according to the input content. For example, in image-text conflict scenarios (such as "smiley faces with negative text"), it is impossible to adaptively select the dominant modality. Typical manifestations: ignoring visual clues when text is dominant, or misjudging text emotions when visual is dominant. Insufficient cross-modal feature fusion leads to increased classification error rates in conflicting scenarios (such as irony detection accuracy <70%). (2) Low computational efficiency: Models with full parameter activation (such as ViLBERT) need to calculate all modules during inference, resulting in resource waste. For example, when processing high-resolution images, FLOPs (Floating Point Operations) are as high as 6.4G, which is difficult to run in real time on edge devices. Typical manifestations: The single-sample inference delay exceeds 100ms (NVIDIA V100 GPU), which cannot meet real-time interaction requirements (such as live barrage sentiment analysis). High energy consumption limits mobile deployment (such as a sharp increase in smartphone battery consumption). :(3) Limited model capacity: Traditional architectures (such as ViLBERT) have fixed parameter sizes, and the model needs to be reconstructed when expanded, which can easily cause training instability (gradient explosion / vanishing). For example, when the number of parameters exceeds 300 million, the difficulty of training convergence increases significantly. Typical manifestations: In complex emotional scenarios (such as metaphorical expressions caused by cultural differences), the model accuracy drops by more than 15%. When the number of experts is expanded, the utilization rate of some experts is less than 10% (load imbalance). (4) Poor interpretability: Problem description: Most existing models are black box structures, and the basis for decision-making cannot be traced. For example, users cannot know whether the classification result is dominated by text keywords, image expressions, or cross-modal interactions. Typical manifestations: It is difficult to locate the source of the error when the model misjudges (such as mistakenly classifying a "red alert" with a fire picture as "positive"). Lack of visualization tools to assist manual review (such as social media content security review).
[0111] The existing multimodal emotion recognition technology described above faces significant bottlenecks in multimodal dynamic interaction, computational efficiency and model capacity. The present invention provides a multimodal image-text emotion recognition method based on a dynamic routing hybrid expert model. By introducing a dynamic routing hybrid expert model and combining the multimodal encoding capabilities of ViT and BERT, input-adaptive expert selection and sparse computing are realized, which surpasses existing solutions in accuracy, efficiency and interpretability, and has significant technological progress and industrial application value.
[0112] Specifically, to achieve dynamic modal interaction and improve computational efficiency: Design a gating network based on MoE to dynamically select a combination of experts according to the input features. For example, when detecting ambiguous text sentiment, the visual experts (such as the facial expression recognition expert) are preferentially activated. When there is a conflict between text and image, the cross-modal contradiction detection expert is activated. Expand the model capacity: Heterogeneous expert design: Allow different experts to use different structures (such as using a CNN expert to process local textures and a Transformer expert to process long-range dependencies). Dynamic expansion mechanism: Automatically increase the number of experts during training (such as starting with 4 and expanding to 16 according to the task complexity). Increase interpretability: Visualization of routing heatmaps: Show the expert-category associations (such as expert 3 being specialized in the "sad" category). Decision traceability tool: Record the experts activated during the reasoning process and their contribution weights.
[0113] Furthermore, the comparison between the present invention and the prior art is shown in Table 1 as follows:
[0114] Table 1 Comparison between the present invention and the prior art
[0115]
[0116] In summary, compared with the prior art, the present invention has the following advantages: 1. Modal interaction and fusion: Dynamic adaptability: The present invention adaptively selects experts according to the input through the MoE gating network, can handle the complex relationships between text and image, while most of the prior art is static fusion and is prone to inaccuracies when dealing with complex scenarios. Sufficient interaction: The present invention uses cross-attention to deeply align text and image features and then splices them, mining emotional clues more fully, while the fusion methods in the prior art are relatively shallow and the information fusion is insufficient. 2. Computational efficiency: Sparse activation: The MoE layer of the present invention sparsely activates and selects the top-k experts, significantly reducing the computational amount (the FLOPs can be reduced by 40%), meeting real-time applications, while the prior art fully activates all parameters, resulting in high computational costs and slow inference. Hardware optimization: The present invention is convenient for accelerating using hardware sparse computing kernels, while the prior art is limited in hardware optimization, with high energy consumption and difficulty in improving the inference speed. 3. Model capacity and scalability: Scalable architecture: The present invention can increase the number of experts and expand the capacity through heterogeneous design to adapt to complex emotional patterns, while the architectures of the prior art are fixed and require reconstruction for expansion, which is prone to unstable training. Load balancing: The present invention sets a load balancing loss to ensure load balancing among experts and improve performance stability, while the prior art may lack an effective load balancing mechanism. 4. Interpretability: Visualization: The present invention can generate routing heatmaps to show the associations between expert selections and emotions, facilitating manual review and understanding, while most of the prior art is a black-box model and it is difficult to trace the decision-making basis. Decision traceability: The present invention can record the activated experts and their contribution weights to assist in traceability, enhancing trust, while the prior art is difficult to provide such detailed information.
[0117] In an embodiment of the present invention, the present invention provides a multi-modal image-text sentiment recognition method based on a dynamic routing mixture of experts model. First, image data and text data are obtained, and the image data and text data are input into a preset multi-modal image-text sentiment recognition network. The preset multi-modal image-text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier. Then, the image data and text data are encoded by the target encoder to output an image global feature and a text global feature. The image global feature and the text global feature are subjected to multi-modal feature fusion to generate an image-text fusion feature. The dynamic routing mixture of experts model is used to perform dynamic expert calculation on the image-text fusion feature to output a weighted feature. Finally, the weighted feature is input into the sentiment classifier for recognition to generate a target multi-modal image-text sentiment recognition result. Based on the above solution, in the process of processing the obtained image data and text data by the target encoder, the dynamic routing mixture of experts model, and the sentiment classifier to output the target multi-modal image-text sentiment recognition result, the present invention adaptively selects an expert according to the input feature by the dynamic routing mixture of experts model to perform expert calculation, so as to process the complex relationship between images and texts, and further improve the recognition effect of the model.
[0118] For better illustration, refer to Figure 4 , which shows a flowchart of the steps for training the model of the preset multi-modal image-text sentiment recognition network provided in the second embodiment of the present invention. This process may include the following steps:
[0119] Step 401, obtain a training image data set and a training text data set;
[0120] Step 402, use a preset loss function and an initial multi-modal image-text sentiment recognition network to determine a model gradient according to the training image data set and the training text data set;
[0121] Step 403, update the model parameters of the initial multi-modal image-text sentiment recognition network using the model gradient to determine an intermediate multi-modal image-text sentiment recognition network, and count the number of iterative model updates in real time;
[0122] Step 404, determine whether the number of model updates reaches a preset number of training times;
[0123] Step 405, if it reaches, use the intermediate multi-modal image-text sentiment recognition network as the trained preset multi-modal image-text sentiment recognition network.
[0124] It should be noted that the training image dataset includes multiple image data for model training, and the training text dataset includes multiple text data for model training. Through the initial multi-modal image-text sentiment recognition network, based on the training image dataset and the training text dataset, multiple training multi-modal image-text sentiment recognition results and multiple expert activation frequencies are output (including the activation frequency of the vision-dominant expert, the activation frequency of the text-dominant expert, the activation frequency of the cross-modal contradiction detection expert, and the activation frequency of the shared basis expert). The expert activation frequency is the number of times an expert is activated divided by the total number of samples in the training image dataset and the training text dataset. For example, when the sum of the number of samples in the training image dataset and the training text dataset (total number of samples) is 10, and the vision-dominant expert in the mixture of experts layer is selected 5 times, it indicates that the expert activation frequency of the vision-dominant expert is 0.5.
[0125] Furthermore, substituting the multiple training multi-modal image-text sentiment recognition results and the multiple expert activation frequencies into the preset loss function and taking the derivative, the model gradient is obtained; among them, the preset loss function is specifically:
[0126] ;
[0127] Among them, is the loss value corresponding to the total loss function (i.e., the preset loss function).
[0128] Sentiment classification loss function: Cross-entropy loss (main loss):
[0129] ;
[0130] Among them, is the loss value corresponding to the sentiment classification loss function; is the true label of the i-th sample; is the probability of the i-th sample predicted by the model, representing the i-th training multi-modal image-text sentiment recognition result.
[0131] Load balancing loss function: Penalize the variance of expert utilization:
[0132] ;
[0133] Among them, is the loss value corresponding to the load balancing loss function; is a coefficient used to adjust the weight of the load balancing loss in the overall calculation; is the activation frequency of expert i; N is the total number of experts (N = 4).
[0134] Further, if the number of model updates does not reach the preset number of training times, the intermediate multi-modal image-text sentiment recognition network is used as the new initial preset multi-modal image-text sentiment recognition network, and the process jumps to step 402 until the number of model updates reaches the preset number of training times. The intermediate multi-modal image-text sentiment recognition network determined when the number of model updates reaches the preset number of training times is used as the trained preset multi-modal image-text sentiment recognition network. Herein, the preset number of training times can be set as needed, and the present invention does not make specific limitations thereto.
[0135] It is worth mentioning that for the learning rate scheduling: the initial learning rate is 3e-5, and it decays by 0.9 times every 10 Epochs. For gradient clipping: the threshold is 1.0 to prevent gradient explosion. For hardware acceleration: NVIDIA A100 GPU (NVIDIA Ampere Architecture 100) is used for mixed-precision training (FP16, 16-bit Floating Point).
[0136] In the embodiments of the present invention, the present invention designs a multi-task loss function including a sentiment classification loss and a load balancing loss. The sentiment classification loss uses a cross-entropy loss function to measure the accuracy of the model for sentiment classification, while the load balancing loss ensures the load balance of each expert by penalizing the variance of the expert utilization rate, preventing some experts from being idle or overloaded. By weighted summing the two, the total loss function is obtained, enabling the model to not only pursue accurate sentiment classification results during training but also take into account the reasonable utilization of expert resources, thereby improving the stability and overall performance of the model.
[0137] Please refer to Figure 5 , Figure 5 which is the structural block diagram of a multi-modal image-text sentiment recognition system based on a dynamic routing mixture of experts model provided in Embodiment III of the present invention.
[0138] A multi-modal image-text sentiment recognition system based on a dynamic routing mixture of experts model provided by the present invention includes:
[0139] An acquisition module 501, configured to acquire image data and text data, and input the image data and text data into a preset multi-modal image-text sentiment recognition network, where the preset multi-modal image-text sentiment recognition network includes a target encoder, a dynamic routing mixture of experts model, and a sentiment classifier;
[0140] An encoding module 502, configured to encode the image data and text data through the target encoder, and output an image global feature and a text global feature;
[0141] A fusion module 503, configured to perform multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature;
[0142] An expert calculation module 504, which is used to perform dynamic expert calculation on the image-text fusion features by using a dynamic routing mixture of experts model, and output weighted features;
[0143] An identification module 505, which is used to input the weighted features into a sentiment classifier for identification, and generate a target multi-modal image-text sentiment recognition result.
[0144] Furthermore, the target encoder includes an image encoder and a text encoder; the encoding module 502 includes:
[0145] A first sub-module, which is used to perform image encoding on the image data through the image encoder, and output image global features;
[0146] A second sub-module, which is used to perform text encoding on the text data by using the text encoder, and generate text global features.
[0147] Furthermore, the first sub-module is specifically used for:
[0148] Segment the image data to generate multiple image patches;
[0149] Perform linear embedding on each image patch to generate multiple image embeddings;
[0150] Use multiple image embeddings to form an image embedding sequence, and insert a preset first classification token into the image embedding sequence, and output a marked image embedding sequence;
[0151] Perform Transformer encoding on the marked image embedding sequence to generate image global features.
[0152] Furthermore, the second sub-module is specifically used for:
[0153] Perform word segmentation on the text data to generate multiple text words;
[0154] Perform linear embedding on each text word to generate multiple text word embeddings;
[0155] Use multiple text word embeddings to form a text word embedding sequence, and insert a preset second classification token into the text word embedding sequence, and output a marked text word embedding sequence;
[0156] Perform Transformer encoding on the marked text word embedding sequence, and output text global features.
[0157] Furthermore, the fusion module 503 is specifically used for:
[0158] Perform cross-attention calculation on the image global features and the text global features to generate cross-attention alignment features;
[0159] Concatenate the cross-attention alignment features, the global image features, and the global text features, and output the concatenated features;
[0160] Reduce the dimension of the concatenated features to generate the image-text fusion features.
[0161] Furthermore, the dynamic routing mixture-of-experts model includes two serially connected fully connected layers, a Softmax activation function layer, and a mixture-of-experts layer; the expert calculation module 504 is specifically used for:
[0162] Perform feature transformation on the image-text fusion features through two serially connected fully connected layers to generate the reduced-dimension features;
[0163] Use the Softmax activation function layer to calculate weights based on the reduced-dimension features, and determine the weights corresponding to multiple experts in the mixture-of-experts layer;
[0164] Sort the weights corresponding to each expert in descending order, select the experts corresponding to the top pre-set number of weights to perform expert calculations on the multi-modal fusion features, and output multiple expert results;
[0165] Perform weighted summation on each expert result according to the top pre-set number of weights to generate the weighted features.
[0166] Furthermore, the sentiment classifier includes two serially connected fully connected layers and an activation function layer; the recognition module 505 is specifically used for:
[0167] Perform feature transformation on the weighted features through two serially connected fully connected layers, and output the transformed features;
[0168] Perform non-linear mapping on the transformed features through the activation function layer to generate the target multi-modal image-text sentiment recognition result.
[0169] In an optional system embodiment, it further includes:
[0170] The first module is used to obtain the training image dataset and the training text dataset;
[0171] The second module is used to determine the model gradient according to the training image dataset and the training text dataset by using the pre-set loss function and the initial multi-modal image-text sentiment recognition network;
[0172] The third module is used to update the model parameters of the initial multi-modal image-text sentiment recognition network by using the model gradient, determine the intermediate multi-modal image-text sentiment recognition network, and count the number of iterative model updates in real time;
[0173] The fourth module is used to determine whether the number of model updates reaches the pre-set number of training times;
[0174] The fifth module is configured to use the intermediate multi-modal image-text sentiment recognition network as the trained pre-set multi-modal image-text sentiment recognition network if the condition is met.
[0175] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and sub-modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0176] An embodiment of the present invention further provides a computer device, including a memory and a processor, where a computer program is stored in the memory; when the computer program is executed by the processor, the processor is caused to execute the steps of the multi-modal image-text sentiment recognition method based on the dynamic routing mixture-of-experts model as described in any of the foregoing embodiments.
[0177] In several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0178] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0179] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-modal image-text sentiment recognition method based on a dynamic routing mixture of experts model, characterized in that Including: Obtain image data and text data, and input the image data and the text data into a preset multi-modal image-text sentiment recognition network, where the preset multi-modal image-text sentiment recognition network includes a target encoder, a dynamic routing mixture-of-experts model, and a sentiment classifier; Encode the image data and the text data through the target encoder, and output an image global feature and a text global feature; Perform multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature; Perform dynamic expert calculation on the image-text fusion feature by using the dynamic routing mixture-of-experts model, and output a weighted feature; Input the weighted feature into the sentiment classifier for recognition to generate a target multi-modal image-text sentiment recognition result.
2. The multimodal graphic-text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 1, wherein The target encoder includes an image encoder and a text encoder; the encoding of the image data and the text data through the target encoder to output an image global feature and a text global feature includes: Perform image encoding on the image data through the image encoder to output an image global feature; Perform text encoding on the text data by using the text encoder to generate a text global feature.
3. The multimodal graphic-text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 2, characterized in that, The performing image encoding on the image data through the image encoder to output an image global feature includes: Segment the image data to generate a plurality of image patches; Perform linear embedding on each of the image patches to generate a plurality of image embeddings; Use the plurality of image embeddings to form an image embedding sequence, and insert a preset first classification token into the image embedding sequence to output a token image embedding sequence; Perform Transformer encoding on the token image embedding sequence to generate an image global feature.
4. The multimodal image-text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 2, wherein The performing text encoding on the text data by using the text encoder to generate a text global feature includes: Perform word segmentation on the text data to generate a plurality of text words; Perform linear embedding on each of the text words to generate a plurality of text word embeddings; Use the plurality of text word embeddings to form a text word embedding sequence, and insert a preset second classification token into the text word embedding sequence to output a token text word embedding sequence; Perform Transformer encoding on the token text word embedding sequence to output a text global feature.
5. The multimodal graphic and text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 1, wherein, The performing multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature includes: Perform cross-attention calculation on the image global feature and the text global feature to generate a cross-attention alignment feature; Concatenate the cross-attention alignment feature, the image global feature, and the text global feature, and output a concatenated feature; Reduce the dimension of the concatenated feature to generate an image-text fusion feature.
6. The multimodal graphic-text sentiment recognition method based on the dynamic routing mixture-of-experts model according to claim 1, wherein, The dynamic routing mixture-of-experts model includes two serially connected fully connected layers, a Softmax activation function layer, and a mixture-of-experts layer; the performing dynamic expert calculation on the image-text fusion feature by using the dynamic routing mixture-of-experts model to output a weighted feature includes: Perform feature transformation on the image-text fusion feature through the two cascaded fully connected layers to generate a dimensionality-reduced feature; Use the Softmax activation function layer to calculate weights based on the dimensionality-reduced feature to determine the weights corresponding to multiple experts in the mixture-of-experts layer; Sort the weights corresponding to each expert in descending order, select the experts corresponding to the top preset number of weights to perform expert calculations on the multi-modal fusion feature, and output multiple expert results; Perform weighted summation on each of the expert results according to the top preset number of weights to generate a weighted feature.
7. The multimodal graphic and text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 1, wherein The sentiment classifier includes two cascaded fully connected layers and an activation function layer; inputting the weighted feature into the sentiment classifier for recognition to generate a target multi-modal image-text sentiment recognition result includes: Perform feature transformation on the weighted feature using the two cascaded fully connected layers to output a transformed feature; Perform non-linear mapping on the transformed feature through the activation function layer to generate a target multi-modal image-text sentiment recognition result.
8. The multimodal graphic-text sentiment recognition method based on the dynamic routing mixture of experts model according to claim 1, wherein The model training process of the preset multi-modal image-text sentiment recognition network includes: Obtain a training image data set and a training text data set; Use a preset loss function and an initial multi-modal image-text sentiment recognition network to determine model gradients based on the training image data set and the training text data set; Use the model gradients to update the model parameters of the initial multi-modal image-text sentiment recognition network to determine an intermediate multi-modal image-text sentiment recognition network, and count the number of iterative model updates in real time; Judge whether the number of model updates reaches a preset number of training times; If it reaches, use the intermediate multi-modal image-text sentiment recognition network as the trained preset multi-modal image-text sentiment recognition network.
9. A multi-modal text-image sentiment recognition system based on a dynamic routing mixture-of-experts model, characterized in that Includes: An acquisition module for acquiring image data and text data and inputting the image data and the text data into a preset multi-modal image-text sentiment recognition network, where the preset multi-modal image-text sentiment recognition network includes a target encoder, a dynamic routing mixture-of-experts model, and a sentiment classifier; An encoding module for encoding the image data and the text data through the target encoder to output an image global feature and a text global feature; A fusion module for performing multi-modal feature fusion on the image global feature and the text global feature to generate an image-text fusion feature; An expert calculation module for performing dynamic expert calculation on the image-text fusion feature using the dynamic routing mixture-of-experts model to output a weighted feature; A recognition module for inputting the weighted feature into the sentiment classifier for recognition to generate a target multi-modal image-text sentiment recognition result.
10. A computer device, characterized in that, Includes a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor executes the steps of the multi-modal image-text sentiment recognition method based on a dynamic routing mixture-of-experts model according to any one of claims 1-8.
Citation Information
Patent Citations
Social media sentiment analysis method and system based on multi-modal feature fusion
CN112508077A
Pulmonary nodule image detection method and system based on CT image
CN113888466A
Cross-modal multilayer fusion emotion recognition method and system
CN118861773A
Multi-modal named entity identification method and device
CN119337883A
Cited By
Hierarchical hybrid expert model-based reasoning method and system, and storage medium
CN120471184A
Reasoning methods, systems, and storage media based on hierarchical hybrid expert models
CN120471184B
Multi-modal data prediction method and device based on hybrid expert attention network
CN120974259A
Target positioning method, target positioning device and computer storage medium
CN121392231A