High-performance sentiment analysis method and device based on multi-modal data alignment and medium

Through the combination of deep reinforcement learning and dynamic time regularization algorithm, the problem of modal features is not synchronized in multimodal sentiment analysis is solved, efficient and accurate multimodal alignment is achieved, and the performance of the sentiment analysis model is improved.

CN120162622APending Publication Date: 2025-06-17HEFEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510221467.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In multimodal sentiment analysis, there are problems such as abnormality or inconsistency in modal features such as text, audio and video, which leads to a degradation of model performance. It is difficult for traditional alignment methods to effectively solve this problem.

Method used

Deep reinforcement learning (DQN) combined with dynamic time regularization (DTW) algorithm is used to construct action selection strategies through DQN, generate the optimal DTW distance, guide multimodal time series state vector alignment, and achieve efficient alignment between modals.

Benefits of technology

It improves the accuracy and efficiency of multimodal data alignment, reduces the computational complexity and time overhead, is suitable for multi-type application scenarios, and enhances the accuracy and robustness of alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162622A_ABST
    Figure CN120162622A_ABST
Patent Text Reader

Abstract

The invention discloses a high-performance sentiment analysis method and device based on multi-modal data alignment and a medium, and the method comprises the steps: taking ResNet-50 as a backbone network, and extracting multi-modal time sequence features; inputting the extracted time sequence features into a state presentation layer of the trained DQN to generate a time sequence state vector; and constructing an action selection strategy by using DQN to obtain an optimal DTW distance, and guiding state vector alignment of the multi-modal time sequence. According to the efficient multi-modal alignment method based on deep reinforcement learning, compared with a traditional alignment method, the DQN algorithm can intelligently adjust the alignment path, the calculation requirement is lower, and the intelligent alignment algorithm is more suitable for multiple types of application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data alignment sentiment analysis, and in particular, to a high-performance sentiment analysis method, device and medium based on multimodal data alignment. Background Art

[0002] Multimodal sentiment analysis is an extended field of traditional text-based sentiment analysis. In addition to text features, other modal features such as speech and vision are also considered. Multimodal sentiment analysis is no longer limited to traditional single-modal tasks, but integrates multiple modal features, which is closer to the process of human understanding of the world, and thus has attracted more and more attention. In the process of multimodal sentiment analysis, various modal features such as text, audio, and video are involved, and there are phenomena of out-of-sync or inconsistency, which may have a negative impact on the model performance. The introduction of multimodal alignment methods can effectively alleviate the occurrence of such problems.

[0003] In recent years, the emergence of deep learning has provided new ideas for multimodal alignment problems. Multimodal alignment is divided into explicit alignment and implicit alignment. In contrast, the implicit alignment method provides a flexible and general modal alignment method for multimodal sentiment analysis tasks and is widely used in various multimodal sentiment analysis fields. The attention mechanism is used to achieve the alignment between the opinion target and the image, so as to obtain a visual representation sensitive to the opinion target.

[0004] However, the same emotion may be expressed by different visual representations for different opinion targets. This method ignores the granularity differences between multimodals (such as text and image), and thus the inconsistency of opinion target granularity may cause the attention mechanism to fail to capture the corresponding visual representation.

[0005] Extracting coarse-grained and fine-grained opinion targets from scene images can enable the model to better learn the alignment relationship between the opinion target and the image features. The diversity of visual representations inevitably leads to its sparsity, which makes it difficult to learn the precise mapping relationship between visual representations and sentiment labels. In addition, the semantic relationship inconsistency between the text and image modalities will make the information of the image modality become noise, resulting in a decline in the overall performance of the model. Currently, when performing multimodal alignment, directly reducing the dimension of high-dimensional data using PCA or other methods may lead to incomplete data retention and inaccurate alignment. Using the DTW distance between modalities as a metric can enable the agent of reinforcement learning to learn a data reduction method by itself to achieve the alignment operation between modalities to retain data to the greatest extent.

[0006] For example, the patent document with the application number 202410901937.X discloses a data alignment processing method and device for multimodal fusion. Using the solution of this application, it is possible to achieve that in the mutual generation of images and texts, generating samples with high similarity can also promote the data alignment between the two modalities and improve the sufficiency of multimodal feature fusion. However, there are also problems in its solution: the model cannot better learn the alignment relationship between the opinion target and the image features. Summary of the Invention

[0007] The purpose of the present invention is to provide a high-performance sentiment analysis method, device and medium based on multimodal data alignment, using the DTW distance between modalities as a metric to achieve the alignment operation between multimodals.

[0008] Embodiments of the present invention provide a high-performance sentiment analysis method, device and medium based on multimodal data alignment.

[0009] First aspect: A high-performance sentiment analysis method based on multimodal data alignment, including:

[0010] S1. Using ResNet-50 as the backbone network, extract multimodal time series features;

[0011] S2. Input the extracted time series features into the state representation layer of the trained DQN to generate a time series state vector;

[0012] S3. Use DQN to construct an action selection strategy to obtain the optimal DTW distance and guide the alignment of the multimodal time series state vector.

[0013] Further, the step of using ResNet-50 as the backbone network to extract multimodal time series features includes:

[0014] Extract multimodal features from the output of the Conv4c layer, stack the context frame features along the time dimension to obtain a combined feature. After the combined feature aggregates time information through a convolutional layer and is pooled by a pooling layer, it outputs multimodal time series features through 2 fully connected layers and 1 linear projection layer.

[0015] Further, the DQN structure includes:

[0016] Multiple layers of fully connected layers. After each layer of fully connected layer, the ReLU activation function is used as the activation layer to output the Q-value estimation of each step of action, and obtain the cumulative return of the Q-value estimation of each step of action.

[0017] Further, the step of using DQN to construct an action selection strategy in S3 includes:

[0018] Based on the Q-value estimation output by DQN, use the ε-greedy strategy to select actions of moving up, down or diagonally.

[0019] Further, the guiding of the multi-modal time series state vector alignment in S3 includes:

[0020] If the current alignment state minimizes the local distance between the two time series, a positive reward is given; otherwise, a negative reward is given. Based on the immediate reward value of each step of the action, the accuracy of the alignment of the two time series is designed; and the Q-value estimate of each state-action pair is updated through the Bellman equation.

[0021] Further, when the DQN is trained, the joint loss function includes the DTW loss and the reinforcement learning reward function, and the formula is expressed as:

[0022] L = L DTW + λL RL

[0023] where L DTW is the DTW loss, and λL RL is the reinforcement learning reward function, and λ is used to balance the DTW loss and the reinforcement learning reward function;

[0024] Further, the:

[0025] DTW loss, the formula is expressed as:

[0026]

[0027] where N represents the number of time steps of modality 1, M represents the number of time steps of modality 2, x i represents the feature vector at the i-th time step of modality 1, and y j represents the feature vector at the j-th time step of modality 2.

[0028] The reinforcement learning reward function, the formula is expressed as:

[0029]

[0030] where T represents the number of time steps, γ represents the discount factor, and R t represents the reward at the t-th time step. Specifically, the DTW loss learns the alignment path of each pair of modalities by minimizing the time distance between modalities.

[0031] Further, the: In S2, generating the time series state vector, the formula is expressed as:

[0032]

[0033] where is the global average pooling feature map, is the max pooling feature map, and C1D k1D convolution with a convolution kernel size of k.

[0034] Second aspect: An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method provided in the first aspect.

[0035] Third aspect: A non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method provided in the first aspect.

[0036] Advantages of the present invention:

[0037] 1. The efficient multi-modal alignment method based on deep reinforcement learning in the present invention. Compared with traditional alignment methods, the DQN algorithm can intelligently adjust the alignment path with lower computational requirements. For datasets for multi-modal alignment that usually contain complex data of different modalities and have a risk of alignment errors, the intelligent alignment algorithm of the present invention is more applicable in various types of application scenarios.

[0038] 2. The method of the present invention combines the dynamic time warping (DTW) algorithm and deep reinforcement learning technology at the same time, effectively improving the accuracy and efficiency of multi-modal data alignment, while reducing computational complexity and time overhead. For the consistency problem in the processing of different modal features by traditional methods in multi-modal alignment, the method of the present invention proposes an efficient alignment module based on the attention mechanism. Without increasing additional computational costs, through the introduction of the DQN algorithm for dynamic adjustment, efficient alignment is achieved. Embedding the algorithm into the proposed DQN-DTW network realizes the dual benefits of the high efficiency and flexibility of the network.

[0039] 3. The present invention introduces a joint loss function, which combines the DTW loss and the reinforcement learning reward function. The DTW loss considers the time series matching between modalities, while the reinforcement learning reward function not only considers the optimization of the alignment path by dynamically adjusting the alignment strategy, but also further enhances the accuracy and robustness of the alignment by increasing the differences between modalities. Description of the drawings

[0040] Figure 1 It is a schematic flowchart of the high-performance sentiment analysis method based on multi-modal data alignment of the present invention;

[0041] Figure 2 It is a schematic diagram of the ResNet-50 backbone network structure of the present invention;

[0042] Figure 3 It is a schematic diagram of the deep Q-network (DQN) structure of the present invention;

[0043] Figure 4Schematic diagram of the principle of the DQN-DTW algorithm of the present invention;

[0044] Figure 5 Schematic diagram of the structure of the high-performance sentiment analysis device based on multi-modal data alignment of the present invention;

[0045] Figure 6 Schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners

[0046] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0047] Existing methods ignore the granularity differences between multi-modalities (such as text and image), resulting in the attention mechanism being unable to capture the corresponding visual representations, leading to incomplete data retention and inaccurate multi-modal data alignment.

[0048] In view of the above problems, the present invention provides a high-performance sentiment analysis method based on multi-modal data alignment. Figure 1 Schematic diagram of the flow of the high-performance sentiment analysis method based on multi-modal data alignment provided by the embodiment of the present invention. The method includes:

[0049] S1. Using ResNet-50 as the backbone network, extract multi-modal time series features.

[0050] As Figure 2 shown, using ResNet-50 as the backbone network, extract multi-modal features from the output of the Conv4c layer. The extracted feature size is 14, 14, 14, 1024. Then, stack k context frame features along the time dimension to obtain combined features. Aggregate the time information of the combined features through two 3D convolutional layers, followed by a 3D global max pooling layer, and output multi-modal time series features through 2 fully connected layers and 1 linear projection layer. Each multi-modal time series feature has 128 dimensions.

[0051] S2. Input the extracted time series features into the state representation layer of the trained DQN to generate a time series state vector.

[0052] Construct DQN and train it. DQN includes a state representation layer, and the extracted multi-modal time series features can be input into the state representation layer of DQN.

[0053] The state representation layer represents the local features of the multi-modal time series as vectors according to the multi-modal time series alignment state at the current time step, and generates a time series state vector.

[0054] Specifically, the state vector is composed of the index of the current multi-modal time series alignment position and the features of the previous alignment state. The state includes the time step t and the current cumulative alignment path.

[0055] When generating the time series state vector, DQN first uses global average pooling and max pooling to generate two different feature maps and Then, they are processed separately through a shared one-dimensional convolution to learn the weights of each modality; then, element-wise summation is used to merge the output feature vectors, and after a Sigmoid activation operation, the final modality time series state vector graph M is generated m .

[0056] The process formula is expressed as:

[0057]

[0058] Among them, is the global average pooling feature map, is the max pooling feature map, and C1D k is a 1D convolution with a convolution kernel size of k.

[0059] S3. Use DQN to construct an action selection strategy to obtain the optimal DTW distance and guide the alignment of the multi-modal time series state vector.

[0060] As Figure 3 shown, DQN is a decision-making network based on a deep neural network, and its input is the multi-modal time series state vector.

[0061] DQN is composed of multiple fully connected layers (for example, 3 layers as shown in the figure). After each layer, the ReLU activation function is used as the activation layer, and the Q-value estimates for each step of the action (such as moving up, down, or diagonally) are output. Use DQN to estimate the cumulative rewards that may be obtained after taking different actions in the current state.

[0062] Based on the Q-value estimates output by DQN at each step, DQN constructs an action selection strategy, and the action selection strategy uses the ε-greedy strategy to select actions.

[0063] Specifically, in the initial stage, the algorithm will randomly select actions with a higher probability to promote exploration; in the later stage, it will select the action with the largest current Q-value estimate with a higher probability to guide the alignment of the multi-modal time series state vector; the ε value gradually decreases as the training progresses to ensure that more learned strategies are used in the later stage of training.

[0064] Update the Q value based on the Bellman equation to guide the alignment of the multi-modal time series state vector. Specifically:

[0065] Each step of the action generates an immediate reward value, and the design of the reward value is based on the alignment accuracy. If the current alignment state minimizes the local distance between the two time series, a higher positive reward is given; otherwise, a negative reward is given. The experience replay mechanism is adopted, and samples drawn from the experience pool are used to update the DQN parameters, and the Q-value estimates of each state-action pair are updated through the Bellman equation.

[0066] Using the trained DQN, based on the time series alignment path generated by the optimal action selection strategy, calculate the dynamic time warping distance (DTW) between the two time series.

[0067] Figure 4 It is the framework for solving the DTW process of the entire DQN. Among them, the DQN is used to guide the selection of the alignment path. By selecting the path with the largest Q value, it is ensured that the selected path minimizes the alignment error between the two time series, and finally an optimal path is generated and the DTW distance is output.

[0068] The present invention solves dynamic time warping (DTW) through a deep Q-network (DQN), and uses frame-level and video-level cues for self-supervised video representation learning.

[0069] During DQN training, the joint loss function includes the DTW loss and the reinforcement learning reward function.

[0070] The DTW loss maintains an alignment path for each pair of modalities, and optimizes the model by minimizing the time distance between each pair of modalities. The DTW loss formula is expressed as:

[0071]

[0072] Among them, N represents the number of time steps of modality 1, M represents the number of time steps of modality 2, x i represents the feature vector of the i-th time step of modality 1, and y j represents the feature vector of the j-th time step of modality 2. By minimizing the DTW loss, the feature vectors of different modalities will be pulled towards their corresponding time steps, thereby reducing the overall alignment error.

[0073] The reinforcement learning reward function aims to enhance the alignment accuracy and robustness between different modalities by dynamically adjusting the alignment strategy. The reinforcement learning reward function formula is expressed as:

[0074]

[0075] Among them, T represents the number of time steps, γ represents the discount factor, and R t represents the reward at the t-th time step. Specifically, the DTW loss learns the alignment path of each pair of modalities by minimizing the time distance between modalities.

[0076] This encourages the modal features belonging to the same time step to be close in the feature space. The reinforcement learning reward function aims to increase the alignment accuracy between modal features in the feature space by dynamically adjusting the alignment strategy, so that they are well separated in the feature space. This helps to improve the discriminative ability of the feature representation and makes it easier to distinguish different modalities in the dataset.

[0077] The joint loss function is expressed by the formula:

[0078] L = L DTW + λL RL

[0079] where L DTW is the DTW loss, and λL RL is the reinforcement learning reward function. λ is used to balance the DTW loss and the reinforcement learning reward function. In the experiments of this method, λ can be set to 0.01 according to experience.

[0080] The method of the present invention combines the time alignment loss and reinforcement learning. The former promotes the improvement of model performance, while the latter can preserve data. In practical application scenarios, such as in social media monitoring, brand and marketing teams often need to monitor the emotional reactions of users on social media to quickly respond to market trends.

[0081] For example: first collect user comments, posts and pictures on social media platforms; then perform sentiment analysis using multimodal alignment methods by extracting text, audio (if there are video comments) and image features. Use visualization tools to generate sentiment trend charts to help brands identify user mood fluctuations and adjust market strategies in a timely manner. By real-time monitoring of sentiment trends, brands can better understand user needs and improve the targeting and effectiveness of marketing campaigns. After the release of new products, brands can quickly evaluate user feedback and make corresponding adjustments.

[0082] The present invention also provides a high-performance computer emotion analysis device based on multimodal data alignment, as Figure 5 shown. The device includes:

[0083] A data acquisition module, using ResNet-50 as the backbone network to extract multimodal time series features;

[0084] A feature alignment module, based on the trained DQN, constructs an action selection strategy to obtain the optimal DTW distance, and guides the alignment of multimodal time series state vectors.

[0085] Among them, DQN is composed of multiple fully connected layers. After each layer, the ReLU activation function is used as the activation layer, and the Q-value estimates of each step of actions (such as moving up, down, or diagonally) are output; DQN is used to estimate the cumulative rewards that may be obtained after taking different actions in the current state, and finally generate the optimal path and output the DTW distance.

[0086] The DQN-DTW network structure designed by the present invention combines the dynamic time warping (DTW) algorithm and the deep reinforcement learning technology, and proposes the multi-modal alignment block DQN-AlignBlock, which effectively improves the accuracy and efficiency of multi-modal data alignment, and at the same time reduces the computational complexity and time overhead.

[0087] Aiming at the consistency problem in the processing of different modal features by traditional methods in multi-modal alignment, the present invention realizes efficient alignment by introducing the DQN algorithm for dynamic adjustment; embedding it into the proposed DQN-DTW network realizes the dual benefits of the efficiency and flexibility of the network.

[0088] The present invention also provides an electronic device. Figure 6 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention, as Figure 6 shown. The electronic device may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory, for example, to execute the following method:

[0089] S1. Use ResNet-50 as the backbone network to extract multi-modal time series features;

[0090] S2. Input the extracted time series features into the state representation layer of the trained DQN to generate a time series state vector;

[0091] S3. Use DQN to construct an action selection strategy to obtain the optimal DTW distance, and guide the alignment of the multi-modal time series state vectors.

[0092] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0093] Embodiments of the present invention also provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments, for example, including:

[0094] S1. Using ResNet-50 as the backbone network to extract multi-modal time series features;

[0095] S2. Inputting the extracted time series features into the state representation layer of the trained DQN to generate a time series state vector;

[0096] S3. Using DQN to construct an action selection strategy to obtain the optimal DTW distance and guide the alignment of the multi-modal time series state vectors.

[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A high-performance sentiment analysis method based on multimodal data alignment, characterized in that: include: S1, using ResNet-50 as the backbone network to extract multimodal time series features; S2, input the extracted time series features into the state representation layer of the trained DQN to generate a time series state vector; S3. Use DQN to build an action selection strategy to obtain the optimal DTW distance and guide the alignment of multimodal time series state vectors.

2. The sentiment analysis method according to claim 1, characterized in that: The ResNet-50 is used as the backbone network to extract multimodal time series features, including: Multimodal features are extracted from the output of the Conv4c layer, and the context frame features are stacked along the time dimension to obtain combined features. The combined features are aggregated through the convolution layer to aggregate the time information and pooled through the pooling layer, and then the multimodal time series features are output through 2 fully connected layers and 1 linear projection layer.

3. The sentiment analysis method according to claim 1, characterized in that: The DQN structure includes: Multiple layers of fully connected layers, using the ReLU activation function as the activation layer after each fully connected layer, output the Q value estimate of each step of action, and obtain the cumulative return of the Q value estimate of each step of action.

4. The sentiment analysis method according to claim 1, characterized in that: In S3, DQN is used to construct an action selection strategy, including: Based on the Q-value estimate output by DQN, an ε-greedy strategy is used to select an upward, downward, or diagonal move action.

5. The sentiment analysis method according to claim 1, characterized in that: The S3 guides the multi-modal time series state vector alignment, including: If the current alignment state minimizes the local distance between the two time series, a positive reward is given, otherwise a negative reward is given. Based on the instantaneous reward value of each action, the accuracy of the alignment of the two time series is designed; and the Q value estimate of each state-action pair is updated through the Bellman equation.

6. The sentiment analysis method according to claim 1, characterized in that: When the DQN is trained, the joint loss function includes the DTW loss and the reinforcement learning reward function, and the formula is expressed as: L=L DTW +λL RL Among them, L DTW is the DTW loss, λL RL is the reinforcement learning reward function, and λ is used to balance the DTW loss and the reinforcement learning reward function.

7. The sentiment analysis method according to claim 6, characterized in that: Said: DTW loss, the formula is expressed as: Where N is the time step of mode 1, M is the time step of mode 2, and x i represents the eigenvector of mode 1 at the i-th time step, y j represents the eigenvector of mode 2 at the jth time step; Reinforcement learning reward function, the formula is expressed as: Where T is the number of time steps, γ is the discount factor, and R t represents the reward at the t-th time step. Specifically, the DTW loss learns the alignment path for each pair of modalities by minimizing the temporal distance between the modalities.

8. The sentiment analysis method according to claim 1, characterized in that: The time series state vector is generated in S2, and the formula is expressed as: in, is the global average pooling feature map, is the maximum pooling feature map, C1D k is a 1D convolution with a kernel size of k.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the sentiment analysis method according to any one of claims 1 to 8 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the sentiment analysis method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data alignment processing method and device for multi-modal fusion

    CN118761027A