Multimodal sentiment analysis method and device for negative information filtering and medium

Through comparative learning within and between modes and multimodal fusion of geometric algebra models, negative information is eliminated, and the problem of affecting the accuracy of sentiment analysis in the prior art is solved, achieving higher accuracy and robustness.

CN120067800APending Publication Date: 2025-05-30SHANGHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510136884.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing multimodal sentiment analysis methods ignore negative information present in the data source, resulting in the accuracy and robustness of sentiment analysis.

Method used

Through comparative learning within and between modes, multimodal fusion is performed in combination with geometric algebra models to eliminate negative information and achieve accurate sentiment analysis.

Benefits of technology

It improves the accuracy and robustness of multimodal sentiment analysis, and can more accurately understand and integrate emotional information from different modes, reducing misjudgments caused by negative information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067800A_ABST
    Figure CN120067800A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-mode sentiment analysis method and device for negative information filtering and a medium, and the method comprises the steps: collecting audio, video and text information, and carrying out the feature extraction of each mode; performing comparative learning on each extracted modal feature, selecting other modals from each modal to form a modal pair, calculating a comparative loss function of each modal pair, and obtaining feature values, including a first audio feature value, a first video feature value and a first text feature value, of which negative information in the modals and between the modals is removed; inputting the characteristic values into a geometric algebraic model to obtain fused multi-modal characteristic data; and performing regression operation on the fused multi-modal feature data, calculating a total loss function, further eliminating negative information between modals, and finally outputting an emotion analysis result. Compared with the prior art, the method has the advantages of being high in sentiment analysis accuracy, high in robustness, more sensitive and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and more particularly to a multi-modal sentiment analysis method, device, and medium for negative information filtering. Background Art

[0002] In the rapidly developing field of human-computer interaction, multi-modal sentiment analysis has become a key technology and has received much attention from the research community. Its core goal is to accurately identify human emotions, thereby bridging the communication gap between humans and machines. By integrating multiple input sources such as facial expressions, audio, text, and physiological signals, an understanding of the emotional state is constructed. Compared with single-modal sentiment analysis methods based on audio, text, and images, multi-vector multi-modal sentiment analysis exhibits higher robustness and accuracy due to its ability to fuse multiple signals.

[0003] Current algorithms mainly focus on synergistically integrating multi-modal data to improve recognition accuracy, such as attention-based mechanisms, tensor fusion, and transformer networks. However, they ignore the problem of negative information existing in the data sources. Negative information refers to data that may mislead emotion analysis, usually coming from noise or differences from the main emotion labels. When different modalities, such as text, speech, and images, express inconsistent emotions, data that misleads emotion analysis will be generated. For example, the text conveys positive emotions, while the speech or facial expression shows negative or neutral emotions. Secondly, background noise, blurred images, or unclear recordings will reduce the accuracy of sentiment analysis. If negative information is ignored, it will seriously affect the effectiveness of the sentiment analysis system. The current method to solve this problem is to emphasize the consistency and complementarity of multi-modal inputs, while ignoring the subtle dynamics when differences occur between the inputs. This gap in identifying and integrating negative information limits the model's ability to handle complex emotion expressions and reduces its flexibility in adapting to the diverse and mixed emotions commonly found in the real world. For example, feature extractors such as COVAREP and Openface cannot completely eliminate the inherent noise in single-modal features, causing certain interference to downstream sentiment inference tasks and complicating the analysis. Therefore, how to effectively identify and eliminate negative information within and between modalities to achieve accurate sentiment analysis is a technical problem that needs to be solved. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multi-modal sentiment analysis method, device, and medium for negative information filtering, which eliminates negative information and achieves accurate sentiment analysis through contrastive learning within and between modalities, as well as multi-modal fusion using geometric data.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] According to one aspect of the present invention, a multi-modal sentiment analysis method for negative information filtering is provided, and the specific steps include:

[0007] S1. Collect audio, video, and text information, and extract features for each modality respectively;

[0008] S2. Perform contrastive learning on the extracted features of each modality. For each modality, other modalities are respectively selected to form modality pairs, and the contrastive loss function of each modality pair is calculated to obtain eigenvalue of features removing negative information within and between modalities, including the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue;

[0009] S3. Input the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue into a geometric algebra model to obtain fused multi-modal feature data;

[0010] S4. Perform a regression operation on the fused multi-modal feature data, calculate the total loss function, further eliminate negative information between modalities, and finally output the sentiment analysis result.

[0011] Further, the negative information includes inconsistent emotion data expressed between different modalities and noise data of each modality signal.

[0012] Further, in S2, the contrastive learning of each modality includes:

[0013] Initialize the feature data parameters, generate a positive sample mask, align the dimensions through a linear layer and perform different vector normalization processes; calculate the similarity matrix of the intra-modal features of each modality; determine whether the sample is a positive sample. If so, remove the negative sample part in the positive sample through the similarity matrix. If not, combine the logical values of the negative samples to generate a mask of the same shape; combine the determined positive and negative samples; for each modality, other modalities are respectively selected to form modality pairs, that is, the combination of the current modality - other modalities, calculate the contrastive loss function of each modality pair, remove negative information within and between modalities, maximize the similarity of the same sentiment category, and minimize the similarity of different sentiment categories.

[0014] Further, the expression of the contrastive loss function L(x i ) of the modality pair is respectively:

[0015]

[0016] Among them, x i is the current modality, and the samples of the two modalities in the modality pair; δ(x i , y i ) is the similarity measure of the same samples of two different modalities; δ(x i , y k) is the similarity measure for different samples of two different modalities; δ(x i , x j ) is the similarity measure for different samples of the same modality; θ is the hyperparameter that controls the within-modal alignment.

[0017] Furthermore, for the said δ(x i , y i ), the specific expression is:

[0018] δ(x i , y i ) = exp(f x (x i )) T f y (y i ))),

[0019] where f x (x i ) is the representation of sample x, and f y (y i ) is the representation of sample y.

[0020] Furthermore, in the text modality contrastive learning model, the modality pairs include text-audio modality pairs and text-video modality pairs; in the audio modality contrastive learning model, the modality pairs include audio-text modality pairs and audio-video modality pairs; in the video modality contrastive learning model, the modality pairs include video-text modality pairs and video-audio modality pairs.

[0021] Furthermore, in the said S3, the specific steps for modality fusion through the geometric algebra model include:

[0022] Perform geometric convolution calculations on the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue to obtain the sum of feature vectors, and perform convolution operations in the geometric algebra space. The convolution operations include left convolution and right convolution, and process discrete and non-fully symmetric vector features; on each layer of features in the scale space of geometric algebra, detect feature points through the Hessian matrix; select a cube neighborhood centered on the feature points to obtain the modality eigenvalues of each feature point region; sum up the modality eigenvalues to obtain the multi-modal fusion eigenvalue.

[0023] Furthermore, the expression of the total loss function is:

[0024] L = L predictive + αL(x i ) + βL Algebra ,

[0025] where L predictive is the regression loss function, L(x i ) is the contrast loss function, LAlgebra is a geometric algebra loss function, and α and β are loss parameters;

[0026] The expression of the regression loss function is:

[0027]

[0028] where n is the number of samples; is the feature prediction value of sample y j ;

[0029] The expression of the geometric algebra loss function is:

[0030]

[0031] where F fusion is the fusion function of three modalities, n is the number of samples, c is the number of emotion categories, and y ij is sample i in emotion j.

[0032] According to the second aspect of the present invention, there is provided an electronic device, including a memory and a processor, where a computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.

[0033] According to the third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] (1) Improve the accuracy and robustness of multi-modal sentiment analysis: Through contrast learning, negative information is effectively removed from multi-modal data such as audio, video, and text. Contrast learning is applied in the feature spaces within and between each modality to eliminate negative information within and between modalities, thereby improving the accuracy of multi-modal sentiment analysis, more precisely understanding and integrating sentiment information from different modalities, and reducing misjudgments caused by negative information.

[0036] (2) Achieve a unified multi-dimensional vector space representation and further eliminate negative information between modalities: Through the mathematical framework of geometric algebra, feature representations of different modalities can be transformed into unified geometric objects, expressed in a unified way, and their relationships are expressed in a higher-level mathematical structure, solving the problem of inconsistent sentiment between different modalities, further eliminating negative information between modalities, and thus further enhancing the stability, accuracy, and sensitivity of sentiment analysis.

[0037] (3) Optimize feature representation and improve sentiment analysis performance: By combining contrastive learning techniques and geometric algebra frameworks, effective elimination of negative information and optimization of feature representation in multimodal sentiment analysis are achieved. Geometric algebra provides a unified feature representation space for contrastive learning, enabling contrastive learning to more accurately identify and eliminate negative samples and more subtly capture and analyze emotional features. At the same time, the optimization mechanism of geometric algebra further enhances the stability and consistency of feature representation. This combined technology not only improves the accuracy of sentiment analysis but also enables the model to demonstrate balanced and efficient performance on multiple key metrics. Brief Description of the Drawings

[0038] Figure 1 Flowchart of the multimodal sentiment analysis method for negative information filtering;

[0039] Figure 2 Flowchart of contrastive learning;

[0040] Figure 3 Flowchart of the geometric algebra module. Detailed Implementation Manner

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] As Figure 1 shown, it is a multimodal sentiment analysis method for negative information filtering, and the specific steps include:

[0043] S1. Collect audio, video, and text information, and perform feature extraction on each modality separately;

[0044] S2. Perform contrastive learning on the extracted features of each modality separately. For each modality, other modalities are respectively selected to form modality pairs, and the contrastive loss function of each modality pair is calculated to obtain eigenvalue of features after removing negative information within and between modalities, including the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue;

[0045] S3. Input the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue into the geometric algebra model to obtain fused multimodal feature data;

[0046] S4. Perform a regression operation on the fused multimodal feature data, calculate the total loss function, further eliminate negative information between modalities, and finally output the sentiment analysis result.

[0047] Among them, the negative information includes inconsistent emotion data expressed between different modalities and noise data of each modality signal.

[0048] In S1, a pre-trained text model is used to extract features from the information of the text modality. Through deep bidirectional encoding representation, semantic and syntactic information in the text is captured. For the information of the video modality, a facial action analysis tool is used to identify and measure facial action units defined in the Facial Action Coding System, and facial postures and related expression features are captured to support emotion analysis. For the information of the audio modality, an audio feature processing unit is used to perform comprehensive sound processing functions on the audio modality. By extracting various low-level acoustic features, including 12 Mel-frequency cepstral coefficients, pitch, segmentation features between voiced and unvoiced segments, and glottal source parameters. In addition, in order to better capture the long-term dependencies in the speech and facial expression time series data, a long short-term memory network is used to manage, store, forget, and update information through a gating mechanism, and effectively extract dynamic features from audio and visual signals. Finally, the visual feature dimension is 708, the text feature dimension is 768, and the acoustic feature dimension is 33.

[0049] As Figure 2 shown, in S2, the contrastive learning of each modality includes: initializing the feature data parameters, generating a positive sample mask, aligning the dimensions through a linear layer and performing different vector normalization processes; calculating the similarity matrix of the intra-modal features of each modality; determining whether the sample is a positive sample. If it is, the negative sample part in the positive sample is removed through the similarity matrix. If not, the logical values of the negative samples are combined to generate a mask of the same shape; combining the positive and negative samples obtained from the determination; selecting other modalities for each modality to form modality pairs, that is, the combination of the current modality - other modalities, calculating the contrastive loss function of each modality pair, removing negative information within and between modalities, maximizing the similarity of the same emotion category, and minimizing the similarity of different emotion categories.

[0050] In the text modality contrastive learning model, the modality pairs include text-audio modality pairs and text-video modality pairs; in the audio modality contrastive learning model, the modality pairs include audio-text modality pairs and audio-video modality pairs; in the video modality contrastive learning model, the modality pairs include video-text modality pairs and video-audio modality pairs.

[0051] At the same time, a small batch M of samples is drawn for contrastive learning, and the positive sample set P i and the negative sample set N i are defined as:

[0052] P i ={y i},

[0053]

[0054] Among them, y i is each sample, and the contrast loss functions L(x i ) of the modality pairs are respectively expressed as:

[0055]

[0056] Among them, x i is the current modality, the samples of the two modalities in the modality pair; δ(x i , y i ) is the similarity measure of the same samples of two different modalities; δ(x i , y k ) is the similarity measure of different samples of two different modalities; δ(x i , x j ) is the similarity measure of different samples of two same modalities; θ is the hyperparameter that controls the within-modality alignment.

[0057] The specific expression of δ(x i , y i ) is:

[0058] δ(x i , y i ) = exp(f x (x i )) T f y (y i ))

[0059] Among them, f x (x i ) is the representation of sample x, and f y (y i ) is the representation of sample y.

[0060] In this embodiment, through the contrast loss function, contrast learning is performed both within the modality and between modalities. By maximizing the similarity between modality pairs from the same emotion category and minimizing the similarity between modality pairs from different emotion categories, the negative information between modalities and within modalities is eliminated, enabling the model to more accurately understand and integrate information from different modalities when dealing with complex emotion expressions, and improving its ability to recognize complex emotions in real-world scenarios.

[0061] Geometric algebra effectively eliminates the negative information impact between vectors by establishing a unified expression method among vector objects in different dimensions. In sentiment analysis, different modalities usually correspond to different feature spaces, and there are also differences in the emotional expression forms of each modality. Such differences and inconsistencies may lead to negative information between modalities. For example, the emotional expression of one modality may conflict with that of other modalities, affecting the accuracy of overall sentiment analysis. Geometric algebra can handle the conflicts between vectors through symmetry and geometric transformation. For example, the inner product can be used to measure the similarity between vectors, and the outer product can capture the interaction between different emotions. By performing geometric algebra operations on the vectors of each modality, the emotional expression of each modality can be balanced and adjusted, eliminating the interference caused by negative information between modalities, such as emotional inconsistency or noise. The representation method based on geometric algebra can enhance the robustness of the multi-modal sentiment analysis model, enabling it to maintain a high recognition accuracy when facing complex and mixed emotional expressions.

[0062] As Figure 3 shown, in S3, the specific steps for modality fusion through the geometric algebra model are as follows:

[0063] Perform geometric convolution calculations on the first audio eigenvalue, the first video eigenvalue, and the first text eigenvalue to obtain the sum of feature vectors. Conduct convolution operations in the geometric algebra space. The convolution operations include left convolution and right convolution to process discrete and non-fully symmetric vector features. Detect feature points through the Hessian matrix on each layer of features in the scale space of geometric algebra. Select a cubic domain centered on the feature points to obtain the eigenvalues of each modality in each feature point region. Sum up the eigenvalues of each modality to obtain the eigenvalue of multi-modal fusion.

[0064] In the feature vectors of the three modalities of audio, video, and text, the convolution expression is:

[0065] v 3 (r) = ∫∫∫Q N (μ)V(r - μ)dμ,

[0066] where μ and r are the respective eigenvalues, and Q N (μ) and V(r - μ) are two three-dimensional spatial vectors in space, and the product of Q N (μ)V(r - μ) is a geometric product. In discrete vectors, the expression for the convolution of two 3D vectors is:

[0067] v 3 (r 1 r 2 r 3 ) = ∑∑∑Q(μ 1 μ 2 μ 3 )V(r1 -μ 1 ,r 2 -μ 2 ,r 3 -μ 3 ),

[0068] where Q(μ 1 μ 2 μ 3 ) and V(r 1 -μ 1 ,r 2 -μ 2 ,r 3 -μ 3 ) are two three - dimensional multivectors in a geometric algebra space, and the product of Q(μ 1 μ 2 μ 3 ) and V(r 1 -μ 1 ,r 2 -μ 2 ,r 3 -μ 3 ) is a geometric product. Since the characteristics of the three vectors extracted are discrete, and the geometric product is neither completely symmetric nor completely antisymmetric, there are two convolutions, the left convolution F left and the right convolution F right respectively as follows:

[0069]

[0070] where h(j 1 ,j 2 ,j 3 ) is an equation in a geometric algebra space, and h(j 1 ,j 2 ,j 3 )I(f 1 -j 1 ,f 2 -j 2 ,f 3 -j 3 ) is a geometric product. The unification between different vectors in a high - dimensional space is achieved through the above convolution method.

[0071] In this embodiment, for multimodal emotion analysis, the extraction of geometric features is realized through the SURF algorithm. For the data stream of multiple vectors, its geometric algebra structure is equivalent to a multi - valued signal, which is described by a linear vector as:

[0072] I(x) = f 1 (x)e 1 + f 2 (x)e 2 + f3 (x)e 3 ,

[0073] On the other hand, multi-modal emotion analysis is reconstructed from many emotion vectors of low-dimensional projections, similar to the reconstruction process of common multi-modal data. The representation of the scale space of geometric algebra first performs the convolution operation mentioned above. Then, on each layer of features in the scale space, extreme points are detected according to the Hessian matrix. The specific formula is as follows:

[0074]

[0075] Among them, is the second-order partial derivative of the convolution of X point and I. After judging the feature points, first judge whether the determinant value of the above formula of the point satisfies the condition of being greater than the set threshold. For the points that meet the conditions, in the 4D scale space (X, S) = f 1 (x)e 1 +f 2 (x)e 2 +f 3 (x)e 3 +Se 4 . Only when the response value of this point is greater than the response value of its neighboring points, this point is regarded as a feature point. Through this method, geometric features can not only describe the spatial characteristics of multi-modal emotion analysis, but also be closely combined with its projection structure, thus enhancing the interpretability and analysis accuracy of the data. After detecting the feature points, a cubic neighborhood centered on the feature points is selected. Each feature point region can correspond to three-dimensional eigenvalue respectively as: d 1 e 1 , d 2 e 2 , d 3 e 3 . Then sum the eigenvalues of this region, which is the vector combination of the last three modalities. The specific formula is as follows:

[0076] F fusion = ∑d 1 e 1 + ∑d 2 e 2 + ∑d 3 e 3 + ∑|d 1 e 1 | + ∑|d 2 e 2 | + ∑|d 3 e 3 ,

[0077] The present invention introduces a unified multi-dimensional vector space through geometric algebra, which can transform feature representations of different modalities into unified geometric objects and express the relationships between them with a higher-level mathematical structure. Under this framework, each vector (text, audio, image) in sentiment analysis can be described and fused through operations such as matrices and convolutions in geometric algebra. Through these operations, vectors of different modalities can be compared and combined in a unified way, thus solving the problem of emotional inconsistency between different modalities. In summary, the geometric algebra framework introduced by the present invention provides an effective framework for multi-vector sentiment analysis, which can eliminate negative information between modalities through unified geometric representations, thereby improving the accuracy and stability of sentiment analysis.

[0078] The expression of the total loss function is:

[0079] L = L predictive + αL(x i ) + βL Algebra ,

[0080] where L predictive is the regression loss function, L(x i ) is the contrast loss function, L Algebra is the geometric algebra loss function, and α and β are loss parameters;

[0081] The expression of the regression loss function is:

[0082]

[0083] where n is the number of samples; is the feature prediction value of sample y j ;

[0084] The expression of the geometric algebra loss function is:

[0085]

[0086] where F fusion is the fusion function of three modalities, n is the number of samples, c is the number of emotion categories, and y ij is sample i in emotion j.

[0087] In this embodiment, comparative experiments were conducted on the aligned datasets CMU-MOSI and CMU-MOSEI, and the experimental results are shown in Tables 1, 2, and 3. On the CMU-MOSI dataset, the evaluation index results for binary classification are 84.1 / 86.1, the mean squared error is 0.699, and the correlation is 0.799. On the CMU-MOSEI dataset, the evaluation index results for binary classification are 83.7 / 86.0, the mean squared error is 0.538, and the correlation is 0.768. These indicators show that this embodiment performs excellently in accurately identifying different emotional intensities and distinguishing positive and negative emotions, and the method of this embodiment demonstrates balanced and efficient performance on multiple key indicators. The accuracy of sentiment analysis is improved by eliminating negative information from speech, text, and image data. Contrastive learning is used to filter out negative information in the data, and by introducing a geometric algebra framework, the filtered features are integrated into a unified representation for sentiment analysis. This framework enables the model to dynamically focus on cross-modal relevant features, ensuring the sensitivity of the fusion process to the nuances of emotional expression.

[0088] Table 1 Experimental data on CMU-MOSI and CMU-MOSEI datasets

[0089] dataset accuracy F1 mean squared error correlation CMU-MOSI 84.1 / 86.1 84.1 / 86.2 0.699 0.799 CMU-MOSEI 83.7 / 86.0 83.7 / 86.1 0.538 0.768

[0090] Table 2 Experimental data on CMU-MOSI dataset for different modules

[0091]

[0092] Table 3 Experimental data on CMU-MOSEI dataset for different modules

[0093]

[0094] In the performance comparison experiment on the aligned datasets, this embodiment first evaluated the proposed model based on the CMU-MOSI and CMU-MOSEI aligned datasets. As shown in Table 1, this embodiment shows significant advantages in multiple indicators. Specifically,

[0095] In the qualitative analysis, the binary classification accuracy (Acc2) of this embodiment is particularly prominent, reaching 84.1 / 86.1 on CMU-MOSI and 83.7 / 86.0 on CMU-MOSEI. In the quantitative analysis, the mean absolute error (MAE) of the model proposed in this embodiment is 0.699 on CMU-MOSI and 0.538 on CMU-MOSEI, both of which are the lowest values among similar models, indicating that this embodiment has high-precision characteristics in the emotional intensity estimation task, which is crucial for capturing subtle emotional differences.

[0096] In addition, the model achieved high scores of 0.799 and 0.768 respectively in terms of the correlation coefficient (Corr) index, further verifying its ability to effectively capture the internal laws of emotional expression and being highly consistent with human subjective judgments.

[0097] Tables 2 and 3 show the result data of the module ablation experiments conducted in this embodiment on the CMU - MOSI and CMU - MOSEI datasets, aiming to verify the impact of the contrast learning and geometric algebra filtering fusion module on the performance of the emotion recognition model. The experimental results show that the synergistic effect of contrast learning and geometric algebra mechanism significantly improves the model's classification ability for complex emotional states. Contrast learning reduces the features of inconsistent emotional expressions by suppressing the influence of negative samples with modality and label conflicts; geometric algebra eliminates redundant noise by optimizing the feature space alignment. After combining contrast learning with geometric algebra filtering fusion, the accuracy of the model on CMU - MOSI is 84.1 / 86.1, the correlation is 0.799, both reaching the highest values; the average mean square error is 0.699, reaching the lowest value, indicating that the solution of this embodiment has the best joint elimination effect on negative information.

[0098] Generally speaking, the present invention not only performs excellently in multi - classification tasks, but also can accurately quantify the emotional intensity, and at the same time maintains a strong correlation with the real emotional distribution, reflecting its comprehensive advantages in complex emotion recognition scenarios.

[0099] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0100] The electronic device of the present invention includes a central processing unit (CPU), which can execute various appropriate actions and processes according to the computer program instructions stored in the read - only memory (ROM) or the computer program instructions loaded from the storage unit into the random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.

[0101] Multiple components in the device are connected to the I / O interface, including: an input unit, such as a keyboard, a mouse, etc.; an output unit, such as various types of displays, speakers, etc.; a storage unit, such as a disk, an optical disc, etc.; and a communication unit, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks. The processing unit executes the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of the method of the present invention described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute the method of the present invention by any other suitable means (e.g., by means of firmware).

[0102] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0103] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0104] In the context of the present invention, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0105] As described above, the foregoing are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A multimodal sentiment analysis method for filtering negative information, characterized in that: The specific steps include: S1, collect audio, video and text information, and extract features for each modality; S2, performing comparative learning on the extracted features of each modality respectively, each modality selects other modalities to form a modality pair, and calculates the comparative loss function of each modality pair to obtain feature values ​​for removing negative information within the modality and between modalities, including a first audio feature value, a first video feature value, and a first text feature value; S3, inputting the first audio feature value, the first video feature value, and the first text feature value into a geometric algebra model to obtain fused multimodal feature data; S4. Perform regression operation on the fused multimodal feature data, calculate the total loss function, further eliminate negative information between modalities, and finally output the sentiment analysis result.

2. The multimodal sentiment analysis method for filtering negative information according to claim 1, characterized in that: The negative information includes inconsistent emotion data expressed between different modalities and noise data of each modal signal.

3. The multimodal sentiment analysis method for filtering negative information according to claim 1, characterized in that: In S2, the comparative learning of each modality includes: Initialize the feature data parameters, generate a positive sample mask, align the dimensions through the linear layer and perform different vector normalization processing; calculate the similarity matrix of the intra-modal features of each modality; determine whether the sample is a positive sample. If so, remove the negative sample part in the positive sample through the similarity matrix. If not, combine the logical values ​​of the negative samples to generate a mask of the same shape; combine the positive and negative samples obtained by the judgment; select other modalities for each modality to form a modal pair, that is, a combination of the current modality and other modalities, calculate the contrast loss function of each modality pair, remove negative information within and between modalities, maximize the similarity of the same emotion category, and minimize the similarity of different emotion categories.

4. The multimodal sentiment analysis method for filtering negative information according to claim 3, characterized in that: The contrast loss function L(x i ) is: Among them, x i is the current mode, the samples of the two modes in the mode pair; δ(x i ,y i ) is the similarity measure of the same samples in two different modalities; δ(x i ,y k ) is the similarity measure of two different modal samples; δ(x i ,x j ) is the similarity measure between two samples of the same modality; θ is the hyperparameter that controls the intra-modality alignment.

5. The multimodal sentiment analysis method for filtering negative information according to claim 4, characterized in that: The δ(x i ,y i ) is expressed as: δ(x i ,y i )=exp(f x (x i ) T f y (y i )), Among them, f x (x i ) is the representation of sample x, f y (y i ) is the representation of sample y.

6. The multimodal sentiment analysis method for filtering negative information according to claim 4, characterized in that: In the text modality contrast learning model, the modality pairs include text-audio modality pairs and text-video modality pairs; in the audio modality contrast learning model, the modality pairs include audio-text modality pairs and audio-video modality pairs; in the video modality contrast learning model, the modality pairs include video-text modality pairs and video-audio modality pairs.

7. The multimodal sentiment analysis method for filtering negative information according to claim 1, characterized in that: In S3, the specific steps of performing modal fusion through the geometric algebraic model include: The first audio eigenvalue, the first video eigenvalue and the first text eigenvalue are subjected to geometric convolution calculation to obtain the sum of the eigenvectors, and a convolution operation is performed in the geometric algebraic space. The convolution operation includes left convolution and right convolution to process discrete and non-completely symmetric vector features. The feature points are detected by the Hessian matrix on each layer of the feature in the geometric algebraic scale space. The cubic area centered on the feature point is selected to obtain the eigenvalues ​​of each modality of each feature point area. The eigenvalues ​​of each modality are summed to obtain the eigenvalues ​​of multimodal fusion.

8. The multimodal sentiment analysis method for filtering negative information according to claim 1, characterized in that: The expression of the total loss function is: L=L predictive +αL(x i )+βL Algebra , Among them, L predictive is the regression loss function, L(x i ) is the contrast loss function, L Algebra is the geometric algebraic loss function, α and β are loss parameters; The expression of the regression loss function is: Where n is the number of samples; For sample y j The feature prediction value of The expression of the geometric algebraic loss function is: Among them, F fusion is the fusion function of the three modes, n is the number of samples, c is the number of emotion categories, y ij is sample i in emotion j.

9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.