Video Emotion Analysis Method Based on Multi-Scale Feature Extraction and Multi-Task Learning

By adopting multi-scale feature extraction and multi-task learning methods in video sentiment analysis, the problem of insufficient modal feature extraction and information complementarity capabilities in the prior art is solved, and more efficient multi-modal video sentiment analysis is achieved.

CN116597353BActive Publication Date: 2025-05-30FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310558675.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-05-30
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

The existing multimodal video sentiment analysis methods are difficult to effectively extract the characteristics of different modes and model the connections between modes, resulting in the inadequacy of information complementarity.

Method used

The video sentiment analysis method based on multi-scale feature extraction and multi-task learning is adopted. Through the combination of single-modal feature extraction, multi-scale feature representation, cross-modal cross-attention mechanism and multi-task learning, the features of different modes are extracted and the connection between modes is modeled.

Benefits of technology

It improves the accuracy of multimodal video sentiment analysis, can better extract the characteristics of different modes and model the connections between modes, and enhances the ability to complement each other.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116597353B_ABST
    Figure CN116597353B_ABST
Patent Text Reader

Abstract

The present invention proposes a video emotion analysis method based on multi-scale feature extraction and multi-task learning. First, in order to extract multi-scale features, the present invention proposes a multi-scale feature extraction method, which uses channel attention to model the outputs of different hidden layers. Secondly, a multi-modal fusion strategy based on key modalities is proposed, which uses the attention mechanism to increase the proportion of key modalities and explores the relationship between key modalities and other modalities. Finally, the proposed model is trained using the multi-task learning method to ensure that the model can learn better feature representations. Through experiments, the present invention achieves better emotion analysis accuracy on the publicly available multi-modal emotion analysis benchmark dataset than existing algorithms and inventions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-modal video emotion analysis in the field of emotion analysis, and particularly to a video emotion analysis method based on multi-scale feature extraction and multi-task learning. Background Art

[0002] Currently, deep learning has been widely applied in fields such as computer vision, natural language processing, and text signal processing. In addition, deep learning and neural networks have also shown levels close to or even exceeding that of humans in fields such as emotion analysis, image classification, and text translation. To better understand the present invention, the following introduces several basic concepts in the field of deep learning.

[0003] Deep learning: Deep learning is a new research branch in the field of machine learning. By learning the internal distribution law of data samples and the hierarchical representation of data, deep learning enables a computer to have powerful analysis capabilities for data such as text, pictures, sounds, and videos. Its purpose is to hope that the computer can have certain intelligent behaviors like humans.

[0004] Video emotion analysis: Emotion analysis is an important research direction in video understanding tasks. By analyzing videos, the emotional states of the people in the pictures are discovered, and it has important application values in fields such as human-computer interaction, medical diagnosis, and investigation and interrogation. With the rapid development of deep learning technology, people are no longer satisfied with only using one modality to analyze the emotion of videos, but use the data of multiple modalities for fusion to analyze the emotion of videos. Compared with single-modal data, multi-modal data contains richer information and can better reveal the true emotions of users. Analyzing the emotions of these massive multi-modal data helps to better understand people's attitudes and viewpoints and has a wide range of application scenarios.

[0005] Multi-scale feature extraction: Multi-scale feature extraction refers to different feature extraction methods or different scales of the extracted features at different stages of feature extraction.

[0006] Multi-task learning: Multi-task learning refers to improving the generalization performance of all tasks by utilizing the relevant information between multiple tasks during the learning process of a neural network.

[0007] Since multi-source heterogeneous data is distributed in different modal representation spaces, the barriers between different spaces make it difficult to tightly couple the feature vectors of each modality, and less information interaction inside and outside the modality is considered during the modality fusion process, which may ignore the information complementary ability between different modalities. Therefore, there is a need for a more efficient multi-modal based video emotion analysis method that can better extract the features of different modalities and model the connections between modalities. Summary of the Invention

[0008] In view of this, the purpose of the present invention is to provide a video emotion analysis method based on multi-scale feature extraction and multi-task learning, which can better extract features of different modalities and model the relationship between modalities.

[0009] To achieve the above object, the present invention adopts the following technical solutions: A video emotion analysis method based on multi-scale feature extraction and multi-task learning, the specific steps are as follows:

[0010] Step S1: Extract single-modal feature representations; on the basis of the original video clip data, use a single-modal feature extraction sub-model to extract single-modal feature representations from video and audio modality data using a feed-forward neural network, and use a BERT model to extract feature vectors for the text modality;

[0011] Step S2: Obtain single-modal multi-scale feature representations; respectively use channel attention mechanisms for the video and audio feature representations in Step S1 to obtain multi-scale feature representations, and then use 1×1 convolution operations to perform channel integration on them;

[0012] Step S3: Obtain the multi-modal fusion result: First, use cross-modal cross-attention to capture the interaction between text-video and text-audio; Second, for the text modality, use self-attention mechanism to capture its own internal information; Finally, splice the obtained feature representations to obtain multi-modal fusion features;

[0013] Step S4: Obtain the video emotion analysis result: Input the fusion features obtained in Step S3 into a multi-modal feature emotion analysis sub-model composed of a fully connected layer and an activation function to obtain the video emotion analysis result; Input the obtained single-modal features into a single-modal emotion analysis sub-model to obtain the single-modal emotion analysis result;

[0014] Step S5: Define the loss function; train the proposed model through multi-task learning, and the loss function of the model consists of four parts, namely: 1) Text emotion analysis loss function; 2) Audio emotion analysis loss function; 3) Video emotion loss function; 4) Multi-modal fusion emotion analysis loss function; The four loss functions are combined into the final loss function by weighting.

[0015] In a preferred embodiment, the sub-model structure for extracting single-modal feature vectors in Step S1 is as follows:

[0016] Step S11: First, for the feature extraction work of the text modality, use a pre-trained BERT model to extract text feature representations;

[0017] Step S12: For the feature extraction of video and audio modalities, a feature extraction sub-model consisting of five network layers is used to extract features for each modality; for the first network layer, a unidirectional long short-term memory network (LSTM) is used to capture the temporal features of the modality data; for the subsequent network layers, four linear layers are used for feature dimension transformation and regularization to ensure data stability; the output of the first network layer is shown in formula (1):

[0018]

[0019] Meanwhile, the ReLu function is used as the activation function, and the output of the second network layer is shown in formulas (2), (3), and (4):

[0020]

[0021]

[0022]

[0023] In a preferred embodiment, the obtaining of the single-modal feature representation in step S2 includes the following steps:

[0024] Step S21: First, after the video and audio modalities are processed by the feature extraction sub-model in step S1, the output of each network layer is obtained; subsequently, a new channel dimension is added to the output features of each layer; finally, stacking is performed by channel to obtain a stacked feature; as shown in formula (5):

[0025]

[0026] Step S22: The stacked feature is processed using the channel attention mechanism. First, the Squeeze operation is used to compress the data of each channel and summarize the information to obtain a global receptive field. The calculation process is shown in formula (6):

[0027]

[0028] Secondly, the Excitation operation is used to generate weights for each dimension of the stacked vector; the calculation formula is shown in (7):

[0029]

[0030] Finally, the Scale operation is used to multiply the stacked vector by the weights and a 1×1 convolutional kernel is used to perform channel integration on the result, thereby obtaining the final video and speech single-modal multi-scale features, as shown in formulas (8) and (9):

[0031]

[0032]

[0033] Step S23: For the text modality, use the initial 768-dimensional text features after being processed by BERT, and then use a linear layer to reduce its dimension to be equal to the dimension sizes of the video and speech modalities.

[0034] In a preferred embodiment, obtaining the video sentiment analysis result in the said step S3 includes the following steps:

[0035] Step S31: Set the three modalities as text-video modality pairs, text-audio modality pairs, and text single modality. For the first two, use cross-modal cross-attention. The specific operation is that when calculating the attention, set the text features as Key, and the features of the other modality as Value and Query, and at the same time use multi-head attention to calculate the fusion results of text-video and text-audio;

[0036] Step S32: For the text single modality, use the self-attention mechanism to capture the information useful for the sentiment analysis result in itself;

[0037] Step S33: Concatenate the results obtained in the above steps to get the final multi-modal fusion features.

[0038] In a preferred embodiment, obtaining the video sentiment analysis result in the said step S4 includes the following steps:

[0039] Input the final multi-modal fusion features into a multi-modal sentiment analysis sub-model composed of three fully connected layers and a tanh activation function, and further obtain the final video sentiment analysis result.

[0040] In a preferred embodiment, defining the loss function in the said step S5 includes the following steps:

[0041] Step S51: Use the single-modal sentiment analysis sub-models corresponding to their respective modalities to process the single-modal features obtained in step S2 to predict the single-modal sentiment analysis results; calculate the loss values between the respective single-modal sentiment analysis prediction values and the true single-modal sentiment labels;

[0042] Step S52: Use the multi-modal sentiment analysis sub-model to process the multi-modal fusion features obtained in step S4 to predict the video sentiment analysis result; calculate the loss value between the predicted video sentiment analysis result and the true video sentiment label;

[0043] Step S53: The above single-modal prediction results and fusion prediction results simultaneously affect the loss function of the entire model to obtain the final loss function, and optimize this loss function to achieve the model optimization effect.

[0044] Compared with the prior art, the present invention has the following beneficial effects: Through experiments, the sentiment analysis accuracy rate achieved by the present invention on the publicly available multi-modal sentiment analysis benchmark dataset is better than that of existing algorithms and inventions. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of a preferred embodiment of the present invention.

[0046] Figure 2 is a schematic diagram of obtaining modal multi-scale features in a preferred embodiment of the present invention.

[0047] Figure 3 is a schematic diagram of channel attention in a preferred embodiment of the present invention.

[0048] Figure 4 is a schematic diagram of cross-modal cross-attention in a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0050] It should be noted that the following detailed description is illustrative and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0051] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0052] A video sentiment analysis method based on multi-scale feature extraction and multi-task learning includes the following steps.

[0053] Step S1: Data preparation. Record a video file containing user emotions; obtain the text modality (t), audio modality (a), and visual modality (v).

[0054] Step S2: Data separation. Separate the three required modalities from the video file, namely the text modality, audio modality, and visual modality, where the visual modality refers to the video picture.

[0055] Step S3: Data segmentation. Segment the visual modality, audio modality, and text modality obtained in step S2 according to different window sizes and time intervals.

[0056] Step S4: Data processing. Perform data processing on each modal data segment obtained by segmentation in step S3. For the visual modality, use the OpenFace tool for processing. For the audio modality, use the COVAREP tool for processing. For the text modality, use the pre-trained Chinese BERT model for processing;

[0057] Step S5: Obtain multi-scale features. Extract multi-scale features from the visual modality data and audio modality data processed in step S4 to obtain visual modality multi-scale features and audio modality multi-scale features;

[0058] Step S6: Obtain text modality features. Process the text word vectors extracted in step S4 to obtain the feature representation of the text modality;

[0059] Step S7: Multimodal data fusion. Fuse the visual modality multi-scale features, audio modality multi-scale features, and text modality features obtained in step S5 and step S6 to obtain multimodal fusion features;

[0060] Step S8: Obtain the modal sentiment analysis results. Input the single-modal multi-scale feature representations obtained in step S5 and S6 into their respective single-modal sentiment classifiers to obtain the multimodal sentiment analysis results. Input the multimodal fusion features obtained in step S7 into the multimodal sentiment classifier to obtain the multimodal sentiment analysis results;

[0061] Step S9: Obtain the video sentiment analysis result. Perform weighted combination on the three different single-modal sentiment analysis results and the multimodal sentiment analysis result obtained in step S8 to obtain the video sentiment analysis result.

[0062] In this embodiment, step S1 specifically includes the following steps: Use devices with video recording and sound collection functions such as mobile phones, computers, and cameras to record a video containing a person's appearance. The video content has the person's emotion, such as the user's evaluation of a movie, etc. The selected video format is.mp4.

[0063] In this embodiment, step S2 specifically includes the following steps: For each video file, use the OpenCV tool to complete the separation of the audio, and separate the.mp3 audio file from the.mp4 video file. Use the audio file transcribing api provided by Baidu Smart Cloud to transcribe the text from the.mp3 file.

[0064] In this embodiment, step S3 specifically includes the following steps: cutting the visual modality, audio modality, and text modality according to different window sizes respectively. For the visual modality, the present invention extracts frame images from the video at a rate of 1 FPS; for the audio modality, the present invention divides the audio data in the video with a window size of 20 ms; for the text modality, the text is divided sentence by sentence.

[0065] In this embodiment, step S4 specifically includes the following steps: the present invention uses the OpenFace tool to extract 16 facial action units, 68 facial landmarks, head pose and direction, and 6 types of eye gaze features for each frame. For the audio modality, the COVAREP tool is used to extract 32-dimensional audio features for each audio segment, including 20-dimensional Mel-frequency cepstral coefficients (MFCCs) and 12-dimensional Constant-Q chromagram (CQT). For the text modality, for each sentence, pre-trained Chinese BERT is used to obtain 768-dimensional word vectors.

[0066] In this embodiment, taking the audio modality as an example, step S5 specifically includes the following steps:

[0067] Step S51: Build an audio modality feature extractor composed of n neural network layers;

[0068] Step S52: Input the 32-dimensional audio features obtained in step S4 into the feature extractor built in step S51;

[0069] Step S53: First, extract the output of each neural network layer in step S52. According to step S51, n data will be obtained here; secondly, add a channel dimension to each data; finally, stack the n data according to the channel dimension to obtain a stacked data composed of n data;

[0070] Step S54: Input the stacked data obtained in step S53 into the channel attention module to model the importance degree of each layer of data, and then use a convolutional kernel to integrate all channels to obtain the multi-scale features of the audio modality.

[0071] Step S55: For the visual modality, use the same processing method as the audio modality to obtain the multi-scale features of the visual modality.

[0072] Further, in step S6, the specific steps for obtaining the text modality features include:

[0073] Step S61: Build a text modality feature extractor composed of m neural network layers;

[0074] Step S62: Input the 768-dimensional text word vectors obtained in Step S4 into the feature extractor built in Step S61.

[0075] Step S63: After being processed in Step S62, a feature representation of the text modality is obtained. Different from the audio modality and the visual modality, the text modality does not extract multi-scale feature representations.

[0076] Furthermore, in the said Step S7, the multi-modal data fusion specifically includes the following steps:

[0077] Step S71: Select one modality from the visual modality, the audio modality, and the text modality as the key modality. The following description takes the text modality as the key modality;

[0078] Step S72: Use cross-modal attention to fuse the key modality (text modality) and the visual modality to obtain a text-visual fusion feature;

[0079] Step S73: Use cross-modal attention to fuse the key modality (text modality) and the audio modality to obtain a text-audio fusion feature;

[0080] Step S74: Use self-attention to emphasize the useful information within the key modality (text modality) itself to obtain a text-text fusion feature;

[0081] Step S75: Concatenate the three features obtained in Steps S72, S73, and S74 to obtain a multi-modal fusion feature.

[0082] Furthermore, in the said Step S8, obtaining the single-modal sentiment analysis result and the multi-modal sentiment analysis specifically includes the following steps:

[0083] Step S81: Input the multi-scale features of the visual modality into the visual modality sentiment classifier to obtain a visual sentiment analysis result;

[0084] Step S82: Input the multi-scale features of the audio modality into the audio modality sentiment classifier to obtain an audio sentiment analysis result;

[0085] Step S83: Input the text modality features into the text modality sentiment classifier to obtain a text sentiment analysis result;

[0086] Step S84: Input the multi-modal fusion feature into the multi-modal modality sentiment classifier to obtain a multi-modal sentiment analysis result;

[0087] Furthermore, in the said Step S9, the steps for obtaining the video sentiment analysis result include:

[0088] Step S91: Weightedly combine the four sentiment analysis results obtained in steps S81, S82, S83, and S84 to obtain the final video sentiment analysis result;

[0089] Step S92: The method loss function consists of the four sub-loss functions described in step S91. During the training process, by continuously iterating and optimizing, minimize the loss value of the method to improve the accuracy of sentiment analysis.

[0090] For describing the specific implementation manners of the present invention, the following definitions are made:

[0091] Text modality (t), audio modality (a), and visual modality (v), denoted as Ui, i ∈ {t, v, a}.

[0092] The detailed steps are as follows:

[0093] Step 1: After the processing of steps S1, S2, S3, and S4, obtain the text modality (t), audio modality (a), and visual modality (v);

[0094] Step 2: As described in step S5, the process of obtaining multi-scale features is as Figure 2 shown. In this embodiment, taking the audio modality a as an example, according to the above method, the process of obtaining multi-scale features of the audio modality is given, which specifically includes the following steps:

[0095] Step 21: The present invention sets five network layers for extracting multi-scale features of the audio modality. For the first network layer, the present invention uses a unidirectional long short-term memory network (LSTM) to capture the temporal features of the audio modality. The output of the first network layer is shown in formula (1):

[0096]

[0097] Step 22: On the basis of obtaining the features F1 a of the first network layer in step S51, in order to keep the method simple and easy to expand, for the subsequent network layers, the present invention uses simple linear layers for feature extraction and dimension transformation, and at the same time uses regularization operations to ensure the stability of the data. Using Relu as the activation function, the output of the second network layer is shown in formulas (2), (3), and (4):

[0098]

[0099]

[0100]

[0101] Step 23: The same operations as in step S52 are adopted for the remaining second to fifth layers. The input audio modality segments are subjected to feature extraction through five hidden layers, obtaining five feature vectors {F1 a, F2 a, F3 a, F4 a, F5 a} at different scales. For the visual modality, the same processing operations as the Audio modality are taken to obtain {F1 v, F2 v, F3 v, F4 v, F5 v}.

[0102] Step 24: According to step S53, a channel dimension is added to the obtained audio modality and visual modality feature sequences Fa and Fv, and then all the hidden layer outputs are stacked according to the channel dimension, as shown in formula (5):

[0103]

[0104] Step 25: According to step S54, the stacked features Fa of the audio modality are obtained, and subsequently, a channel attention mechanism is adopted to calculate the weights of the features of different network layers. The channel attention calculation process is as Figure 3 shown. First, the Squeeze operation is used to compress the data of each channel and summarize the information to obtain a global receptive field. The calculation process is as shown in formula (6):

[0105]

[0106] Secondly, the Excitation operation is used to control the output ratio of different channels, and the calculation formula is as shown in (7):

[0107]

[0108] Finally, a 1×1 convolution kernel is used to integrate the corresponding channels to obtain the final multi-scale features Xa of the audio modality, as shown in formulas (8) and (9):

[0109]

[0110]

[0111] Step 3: As described in step S6, the present invention does not extract multi-scale feature representations for the text modality. In view of the wide application of the Bert model in the text field, the present invention adopts a pre-trained Bert model to extract text modality features;

[0112] Step 31: Use the LSTM model to extract the text modality temporal features, as shown in formula (10):

[0113] F t = LSTM(U t ) Formula (10)

[0114] Step 32: Use the Bert model to extract text modality features, as shown in formula (11):

[0115] X t = Bert(F t ) Formula (11)

[0116] Step 4: As described in step S7, the present invention selects the text modality as the key modality. The selection of the key modality can be based on experimental results and prior knowledge. The modality fusion process is as follows:

[0117] Step 41: The present invention uses cross-modal cross-attention to capture the interaction between the two modalities. The cross-modal cross-attention structure is as Figure 4 shown. In cross-modal cross-attention, the present invention adopts the dot-product attention calculation method:

[0118]

[0119] Step 42: As Figure 3 shown, the cross-modal cross-attention accepts paired modality data. When calculating the text-visual fusion feature, set K = Xt, Q = V = Xv, and calculate Xtv:

[0120]

[0121] Step 43: When calculating the text-visual fusion feature, set K = Xt, Q = V = Xa, and calculate Xta:

[0122]

[0123] Step 44: The text modality uses the self-attention mechanism to capture useful internal information. Set K = Q = V = Xt, and calculate Xtt:

[0124]

[0125] Step 45: According to steps 42, 43, and 44, splice the three calculated fusion features to obtain the multi-modal fusion feature:

[0126] X f = Concat(X tv , X ta , X tt ) Formula (16)

[0127] Step 5: The specific detailed steps of step S8 are as follows:

[0128] Step 51: Calculate the sentiment analysis result of the visual modality:

[0129] y v = classifier v (X v ) Equation (17)

[0130] Step 52: Calculate the sentiment analysis result of the audio modality:

[0131] y a = classifier a (X a ) Equation (18)

[0132] Step 53: Calculate the sentiment analysis result of the text modality:

[0133] y t = classifier t (X t ) Equation (19)

[0134] Step 54: Calculate the sentiment analysis result of the multi-modal fusion feature:

[0135] y f = classifier f (X f ) Equation (20)

[0136] Step 6: According to what is described in Step S9, the steps for obtaining the video sentiment analysis result include:

[0137] Step 61: Perform weighted combination on the four sentiment analysis results calculated in the above Steps 51, 52, 53, and 54 to obtain the sentiment analysis result of the final video:

[0138]

[0139] Step 62: Adopt the method of multi-task learning to iteratively optimize the combined loss function, such as Equation (22), so as to improve the sentiment analysis accuracy of the method:

[0140]

[0141] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the right protection cannot be limited thereby. Those that are only partial changes made according to the claims and the description of the present invention should still fall within the protection scope of the present invention.

Claims

1. Video Emotion Analysis Method Based on Multi-scale Feature Extraction and Multi-task Learning It is characterized in that The specific steps are as follows: Step S1: Extract single-modal feature representations; use a single-modal feature extraction sub-model on the original video clip data to extract single-modal feature representations for video and audio modal data using a feed-forward neural network, and extract feature vectors for the text modality using the BERT model; Step S2: Obtain single-modal multi-scale feature representations; use a channel attention mechanism for the video and audio feature representations in Step S1 to obtain multi-scale feature representations, and then use a 1×1 convolution operation to perform channel integration on them; Step S3: Obtain the multi-modal fusion result: First, capture the interaction between text-video and text-audio in a cross-modal cross-attention manner; Second, for the text modality, use a self-attention mechanism to capture its own internal information; Finally, splice the obtained feature representations to obtain multi-modal fusion features; Step S4: Obtain the video emotion analysis result: Input the fusion features obtained in Step S3 into a multi-modal feature emotion analysis sub-model composed of a fully connected layer and an activation function to obtain the video emotion analysis result; Input the obtained single-modal features into a single-modal emotion analysis sub-model to obtain the single-modal emotion analysis result; Step S5: Define the loss function; train the proposed model through multi-task learning. The loss function of the model consists of four parts, namely: 1) Text emotion analysis loss function; 2) Audio emotion analysis loss function; 3) Video emotion loss function; 4) Multi-modal fusion emotion analysis loss function; The four loss functions are combined into the final loss function by weighting; The obtaining of the single-modal feature representation in Step S2 includes the following steps: Step S21: First, after the video and audio modalities are processed by the feature extraction sub-model in Step S1, the output of each network layer is obtained; Subsequently, a new channel dimension is added to the output features of each layer; Finally, stack them by channel to obtain a stacked feature; As shown in formula (5): Step S22: Process the stacked features using a channel attention mechanism. First, use the Squeeze operation to compress the data of each channel and summarize the information to obtain a global receptive field. The calculation process is as shown in formula (6): Secondly, use the Excitation operation to generate weights for each dimension of the stacked vector; The calculation formula is as shown in (7): Finally, use the Scale operation to multiply the stacked vector by the weights and use a 1×1 convolution kernel to perform channel integration on the result to obtain the final single-modal multi-scale features of video and speech, as shown in formulas (8) and (9): Step S23: For the text modality, use the initial 768-dimensional text features after being processed by bert, and then use a linear layer to reduce its dimension to be equal to the dimension size of the video and speech modalities; The obtaining of the video emotion analysis result in Step S4 includes the following steps: The final multi-modal fusion features are input into a multi-modal sentiment analysis sub-model consisting of three fully connected layers and a tanh activation function, thereby obtaining the final video sentiment analysis result.

2. The video sentiment analysis method based on multi-scale feature extraction and multi-task learning according to claim 1, wherein, the sub-model structure for extracting single-modal feature vectors in step S1 is as follows: Step S11: First, for the feature extraction of the text modality, a pre-trained BERT model is used to extract the text feature representation; Step S12: For the feature extraction of the video and audio modalities, a feature extraction sub-model consisting of five network layers is used to extract the features of each modality; for the first network layer, a unidirectional long short-term memory network LSTM is used to capture the temporal features of the modality data; for the subsequent network layers, four linear layers are used for feature dimension conversion and regularization; the output of the first network layer is shown in formula (1): At the same time, the ReLu function is used as the activation function, and the output of the second network layer is shown in formulas (2), (3) and (4):

3. The video sentiment analysis method based on multi-scale feature extraction and multi-task learning according to claim 1, wherein, the obtaining of the video sentiment analysis result in step S3 includes the following steps: Step S31: Set the three modalities as the text-video modality pair, the text-audio modality pair and the text single modality. For the first two, cross-modal cross-attention is used. The specific operation is that when calculating the attention, set the text feature as the Key, and the features of the other modality as the Value and Query, and at the same time use multi-head attention to calculate the fusion results of text-video and text-audio; Step S32: For the text single modality, use the self-attention mechanism to capture the information useful for the sentiment analysis result in itself; Step S33: Concatenate the results obtained in the above steps to obtain the final multi-modal fusion features.

4. The video sentiment analysis method based on multi-scale feature extraction and multi-task learning according to claim 1, wherein, the definition of the loss function in step S5 includes the following steps: Step S51: Use the single-modal sentiment analysis sub-model corresponding to each modality to process the single-modal features obtained in step S2 to predict the single-modal sentiment analysis result; calculate the loss value between the predicted value of each single-modal sentiment analysis and the true single-modal sentiment label; Step S52: Use the multi-modal sentiment analysis sub-model to process the multi-modal fusion features obtained in step S4 to predict the video sentiment analysis result; calculate the loss value between the predicted video sentiment analysis result and the true video sentiment label; Step S53: The above single-modal prediction results and the fusion prediction result simultaneously affect the loss function finally obtained by the entire model, and the model optimization effect is achieved by optimizing this loss function.