A Video Emotion Description Method Based on Hierarchical Emotion Feature Encoding

Through hierarchical emotion feature coding and multimodal context text generation model, the lack of interaction between emotion classification and description generation in video emotion description is solved, and a more accurate and reliable emotion description is achieved.

CN117292297BActive Publication Date: 2025-07-11UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311251349.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2025-07-11
Estimated Expiration
2043-09-26

AI Technical Summary

Technical Problem

The prior art has the lack of effective interaction and fusion in the emotional classification and description generation subtasks in video emotion descriptions. The emotional categories are too simple and are susceptible to biases in high-frequency common words and low-frequency emotional words in the training corpus, resulting in insufficient description accuracy and robustness.

Method used

Using a hierarchical emotion feature encoding method, video features are extracted through a pre-trained CLIP visual encoder, combined with the pre-trained GloVe network and the Transformer network to fusion of emotional word features, and using the emotion mask mechanism to filter interference words, design a multimodal context text generation model, comprehensively consider video, emotion and text information, and optimize the loss function to improve description accuracy.

Benefits of technology

Effectively filter out irrelevant emotional word interference, improve the accuracy and robustness of the emotional description of the video, and ensure the accuracy and reliability of the generated emotional description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117292297B_ABST
    Figure CN117292297B_ABST
Patent Text Reader

Abstract

The present invention discloses a video emotion description method based on hierarchical emotion feature encoding, and its steps include: 1 video encoding; 2 hierarchical emotion feature encoding; 3 text generation based on multimodal context; 4 model parameter optimization on a video emotion description dataset. The present invention can extract hierarchical fine-grained video emotion clues, filter the interference of irrelevant emotion words on the model, and extract rich context information from three modalities of vision, text, and emotion. Through three emotion-related loss functions, the accuracy of the emotion description, hierarchical emotion encoding, and emotion contrast processes is respectively constrained to generate semantically and emotionally correct video emotion descriptions, thereby improving the accuracy and robustness of the emotion video description model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and relates to technologies such as computer vision, emotion computing, and natural language processing. Specifically, it is a video emotion description method based on hierarchical emotion feature encoding. Background Art

[0002] In the past few decades, researchers in the fields of artificial intelligence and computer vision have been working hard to develop algorithms and models to extract useful information from images and videos, including emotion information. Emotion description is an important research direction in the field of computer vision. Emotion description can be divided into two main directions: text-based emotion description and vision feature-based emotion description. Text-based emotion description refers to using natural language processing technology to analyze and describe emotions in text. Vision feature-based emotion description uses computer vision technology to extract emotion features from images and videos and perform emotion description. The present invention belongs to the latter.

[0003] To improve the accuracy of video emotion analysis, some existing research works have tried different methods, such as extracting richer visual features or modeling different emotion categories. Some other research works have adopted a two-stage training strategy, that is, first using emotion labels for video emotion classification, and then generating video descriptions according to the classification results. However, this method has some defects, such as the lack of effective interaction and fusion between the two subtasks of emotion classification and description generation, and the used emotion categories are too simple and rough, such as positive or negative. There are also some works using attention mechanisms to strengthen the representation and transmission of emotion semantics, but due to the lack of quantitative analysis of emotions, they are easily affected by the deviation of high-frequency common words and low-frequency emotion words in the training corpus. Therefore, there is still great room for improvement and development potential in the task of video emotion description. Summary of the Invention

[0004] In order to overcome the deficiencies of the prior art, the present invention proposes a video emotion description method based on hierarchical emotion feature encoding, in order to automatically extract and encode emotion features and encode them into a hierarchical structure to reflect different aspects and hierarchical characteristics of emotions, so as to better capture and express emotion information, and improve the accuracy and reliability of emotion description.

[0005] The present invention adopts the following technical solutions to solve the technical problems:

[0006] The characteristics of a video emotion description method based on hierarchical emotion feature encoding of the present invention are as follows: It is carried out according to the following steps:

[0007] Step 1, Video Encoding:

[0008] Obtain any video V and its emotional description from the video emotional description dataset, and uniformly sample N video frames from V where f i is the i-th sampled frame; extract the features of the N sampled frames using the pre-trained CLIP visual encoder to obtain the visual features of the video V where v i is the feature of the i-th sampled frame f i ;

[0009] Step 2. Hierarchical emotional feature encoding:

[0010] Step 2.1. Obtain the set of emotional categories E c ={x1,…,x c ,…,xC } , where x c represents the c-th emotional category, and C is the total number of emotional categories;

[0011] Obtain the set of emotional words corresponding to each emotional category in the set of emotional categories E c to form the set of emotional words E w =X1,…,X c ,…,X C}; where X c represents the set of emotional words corresponding to the c-th emotional category, and xc ,j is the j-th emotional word in X c , and M c is the total number of emotional words in X c ;

[0012] Step 2.2. Obtain the text features of each emotional word in the set of emotional words E w through the pre-trained GloVe network where e m is the text feature of the m-th emotional word x m ; M represents the total number of emotional words in the set of emotional words E w , and

[0013] Step 2.3. Use the text feature F e as the key and value of the Transformer network, and use the video feature F v as the query of the Transformer network, so as to obtain the fused feature F e ' output by the Transformer network using Equation (1);

[0014] F e' = Transformer([F v , F e , F e ) (1)

[0015] Input the fused feature F e ' into an average pooling layer and a fully connected layer in sequence, so as to obtain the probability distribution P of the video V over the set of emotion categories E c using Equation (2); c ;

[0016]

[0017] In Equation (2), the output dimension of the fully connected layer is C;

[0018] Step 2.4, Initialize a mask matrix to zero where g i,m represents the element value at the i-th row and m-th column in the mask matrix G;

[0019] Define a parameter K;

[0020] Obtain the emotion categories corresponding to the K largest values in the probability distribution P c to obtain the relevant set of emotion categories E' c ;

[0021] Obtain the relevant set of emotion words E' c corresponding to the relevant set of emotion categories E'; w ;

[0022] If the m-th emotion word x w in the set of emotion words E m is in the relevant set of emotion words E' w , then set the element in the mask matrix G Otherwise, set

[0023] Step 2.5, Use the text feature F e as the key and value of another Transformer network, use the video feature F v as the query of another Transformer network, and use the mask matrix G as the mask of another Transformer network, so as to obtain the emotion feature Fe” output by another Transformer using Equation (3);

[0024] F e ” = Transformer(F v , Fe , F e , G)(3)

[0025] Input the sentiment feature F e ” into another average pooling layer in sequence and another fully connected layer Use Equation (4) to obtain the probability distribution P of video V over the sentiment word set E w ; w ;

[0026]

[0027] In Equation (4), the output dimension of the fully connected layer is M;

[0028] Step 3: Text generation based on multi-modal context:

[0029] Step 3.1: Define the current moment as t and initialize t = 0;

[0030] Step 3.2: Use the pre-trained GloVe network to obtain the text features of the words generated at the previous t moments where w l is the text feature of the l-th generated word;

[0031] Use Equation (5) to obtain the semantic relevance at time t between the i-th visual feature v i of the video V and the text feature w l of the l-th generated word Thus, obtain the alignment matrix at time t

[0032]

[0033] Equation (5), u a , U a , H a , b a are all 4 parameters to be learned; T represents transpose; tanh represents the activation function;

[0034] Step 3.3: Use Equation (6) to obtain the visually aligned text features where w it ' is the i-th visually aligned text feature at time t;

[0035] W t ′ = softmax(A t )W t (6)

[0036] In Equation (6), softmax represents the normalization function;

[0037] Combine the video feature F v , the emotion feature F e ”, and the visually aligned text feature W t ′ to splice into a feature matrix where c it represents the i-th spliced feature at time t;

[0038] Step 3.4: Randomly initialize an LSTM network;

[0039] Use Equation (7) to obtain the attention weight θ it of the i-th spliced feature c t-1 at time t and the hidden state h it of the LSTM network at time t - 1; thus, use Equation (8) to obtain the joint context vector c t ′ at time t;

[0040]

[0041]

[0042] In Equation (10), u θ , U θ , H θ , and b θ are all four learnable parameters in the LSTM network;

[0043] Step 3.5: Use Equation (9) to obtain the hidden state h t of the LSTM network at time t, and thus use Equation (10) to obtain the output probability P t of the video emotion description model at time t;

[0044] h t = LSTM([c t , w t-1 , h t-1 ) (9)

[0045] P t = softmax(W o h t ) (10)

[0046] In Equation (10), W o is the weight matrix to be learned in the LSTM network;

[0047] Step 4: Construct the total loss value composed of the sum of the emotion cross-entropy loss value the hierarchical emotion classification loss value and the emotion contrast loss value And the total loss value is optimized and solved using the stochastic gradient descent method to optimize the model parameters. When it reaches the minimum, the optimal model on the video sentiment description dataset is obtained for realizing the prediction of video sentiment description.

[0048] The feature of the video sentiment description method based on hierarchical sentiment feature encoding according to the present invention also lies in that the total loss value in step 4 is obtained according to the following steps:

[0049] Step 4.1: Calculate the sentiment cross-entropy loss value of the video sentiment description model using formula (11)

[0050]

[0051] In formula (11), T is the total number of time steps calculated by the LSTM network; β is a weight coefficient; is the t-th word in the sentiment description corresponding to video V; represents the indicator function. If is in the sentiment word set E w , then let Otherwise, let be 0;

[0052] Step 4.2: Calculate the hierarchical sentiment classification loss value of the video sentiment description model using formula (12)

[0053]

[0054] In formula (11), is the sentiment category included in the sentiment description corresponding to video V; is the sentiment word included in the sentiment description corresponding to video V;

[0055] Step 4.3: Select a negative video V′ from the video sentiment description dataset that has a different sentiment category from video V; and obtain the semantic relevance between the i-th visual feature and the text feature of the l-th generated word of the negative video V′ according to the process of steps 1 to 3.2

[0056] Calculate the sentiment contrast loss value of the video sentiment description model using formula (13)

[0057]

[0058] In formula (13), σ(·) represents the sigmoid function;

[0059] ​Step 4.4: Sum up the outputs of formulas (11) to (13) to obtain the total loss value.

[0060] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program that supports the processor to execute the video emotion description method, and the processor is configured to execute the program stored in the memory.

[0061] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the video emotion description method when run by a processor.

[0062] Compared with the prior art, the beneficial effects of the present invention are reflected in:

[0063] 1. The present invention proposes a new perspective to solve the emotional video description task, that is, first perceive the emotion of the video, and then generate a description based on the obtained emotional clues, rather than directly mapping the video content to the video description, thereby overcoming the imbalance problem of emotional words and common words in the dataset and improving the accuracy and robustness of the emotional video description model.

[0064] 2. The present invention proposes a hierarchical video emotion feature encoding method. First, determine the emotion category of the video, and then use the emotion mask mechanism to filter emotional words and perform emotion feature encoding, effectively filtering out the interference of irrelevant emotional words to the model, thereby obtaining reliable emotional features and promoting the emotional accuracy of the generated description.

[0065] 3. The present invention comprehensively considers the context information of the video, emotion, and text, and correspondingly designs three emotion-related loss functions to respectively constrain the accuracy of the emotion description, hierarchical emotion encoding, and emotion contrast processes, ensuring the effective training of the emotional video description model, and thus obtaining the correct emotion description. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 It is a flowchart of the video emotion description method based on hierarchical emotion feature encoding according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] In this embodiment, as Figure 1 shown, a video emotion description method based on hierarchical emotion feature encoding includes: 1 video encoding; 2 hierarchical emotion feature encoding; 3 text generation based on multi-modal context; 4 model parameter optimization on the video emotion description dataset. Specifically, it is carried out according to the following steps:

[0068] Step 1: Video encoding:

[0069] Obtain any video V and its emotional description from the video emotional description dataset, and uniformly sample N video frames from V where f i is the i-th sampled frame; in this embodiment, the EmVidCap-L, EmVidCap-S, and EmVidCap video emotional description datasets are used, and N = 30; the features of the N sampled frames are extracted by using a pre-trained CLIP visual encoder to obtain the visual features of video V where v i is the feature of the i-th sampled frame f i .

[0070] Step 2. Hierarchical emotional feature encoding:

[0071] Step 2.1. Obtain the emotional category set E c ={x1,…,x c ,…,x C}, where x c represents the c-th emotional category, and C is the total number of emotional categories; in this embodiment, C = 34;

[0072] Obtain several emotional words corresponding to each emotional category in the emotional category set E c to form the emotional word set E w =X1,…,X c ,…,X C}; where X c represents the emotional word set corresponding to the c-th emotional category, and x c,j is the j-th emotional word in X c , and M c is the total number of emotional words in X c ;

[0073] Step 2.2. Obtain the text features of each emotional word in the emotional word set E w through a pre-trained GloVe network where e m is the text feature of the m-th emotional word x m ; M represents the total number of emotional words in the emotional word set E w , and in this embodiment, M = 179.

[0074] Step 2.3. Use the text feature F e as the key and value of the Transformer network, and use the video feature F vAs the query of the Transformer network, the fused feature F output by the Transformer network is obtained using Equation (1). e '; In this embodiment, the number of layers of the Transformer network is 1;

[0075] F e ' = Transformer([F v , F e , F e ) (1)

[0076] The fused feature F e ' is sequentially input into an average pooling layer and a fully connected layer to obtain the probability distribution P of the video V over the emotion category set E c using Equation (2); c ;

[0077]

[0078] In Equation (2), the output dimension of the fully connected layer is C.

[0079] Step 2.4: Initialize a mask matrix to zero where g i,m represents the element value at the i-th row and m-th column in the mask matrix G;

[0080] Define a parameter K; in this embodiment, K = 5;

[0081] Obtain the emotion categories corresponding to the K largest values in the probability distribution P c to obtain the relevant emotion category set E' c ;

[0082] Obtain the relevant emotion word set E' c corresponding to the relevant emotion category set E' w ;

[0083] If the m-th emotion word x w in the emotion word set E m is in the relevant emotion word set E' w , then set the in the mask matrix G to 1, otherwise set it to 0

[0084] Step 2.5: Use the text feature F e as the key and value of another Transformer network, and use the video feature F vUsing the query of another Transformer network and the mask matrix G as the mask of another Transformer network, the sentiment feature F of the output of another Transformer is obtained by using Equation (3). e ”;

[0085] F e ” = Transformer(F v , F e , F e , G) (3)

[0086] The sentiment feature F e is sequentially input into another average pooling layer and another fully connected layer to obtain the probability distribution P of the video V over the sentiment word set E w using Equation (4); w ;

[0087]

[0088] In Equation (4), the output dimension of the fully connected layer is M.

[0089] Step 3, Text generation based on multimodal context:

[0090] Step 3.1, Define the current time as t and initialize t = 0;

[0091] Step 3.2, Use the pre-trained GloVe network to obtain the text features of the words generated at the previous t moments where w l is the text feature of the l-th generated word.

[0092] Use Equation (5) to obtain the semantic relevance at time t between the i-th visual feature v i of the video V and the text feature w l of the l-th generated word to obtain the alignment matrix at time t

[0093]

[0094] Equation (5), u a , U a , H a , b a are all 4 parameters to be learned; T represents transpose; tanh represents the activation function;

[0095] Step 3.3, Use Equation (6) to obtain the text features of visually aligned where, w it ' is the text feature of the i-th visual alignment at time t;

[0096] W t ′ = softmax(A t )W t (6)

[0097] In Equation (6), Wt represents the normalization function.

[0098] Concatenate the video feature F v , the emotion feature F e ”, and the text feature W t ′ of visual alignment into a feature matrix where, c it represents the i-th concatenated feature at time t.

[0099] Step 3.4: Randomly initialize an LSTM network;

[0100] Use Equation (7) to obtain the attention weight θ it of the i-th concatenated feature c t-1 at time t and the hidden state h it of the LSTM network at time t - 1; thus, use Equation (8) to obtain the joint context vector c t ′ at time t;

[0101]

[0102]

[0103] In Equation (10), u θ , U θ , H θ , b θ are all four learnable parameters in the LSTM network.

[0104] Step 3.5: Use Equation (9) to obtain the hidden state h t of the LSTM network at time t, and thus use Equation (10) to obtain the output probability P t of the video emotion description model at time t;

[0105] h t = LSTM([c t , w t-1 , h t-1 ) (9)

[0106] P t = softmax(W o h t ) (10)

[0107] In formula (10), W o is the weight matrix to be learned in the LSTM network.

[0108] Step 4. Model parameter optimization on the video emotion description dataset:

[0109] Step 4.1. Calculate the emotion cross-entropy loss value of the video emotion description model using formula (11)

[0110]

[0111] In formula (11), T is the total number of time steps calculated by the LSTM network; β is a weight coefficient. In this embodiment, β = 0.2; is the t-th word in the emotion description corresponding to video V; denotes the indicator function. If is in the emotion word set E w then let Otherwise, let be 0.

[0112] Step 4.2. Calculate the hierarchical emotion classification loss value of the video emotion description model using formula (12)

[0113]

[0114] In formula (11), is the emotion category included in the emotion description corresponding to video V; is the emotion word included in the emotion description corresponding to video V.

[0115] Step 4.3. Select a negative video V' from the video emotion description dataset that has a different emotion category from video V; and obtain the semantic relevance between the i-th visual feature and the text feature of the l-th generated word of the negative video V' according to the process of Steps 1 to 3.2

[0116] Calculate the emotion contrast loss value of the video emotion description model using formula (13)

[0117]

[0118] In formula (13), σ(·) represents the sigmoid function;

[0119] Step 4.4. Sum the outputs of formulas (11) to (13) to obtain the total loss value Use the stochastic gradient descent method to optimize the total loss value Optimize and solve to minimize it, so as to obtain the optimal model on the video emotion description dataset for realizing the prediction of video emotion description.

[0120] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0121] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

Claims

1. A video emotion description method based on hierarchical emotion feature encoding, characterized in that It is carried out according to the following steps: Step 1, video encoding: Obtain any video V and its emotional description from the video emotional description dataset, and uniformly sample N video frames from V where f i is the i-th sampled frame; use the pre-trained CLIP visual encoder to extract the features of the N sampled frames to obtain the visual features of the video V where v i is the feature of the i-th sampled frame f i ; Step 2, hierarchical emotion feature encoding: Step 2.1: Obtain the set of emotion categories E c ={x1,…,x c ,…,x C}, where x c represents the c-th emotion category, and C is the total number of emotion categories; Obtain the set of emotion categories \(E\) c A number of emotion words corresponding to each emotion category in it form the emotion word set \(E\) w =\(\{X_1, \ldots, X\) c , \ldots, X\) C \}\); where \(X\) c represents the emotion word set corresponding to the \(c\)-th emotion category, and x\) c,j is the \(j\)-th emotion word in \(X\) c , and \(M\) c is the total number of emotion words in \(X\) c ; Step 2.2: Obtain the set of sentiment words E through the pre-trained GloVe network w The text features of each sentiment word in where e m is the text feature of the m-th sentiment word x m ; M represents the total number of sentiment words in the set of sentiment words E w and Step 2.3: Take the text feature F e as the key and value of the Transformer network, and take the video feature F v as the query of the Transformer network, so as to obtain the fused feature F e ' output by the Transformer network using Equation (1); F e ' = Transformer([F v , F e , F e ) (1) The fused feature F e is sequentially input into an average pooling layer and a fully connected layer so as to obtain the probability distribution P of the video V over the emotion category set E c using Equation (2); c ​ In formula (2), the output dimension of the fully connected layer is C; Step 2.4: Initialize a mask matrix to zero where g i,m represents the element value at the i-th row and m-th column in the mask matrix G; Define a parameter K; Obtain the probability distribution P c for the K largest values in c , thereby obtaining the relevant sentiment category set E′ c ; Obtain the set E' of relevant sentiment categories c The corresponding set E' of relevant sentiment words w ; If the m-th sentiment word x in the sentiment word set E w is in the relevant sentiment word set E' m then, for the mask matrix G w Otherwise, let ​ Step 2.5: Take the text feature F e as the key and value of another Transformer network, take the video feature F v as the query of another Transformer network, and take the mask matrix G as the mask of another Transformer network, so as to obtain the sentiment feature F output by another Transformer using Equation (3) e "; F e ” = Transformer(F v , F e , F e , G)(3) Input the emotional feature F e " into another average pooling layer and another fully connected layer Use Equation (4) to obtain the probability distribution P of the video V over the set of emotional words E w ; w ; In formula (4), the output dimension of the fully connected layer is M; Step 3, text generation based on multimodal context: Step 3.1, define the current moment as t and initialize t = 0; Step 3.2: Obtain the text features of the words generated in the previous t moments using the pre-trained GloVe network where w l is the text feature of the l-th generated word; Obtain the i-th visual feature v of the video V using Equation (5) i with the text feature w of the l-th already generated word l at time t at time t Equation (5), u a , U a , H a , b a are all 4 parameters to be learned; represents transpose; tanh represents the activation function; Step 3.

3. Obtain the visually aligned text features using Equation (6). where w it ' is the i-th visually aligned text feature at time t. W t ′ = softmax(A t )W t (6) In Equation (6), softmax represents the normalization function; Video feature F v and emotional feature F e ", and visually aligned text feature W t ' are concatenated into a feature matrix where c it represents the i-th concatenated feature at time t; Step 3.4, randomly initialize an LSTM network; The splicing feature c of the i-th at time t is obtained using Equation (7) it and the attention weight θ of the hidden state h of the LSTM network at time t-1 t-1 ; thereby, the joint context vector c it ' at time t is obtained using Equation (8) t '; In formula (10), t θ , U θ , H θ , b θ are all four parameters to be learned in the LSTM network; Step 3.

5. Obtain the hidden state h of the LSTM network at time t using Equation (9) t , and thus obtain the output probability P of the video emotion description model at time t using Equation (10) t ; h t = LSTM([c t , w t-1 , h t-1 ) (9) P t = softmax(W o h t ) (10) In formula (10), W o is the weight matrix to be learned in the LSTM network; Step 4: Construct the total loss value composed of the emotional cross-entropy loss value Hierarchical emotional classification loss value Emotional contrast loss value and use the sum of the above losses to form the total loss value And use the stochastic gradient descent method to optimize and solve the total loss value to optimize the model parameters. When it reaches the minimum, obtain the optimal model on the video emotional description dataset for realizing the prediction of video emotional description.

2. The video emotion description method based on hierarchical emotion feature encoding according to claim 1, characterized in that The total loss value in step 4 is obtained according to the following steps: Step 4.

1. Calculate the emotional cross-entropy loss value of the video emotional description model using Equation (11). In Equation (11), T is the total time calculated by the LSTM network; β is a weight coefficient; is the t-th word in the sentiment description corresponding to video V; denotes the indicator function. If is in the sentiment word set E w , then let Otherwise, let be 0; Step 4.2: Calculate the hierarchical sentiment classification loss value of the video sentiment description model using Equation (12). In formula (11), is the emotion category included in the emotion description corresponding to video V; is the emotion word included in the emotion description corresponding to video V; Step 4.3: Select a negative video V′ from the video emotion description dataset, where the emotion category of V′ is different from that of video V; and obtain the semantic relevance between the i-th visual feature of the negative video V′ and the text feature of the l-th generated word according to the process of Steps 1 to 3.2 Calculate the emotion contrast loss value of the video emotion description model using Equation (13). In Equation (13), σ(·) represents the sigmoid function; Step 4.4: Sum the outputs of equations (11) to (13) to obtain the total loss value 3. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program for supporting the processor to execute the video emotion description method according to Claim 1 or 2, and the processor is configured to execute the program stored in the memory.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the video emotion description method according to Claim 1 or 2.

Citation Information

Patent Citations

  • Multi-modal feature fusion video description text generation method

    CN113806587A

  • Aspect-level multi-modal sentiment analysis method based on collaborative attention fusion

    CN115293170A