Short video network public opinion sentiment classification method based on big data
By obtaining and splicing various features of short videos (text, audio, image) in the short video network public sentiment classification, and inputting activation function for judgment, the problems of inaccurate single modal recognition and waste of multimodal data processing resources are solved, and more accurate and robust emotional classification is achieved.
Patent Information
- Application Number
- CN202510319980.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-18
AI Technical Summary
In the classification of public sentiment sentiment in the short video network, the results of single data modal identification are not accurate enough, and the amount of data is huge when multi-modal data is processed, resulting in wasted computing resources.
A method of public sentiment classification for short video networks based on big data is proposed. By obtaining short text features, audio features and image features from video data, splicing them and inputting them into the activation function, and judging the probability of the activation function output and the setting threshold value, the emotional type of the video is judged.
Through multimodal fusion, the accuracy of emotional understanding is improved, the impact of noise is reduced, the robustness is enhanced, complex emotional expressions and dynamic changes are captured, and the accuracy of emotional classification is improved.
Smart Images

Figure CN120182893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to big data technology, deep learning, natural language processing, computer vision technology, and multimodal fields, and particularly relates to a short video network public opinion sentiment classification method based on big data. Background Art
[0002] In social networks, sentiment recognition can monitor the sentiment tendency of users under hot topics in real time, and understand whether the overall attitude of the public towards an event, product or brand is actively supportive, negatively opposed or neutral. For example, after a certain brand launches a new product, by analyzing the comments and discussions of users on social platforms, the brand side can quickly understand the emotional feedback of consumers on the product, and timely adjust the marketing strategy or the direction of product improvement. In addition, through sentiment recognition technology, offensive, insulting or inciting remarks and behaviors can be detected in a timely manner, and potential bad information such as cyber violence and malicious slander can be identified. Once such sentiment-tending content is detected, the platform can quickly take measures, such as blocking and deleting relevant information, warning or punishing the publisher, protecting the legitimate rights and interests of users, and maintaining a healthy online social environment.
[0003] Since the data generated by short video platforms is huge, involving multi-dimensional data such as user behavior, comments, likes, shares, and video content. Traditional data analysis methods often cannot handle such a large-scale data set, so big data technology needs to be used for storage, processing and analysis. In short videos, in addition to text data, the video content itself is also the key to analysis. A video is composed of a large number of frame images. Computer vision technology can extract the visual features of the video by performing image recognition on each frame. Combining information such as facial expressions, body movements, and scene changes in the video can help analyze the emotions conveyed by the video. Common technologies include facial expression recognition, action recognition, scene analysis, etc.
[0004] When the prior art usually uses single data for recognition, the recognition result is not accurate enough; if text data and image data are comprehensively used for recognition, the amount of data is usually huge during the recognition process, resulting in an excessive count and wasting computing resources. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention proposes a short video network public opinion sentiment classification method based on big data. This method obtains short text features, audio features, and image features from video data, splices the obtained features together and inputs them into an activation function, and judges the sentiment type of the video by judging the probability of the output of the activation function and a set threshold. When obtaining the audio features, the audio features are converted into long text, and an intercept threshold N is set. The first N / 2 words at the beginning of the long text and the last N / 2 words at the end of the long text are intercepted, and the intercepted words are spliced together and the audio features are extracted therefrom.
[0006] Further, the calculation process of the truncation threshold N includes:
[0007]
[0008] Among them, min(·) represents finding the minimum value; abs(·) represents taking the absolute value; tanh(·) represents the hyperbolic tangent function; n is the number of sentences in the long text; arsinh(·) represents the inverse hyperbolic sine function; l i represents the length of the i-th sentence; N model represents the text length limit of the pre-trained model embedding vector layer.
[0009] Further, key frames are extracted from the video as image samples, and then image features are extracted from the image samples. When obtaining the key frames, the similarity between the current frame and the previous key frame is judged. If the similarity is less than the set threshold, the current frame is taken as the key frame; otherwise, the next frame of the image is continuously judged.
[0010] Preferably, the calculation of the similarity between the current frame and the previous key frame includes:
[0011]
[0012] Among them, Similarity represents the similarity between the current frame image x and the previous key frame image y; Dict(x, y) represents the distance between the current frame image x and the previous key frame image y in the original video sequence; SSIM(x, y) represents the structural similarity between the current frame image x and the previous key frame image y; ln(·) represents taking the logarithm; M represents the number of frames per second of the current video; MSE(x, y) represents the mean square error between the current frame image x and the previous key frame image y.
[0013] The present invention splices text features, audio features and image features, and then inputs them into an activation function to form an emotional classification model for a short video, and performs multi-modal processing, which can bring the following several significant beneficial effects:
[0014] 1. Integrating multiple information sources can improve the accuracy of emotional understanding. A short video may contain complex emotional expressions, and it may not be possible to fully understand the subtle differences in emotions through a single modality. Multi-modal fusion can ensure that the model captures more emotional dimensions;
[0015] 2. Each modality's features have certain noise or incompleteness. Through multi-modal fusion, the model can supplement and verify with information from other modalities, reduce the influence of noise, and enhance the overall robustness; cross-validation between different modalities helps to eliminate errors caused by modality specificity to improve the understanding of complex emotions;
[0016] 3. Each modality has its own advantages in expressing emotions. After splicing the features of the three modalities, the model can learn the synergistic effects between these modalities, thereby capturing more complex emotional expressions;
[0017] 4. Short videos often have emotional fluctuations, especially in scenarios with obvious emotional turns. Through multimodal fusion, the model can better capture the dynamic changes of emotions.
[0018] 5. In some situations, the emotional classification of a single modality may be biased or contradictory. By fusing the three modalities, the model can establish consistency between the information of different modalities, thereby making better judgments; some emotional expressions may not be easily captured through a single modality, but by combining the information of all modalities, the model can more comprehensively understand the emotional background in the video. Brief Description of the Drawings
[0019] Figure 1 It is a flowchart of a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention;
[0020] Figure 2 It is a schematic diagram of the structure of a deep learning short text emotion classification box in a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention;
[0021] Figure 3 It is a schematic diagram of the structure of a deep learning long text emotion classification box in a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention;
[0022] Figure 4 It is a schematic diagram of named entity recognition in the structure of a deep learning long text emotion classification box in a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention;
[0023] Figure 5 It is a schematic diagram of the structure of a deep learning image collection emotion classification box in a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention;
[0024] Figure 6 It is a schematic diagram of the structure of the overall model emotion classification box in a method for classifying emotions of short-video network public opinion based on big data according to an embodiment of the present invention. Detailed Embodiments
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] The present invention provides a method for classifying the sentiment of short-video network public opinion based on big data. This method obtains short text features, audio features, and image features from video data, splices the obtained features together and inputs them into an activation function, and determines the sentiment type of the video by judging the probability of the output of the activation function compared with a set threshold. When obtaining the audio features, the audio features are converted into long text, and an intercept threshold N is set. The first N / 2 words at the beginning of the long text and the last N / 2 words at the end of the long text are intercepted, and the intercepted words are spliced together and the audio features are extracted therefrom.
[0027] This embodiment provides a method for classifying the sentiment of short-video network public opinion based on big data, as Figure 1 described in detail according to steps S1 to S4 in this embodiment.
[0028] S1. Use the comments and bullet screens of the obtained short videos as the short text existing in the short videos, and obtain the preprocessed short text by removing noise, text tokenization, and removing common words. Then, use the preprocessed short text through the deep learning short text sentiment classification model framework to obtain the short text features in the short videos.
[0029] In this embodiment, the comments and bullet screens in the short videos can obtain the data of the comment area and bullet screens through the API, or can crawl the data through a crawler. The present invention does not limit the sources of the comments and bullet screens in the short videos.
[0030] In the embodiment of the present invention, the preprocessed short text is used through the deep learning short text sentiment classification model framework to obtain the text features in the short videos, as Figure 2 shown, where the deep learning short text sentiment classification model framework includes:
[0031] S11. Convert the original sample into a token sequence according to the DeBERTa model vocabulary;
[0032] Specifically, the DeBERTa model (Decoding-enhanced BERT with Disentangled Attention) is an improved model based on the BERT architecture. It not only inherits the bidirectional encoder structure of BERT, but also enhances the decoding aspect. The decoupled attention mechanism of this model separates the content and position information for processing, so as to better model the context dependence relationship;
[0033] S12. Input the token sequence into the embedding layer of the DeBERTa model to obtain Token Embedding, and extract the [CLS] token embedding to achieve the aggregation of global information;
[0034] Among them, the decoupled attention mechanism and enhanced masked language model (EMLM) based on the DeBERTa model enable the [CLS] token to better aggregate global information. In the sentiment classification task, the [CLS] token embedding can more accurately capture the sentiment tendency of the entire short text;
[0035] S13. Use the [CLS] token embedding as the input and pass it into the Bidirectional Memory Network (Bi-MT layer) for context modeling, so as to capture the information of global and local contexts. Then, use it as the input and pass it into the Global Max Pooling (GM-Pooling layer) to extract the most representative text features. Finally, input it into the All-to-All Connected Layer (AtA-CL layer) to obtain the final short text features;
[0036] Among them, inputting the [CLS] token embedding into the bidirectional memory network can further capture the global and local context information and enhance the modeling ability of sentiment expression. At the same time, combining global max pooling can extract the most representative sentiment features and avoid the interference of irrelevant noise. Finally, through the fully connected layer, the features extracted by DeBERTa can be further fused and optimized to better meet the requirements of the sentiment classification task.
[0037] S2. Convert the audio in the obtained short video into text as the long text existing in the short video. Obtain the preprocessed long text by means of text cleaning, text tokenization, stop word removal, and truncation according to the maximum length of max_seq. Then, pass the preprocessed long text through the deep learning long text sentiment classification framework to obtain the audio features in the short video.
[0038] In the embodiment of the present invention, the method of converting the audio in the short video into text can be implemented through the open-source tool DeepSpeech (an open-source speech recognition engine provided by Mozilla that can convert audio streams into text), Whisper (a multilingual speech recognition model that can not only convert audio into text but also handle various noise situations), or the commercial API Baidu Speech Recognition (supporting Chinese and various dialects). The present invention is not limited thereto.
[0039] In the embodiment of the present invention, in the method of obtaining the preprocessed long text by truncating according to the maximum length of the truncation threshold, the calculation process of the truncation threshold N includes:
[0040]
[0041] Among them, min(·) represents finding the minimum value; abs(·) represents taking the absolute value; tanh(·) represents the hyperbolic tangent function; n is the number of sentences in the long text; arsinh(·) represents the inverse hyperbolic sine function; l i represents the length of the i-th sentence; N model represents the text length limit of the pre-trained model embedding vector layer.
[0042] In this formula, by introducing a limiting factor (including the combination of n + arsinh(n) and the tanh function), this limiting factor ensures that the maximum length of the input sequence will not be too large, thus helping to avoid the problem of insufficient memory during model training; due to the smooth growth characteristic of arsinh(n), it can help the model dynamically adjust the input length according to the complexity of the text data, which enables max_seq to be reasonably controlled according to the actual data volume and will not be fixed at a constant value; through the effective control of the input text length, it can ensure that when the model processes texts of different lengths, it will neither process overly long sequences resulting in unstable training nor lose the capture of important information due to overly short sequences.
[0043] For the truncation method in the preprocessing of long texts, the beginning and end of the long text are respectively truncated by a length of N / 2 and then spliced to obtain the final preprocessed text after truncation; because in short videos, the emotions that the video publisher wants to express are usually at the beginning or end of the video, while the middle of the video is usually the narrative rhythm and the emotions expressed are usually less, so the long text truncated by this method can maximize the excavation of the emotions that the video publisher wants to express in the audio of the short video under the premise of the same text length.
[0044] In the embodiments of the present invention, the preprocessed long text is used to obtain audio features in the short video through a deep learning long text sentiment classification framework, such as Figure 3 shown, where the deep learning long text sentiment classification model framework includes:
[0045] S21. Concatenate and fuse the result obtained by named entity recognition (NER) on the preprocessed text data with the original text to obtain the NER enhanced input;
[0046] Among them, the NER enhanced input data obtained by concatenating and fusing the result obtained by named entity recognition (NER) with the original text can not only enrich the text semantic information, use NER to identify the named entities in the text, and concatenate and fuse these entity information with the original text to enhance the semantic information of the input; at the same time, the NER result can help the model better focus on the key entities in the text, so as to more accurately capture the context information related to emotions;
[0047] In step S21, the result obtained through named entity recognition (NER) is concatenated and fused with the original text to obtain the NER enhanced input, as Figure 4 shown, including:
[0048] S211. Convert the preprocessed long text into a token sequence according to the DeBERTa model vocabulary;
[0049] S212. Input the token sequence into the embedding layer of the DeBERTa model to obtain Token Embedding, and extract the embedding (Token-level Embeddings) of each token, so as to retain the fine-grained information at the lexical level;
[0050] S213. Take the Token-level Embeddings as the input and pass it into the regularization layer (Stochastic DropoutLayer, S-Dropout layer) to prevent overfitting and enhance the generalization effect of the model. Then take it as the input and pass it into the Bi-MT layer (Bidirectional Memory Network) for context modeling. Finally, take it as the input and pass it into the CSM layer (Conditional Sequence Model) to capture the dependencies between labels, and obtain the category label sequence of named entities;
[0051] Among them, by introducing the S-Dropout layer, randomly discarding some neurons during the training process can effectively prevent the model from overfitting and enhance the generalization ability of the model; by introducing the Bi-MT layer, it can capture the forward and backward context information at the same time, so as to more comprehensively understand the semantic relationships in the text and better capture the long-distance dependencies in the long text, thereby improving the recognition ability of complex entities; by introducing the CSM layer, it can model the dependencies between label sequences, thereby improving the accuracy of the NER task and generating more coherent label sequences, thereby improving the overall effect of the NER task;
[0052] S214. Concatenate and fuse the category label sequence of named entities obtained through the CSM layer with the original text to obtain a new text (NER enhanced input text);
[0053] By concatenating and fusing the category label sequence of named entities obtained through the CSM layer with the original text, it can help downstream tasks better understand the key entities and their semantic relationships in the text, thereby improving the performance of the long text sentiment classification task this time;
[0054] S22. Convert the NER enhanced input into a token sequence according to the DeBERTa model vocabulary;
[0055] Among them, the decoupled attention mechanism and relative position encoding of the DeBERTa pre-trained model can better capture the context dependencies in long texts;
[0056] S23. Input the token sequence into the embedding layer of the DeBerta model to obtain Token Embedding, and extract Pooled Output to aggregate the embeddings of the entire sequence;
[0057] Among them, the embedding layer of DeBERTa can generate high-quality token embeddings, providing a solid foundation for subsequent feature extraction and aggregation; through the Pooled Output extracted by the DeBERTa model, the embeddings of the entire sequence can be aggregated to capture global sentiment information;
[0058] S24. Take the Pooled Output as input and input it into the Dual-Gate Recurrent Unit (D-GRU layer) to enhance the context modeling ability and strengthen the feature expression, and finally input it into the AtA-CL layer to obtain the final audio features;
[0059] Among them, by introducing the D-GRU layer and using its dual-gate mechanism (update gate and reset gate), the long-distance dependencies in long texts can be better captured, the context information modeling ability can be enhanced, and at the same time, combined with the Pooled Output of DeBERTa, the feature expression ability can be further enhanced; finally, using the fully connected layer, the features can be deeply fused in a fully connected manner to capture the complex relationships between features, thereby improving the accuracy of sentiment classification.
[0060] S3. Extract the video frames in the obtained short video as the images in the short video, and obtain the preprocessed image set through the methods of extracting key frames, size adjustment, color channel conversion, and normalization processing. Then, use the deep learning image set sentiment classification framework for the preprocessed image set to obtain the image features in the short video.
[0061] In the embodiments of the present invention, the method for extracting key frames is to calculate the similarity Similarity between adjacent key frames based on the change of image content. If the similarity is lower than the threshold t, then select this frame as the key frame. The formula for the similarity Similarity between adjacent key frames is:
[0062]
[0063] Among them, Similarity represents the similarity between the current frame image x and the previous key frame image y; Dict(x, y) represents the distance between the current frame image x and the previous key frame image y in the original video sequence; SSIM(x, y) represents the structural similarity between the current frame image x and the previous key frame image y; ln(·) represents taking the logarithm; M represents the number of frames per second of the current video; MSE(x, y) represents the mean square error between the current frame image x and the previous key frame image y.
[0064] In this formula, by integrating multiple measurement methods, a more comprehensive similarity evaluation can be carried out and the influence of different features can be balanced; by introducing the Dict(x, y) time distance, the temporal relationship between frames can be captured; by introducing the SSIM(x, y) structural similarity, the visual structure information can be better captured and the robustness to illumination and contrast changes can be increased; by introducing the MSE(x, y) mean square error, the differences between pixel levels can be captured and the influence of large error values can be smoothed using logarithmic transformation to avoid the interference of extreme values on the similarity; normalization processing is performed to make the results more stable and easier to compare.
[0065] In the embodiment of the present invention, the pre-processed image collection is passed through a deep learning image collection emotion classification framework to obtain the image features in the short video, such as Figure 5 shown, where the deep learning image collection emotion classification model framework includes:
[0066] S31. Divide the images in the pre-processed image collection into Patches and perform linear embedding to obtain embedding vectors;
[0067] In this embodiment, through Patch division, the local information of the image can be retained, and at the same time, it is converted into a high-dimensional feature representation through linear embedding;
[0068] S32. Input the embedding vectors into three pre-trained models, namely the Vit network (Vision Transformer), the VGG network (Visual Geometry Group), and the ResNet network (Residual Network), respectively, to obtain image features from different angles, and then fuse the image features of the three to obtain an image feature after ensemble learning. Through ensemble learning, the image features are diversified for feature learning and the potential of different models is fully exploited;
[0069] This embodiment realizes multi-feature extraction from three different perspectives (Vit captures global context information through the Transformer architecture, VGG captures local texture and detail information through deep convolutional networks, and ResNet captures multi-level feature representations through residual connections), fully exploiting the potential of different models and compensating for the deficiencies of a single model to enhance the expressive power of features.
[0070] S33. Use the image features after ensemble learning as input and feed them into the Bi-MT layer for temporal information modeling; by introducing the Bi-MT layer, not only can the temporal information in the feature sequence be captured, but also the forward and backward context information can be captured, thus more comprehensively understanding the dependencies in the feature sequence.
[0071] S34. Then use the features containing temporal information as input and feed them into the TriCNN layer (Triple Convolutional Neural Network) to obtain the fusion of local features and global information.
[0072] In this embodiment, by introducing the TriCNN layer, not only can local features of the image be extracted through convolutional operations to capture detail information, but also local features can be fused with global information, thereby generating a more comprehensive feature representation, and the multiple convolutional structures of TriCNN can capture features at different levels, enhancing the expressive power of features.
[0073] S4. Connect the short text features, audio features, and image features obtained from the short video and input them into the activation function, and set a threshold to obtain the result of short video network public opinion sentiment classification prediction.
[0074] In the embodiment of the present invention, after connecting the short text features, audio features, and image features obtained from the short video and inputting them into the activation function, a threshold is set to obtain the result of short video network public opinion sentiment classification prediction. As Figure 6 shown, it includes: connecting the short text features, audio features, and image features obtained from the short video to obtain a short video feature containing all features, inputting the short video feature into the activation function, and finally obtaining the result T finally predicted by the model of the present invention (T is the value obtained by max(X, Y, Z), where X is the probability of negative sentiment, Y is the probability of neutral sentiment, and Z is the probability of positive sentiment). If the value of T is equal to the value of X, the finally predicted result is negative sentiment; if the value of T is equal to the value of Y, the finally predicted result is neutral sentiment; if the value of T is equal to the value of Z, the finally predicted result is positive sentiment.
[0075] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A short video network public opinion sentiment classification method based on big data, which obtains short text features, audio features, and image features from video data, splices the obtained features together and inputs them into an activation function, and judges the emotion type of the video by judging the probability of the activation function output and the set threshold, characterized in that: When obtaining audio features, the audio features are converted into long texts, and a truncation threshold N is set to truncate N / 2 words at the beginning and N / 2 words at the end of the long text, and the truncation words are concatenated together to extract audio features.
2. According to the big data-based short video network public opinion sentiment classification method of claim 1, it is characterized in that: The calculation process of the interception threshold N includes: Among them, min(·) means to find the minimum value; abs(·) means to find the absolute value; tanh(·) means the hyperbolic tangent function; n is the number of sentences in the long text; arsinh(·) means the inverse hyperbolic sine function; l i Indicates the length of the i-th sentence; N model Indicates the text length limit of the embedding vector layer of the pre-trained model.
3. According to a method for classifying short video network public opinion based on big data according to claim 1, it is characterized in that: The barrage and comments that appear in the video are taken as short text data samples, and the DeBerta model is used to obtain short text features from the video data.
4. According to the big data-based short video network public opinion sentiment classification method of claim 3, it is characterized in that: The process of obtaining short text features from video data using the DeBerta model includes: S11, DeBerta model converts the original sample into a token sequence; S12, input the token sequence into the embedding layer to obtain its embedding representation, and extract the hidden state based on the embedding representation; S13, inputting the hidden state into the bidirectional memory network layer to extract context features; S14. The extracted context features are sequentially passed through the global maximum pooling layer and the fully connected layer to obtain short text features.
5. According to a method for classifying short video network public opinion based on big data according to claim 1, it is characterized in that: After the truncated words are concatenated together, named entity recognition is performed, and the recognition result is concatenated with the truncated words as the final long text data. Then, the DeBerta model is used to obtain audio features from the final long text data.
6. According to the method for classifying short video network public opinion based on big data in claim 5, it is characterized in that: The process of obtaining audio features from the final long text data using the DeBerta model includes: S21. The final long text data of the DeBERTa model is converted into a token sequence; S22, input the token sequence into the embedding layer to obtain its embedded representation, and then extract the Pooled Output; S23, input the extracted Pooled Output into the dual gated recurrent unit for data enhancement; S24. The enhanced data is processed through a fully connected layer to obtain audio features.
7. According to a method for classifying short video network public opinion based on big data as described in claim 1, it is characterized in that: Extract key frames from the video as image samples, and then extract image features from the image samples. When obtaining key frames, judge the similarity between the current frame and the previous key frame. If the similarity is less than the set threshold, the current frame is the key frame, otherwise continue to judge the next frame image.
8. According to the big data-based short video network public opinion sentiment classification method of claim 7, it is characterized in that: The calculation of the similarity between the current frame and the previous key frame includes: Where Similarity represents the similarity between the current frame image x and the previous key frame image y; Dict(x,y) represents the distance between the current frame image x and the previous key frame image y in the original video sequence; SSIM(x,y) represents the structural similarity between the current frame image x and the previous key frame image y; ln(·) represents the logarithm; M represents the number of frames per second of the current video; MSE(x,y) represents the mean square error between the current frame image x and the previous key frame image y.
9. A short video network public opinion sentiment classification method based on big data according to claim 1, 7 or 8, characterized in that: The process of obtaining image features includes: S31, preprocessing the image sample, dividing the preprocessed image into patches and performing linear embedding to obtain an embedding vector; S32, input the embedded vectors into the pre-trained Vit network, VGG network, and ResNet network respectively, and concatenate the features obtained respectively to obtain integrated features; S33, inputting the integrated features into the bidirectional memory network layer by layer to model the temporal information, and then passing the data including the temporal information through the TriCNN network to obtain the fusion of local features and global information; The outputs of S34 and TriCNN networks are processed by regularization layer and fully connected layer in turn to obtain image features.
Citation Information
Patent Citations
Mongolian multi-modal sentiment analysis method based on T-M BERT pre-training model
CN114153973A
Text association type short video multi-mode emotion recognition method and system
CN117636196A
Network public opinion text sentiment analysis method, system and equipment and storage medium
CN117874238A