Short Video Content Review System Based on Sentiment Analysis
By combining the self-attention mechanism and the graph convolution network, the problem of insufficient modal interaction in traditional short video content review is solved, more efficient emotion analysis is achieved, and the accuracy of emotion classification is improved.
Patent Information
- Application Number
- CN202510484504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The traditional short video content review method relies on keyword filtering and single modal feature extraction, making it difficult to effectively identify complex and diverse illegal content, ignore the association between text and target words and the interaction between modals, resulting in poor accuracy and reliability of sentiment analysis.
The self-attention mechanism and graph convolution network (GCN) are used to enhance the interaction between words and text, and a graph structure including text, image and audio features is constructed for feature alignment and stitching and fusion to ensure full interaction and complementarity of different modal information.
The accuracy of emotion classification is significantly improved, and the importance of text is dynamically adjusted through self-attention mechanism, the graph convolution network captures contextual relationships, and constructs a multimodal feature map structure to achieve efficient emotion analysis.
Smart Images

Figure CN120014524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent auditing, and specifically refers to a short video content auditing system based on sentiment analysis. Background Art
[0002] With the vigorous development of the Internet, short video platforms have risen rapidly, providing users with a convenient way to obtain information and entertainment. However, the openness and freedom of short video platforms also bring huge challenges to content review. Traditional methods based on keyword filtering are often difficult to effectively identify complex and diverse illegal content. Traditional content review methods mainly rely on technologies based on keyword filtering and single-modal feature extraction. These methods are difficult to meet the needs of effective identification of complex and diverse illegal content. There are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modalities. Existing visual-text sentiment analysis methods usually have poor performance due to limited use of the correlation between different modalities, that is, ignoring the heterogeneity and homogeneity between different modalities. These seriously affect the accuracy and reliability of sentiment analysis. Summary of the invention
[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a short video content review system based on sentiment analysis. Traditional content review methods mainly rely on technologies based on keyword filtering and single modal feature extraction. These methods are difficult to meet the needs of effective identification of complex and diverse illegal content, and there are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modalities. This scheme adopts the self-attention mechanism and graph convolutional network (Graph Convolutional Network, GCN) to fully realize the interaction between aspect words and text to enhance the accuracy of sentiment classification; the existing visual-text sentiment analysis method has limited use of the correlation between different modalities and ignores the heterogeneity and homogeneity between different modalities. This scheme constructs a graph structure including text, image and audio features. By aligning and splicing these features, it ensures that the information between different modalities can fully interact and complement each other. While retaining the uniqueness of each modality, it maximizes the mining of common information between modalities, thereby achieving more efficient sentiment analysis.
[0004] The short video content review system based on sentiment analysis provided by the present invention comprises a data acquisition and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module and a user feedback adjustment module;
[0005] The data collection and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual labeling on the short video data to obtain training set data;
[0006] The text analysis module obtains text data in the training set data, including text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes features of the text data to obtain text features;
[0007] The image content analysis module extracts features of the image and video data in the training set data, identifies objects, scenes, and behaviors in the image and video data, determines whether there is any illegal content, and obtains video image features;
[0008] The audio analysis module uses automatic speech recognition technology to convert the voice and music in the video into audio text data, and uses the method of the text analysis module to extract features from the audio text data to obtain audio features;
[0009] The comprehensive decision module integrates information of text features, video image features, and audio features, performs comprehensive analysis, obtains sentiment scores, makes a final review decision, and feeds back the final review decision to the administrator;
[0010] The administrator of the user feedback adjustment module takes corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for users to submit complaints and feedback to help improve the accuracy of the system and user experience.
[0011] Furthermore, the text analysis module extracts and analyzes the features of the text data, specifically including the following steps:
[0012] Step S1: Global text feature capture, using classification tokens and separation tokens to convert text and aspect words in text data into the form of classification token-text-separation token and token-aspect word-separation token, using the ALBERT model for encoding, and generating global view features through average pooling operation and self-attention mechanism. The aspect words refer to words specifically used to describe different attributes, characteristics and aspects of an object and entity in the sentiment analysis task. The formula used is as follows:
[0013] ;
[0014] ;
[0015] In the formula, Represents text features, Indicates contextual text features, Indicates the starting position index of the aspect word in the text, is a counting variable, represents the text length of the aspect word, represents the aspect word feature, represents the self-attention mechanism, is the global view feature;
[0016] Step S2: Local feature capture. Use the ALBERT model to extract semantic information from the text, capture the context relationship of the aspect word in the text, and obtain a word embedding matrix containing the context relationship;
[0017] Step S3: Position embedding and distance feature. Define the syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of words. The formula used is as follows:
[0018] ;
[0019] ;
[0020] In the formula, represents the syntactic distance feature, is the defined syntactic distance, represents the maximum possible syntactic distance, represents the position-based weight adjustment coefficient, represents the feature of each word in the text, is the position-aware transformation function;
[0021] Step S4: Use GCN to capture syntactic relationships. Use GCN to capture the syntactic relationships between words in the text and update the word features of different layers of GCN. The formula used is as follows:
[0022] ;
[0023] In the formula, represents the adjacency matrix of symmetrically normalized words, represents the output result of the previous layer of GCN, and represent the weight matrix and bias term of the GCN layer, is function, represents the feature of the word after being updated by GCN;
[0024] Step S5: Deep semantic feature extraction. Add an attention layer after the GCN layer, use the word embedding matrix to extract the local features of the aspect word interaction information, and obtain the local interaction features. The formula used is as follows:
[0025] ;
[0026] In the formula, and respectively represent the query, key, and value matrices in the multi-head self-attention mechanism, represents the local interaction feature;
[0027] Step S6: Concatenate the global view feature and the local interaction feature of the text to form the final text feature.
[0028] Furthermore, the image content analysis module extracts the features of the image and video data in the training set data, which specifically includes the following steps:
[0029] Step A1: Initialization, frame the video data into images, and divide the framed images into blocks, dividing them into P image blocks;
[0030] Step A2: Initial image embedding, use the pre-trained VIT model to perform an embedding operation on each image block to obtain the initial image embedding;
[0031] Step A3: Primary and secondary embeddings, randomly select the primary embedding of one image block and the secondary embeddings of other blocks. The primary embedding obtains the Key and Value features through two linear layers, and the secondary embeddings obtain the Query feature through a linear layer connection;
[0032] Step A4: Cross-attention calculation, use the Softmax function to activate the Key and Query features to obtain the cross-attention scores and the Value features;
[0033] Step A5: Feature output, obtain and output the video image features according to the cross-attention scores and the Key, Value, and Query features.
[0034] Furthermore, the comprehensive decision-making module integrates the information of the text features, video image features, and audio features, and conducts comprehensive analysis, which specifically includes the following steps:
[0035] Step B1: Graph structure construction, use the text features, video image features, and audio features to construct a relational graph G;
[0036] Step B2: Multi-modal feature alignment, based on the network of the masked graph self-attention structure, align and fuse the text features, video image features, and audio features according to the relational graph G, and apply the max pooling and average pooling operations to obtain the fused features;
[0037] Step B3: Multi-modal hidden features, use a multi-layer perceptron to encode the fused features, obtain the multi-modal final encoding, and obtain the predicted sentiment score by calculating the cosine similarity between it and each sentiment aspect;
[0038] Step B4: Loss function calculation. Use the cross-entropy loss function to calculate the difference between the predicted sentiment score and the actual sentiment score, and update the parameters of the entire network through backpropagation.
[0039] The beneficial effects achieved by the present invention using the above solution are as follows:
[0040] (1) Aiming at the problem that traditional methods based on keyword filtering and single-modal feature extraction are difficult to meet the needs of effectively identifying complex and diverse illegal content, and there are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modalities. This solution uses the self-attention mechanism and graph convolutional network to fully realize the interaction between aspect words and text. The self-attention mechanism can dynamically adjust the importance of different words in the text, making the model pay more attention to the parts related to the target word, thereby enhancing the ability to capture the sentiment tendency of the text. The graph convolutional network further captures the context relationship and local features in the text, providing richer information for sentiment classification. Through the combination of these two technologies, this solution fully realizes the interaction between aspect words and text, significantly improving the accuracy of sentiment classification.
[0041] (2) Aiming at the problem that existing visual-text sentiment analysis methods have limited utilization of the correlation between different modalities and ignore the heterogeneity and homogeneity between different modalities. This solution constructs a graph structure including text, image, and audio features. By performing feature alignment and splicing fusion operations on these features, it ensures that the information between different modalities can fully interact and complement each other. While retaining the uniqueness of each modality, it maximally mines the common information between modalities, thereby realizing more efficient sentiment analysis. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of a short video content review system based on sentiment analysis proposed by the present invention;
[0043] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. Detailed Embodiments
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Example 1. Refer to Figure 1, the short video content review system based on sentiment analysis provided by the present invention, the system includes a data collection and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module and a user feedback adjustment module;
[0046] The data collection and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual annotation on the short video data to obtain training set data;
[0047] The text analysis module obtains the text data in the training set data, including the text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes the characteristics of the text data to obtain text features;
[0048] The image content analysis module extracts the features of the image and video data in the training set data, identifies the objects, scenes, and behaviors in the image and video data, and determines whether there is any illegal content to obtain video image features;
[0049] The audio analysis module uses automatic speech technology to identify and convert the speech and music in the video into audio text data, and uses the method of the text analysis module to extract the features of the audio text data to obtain audio features;
[0050] The comprehensive decision-making module integrates the information of text features, video image features, and audio features, and conducts comprehensive analysis to obtain a sentiment score, makes a final review decision, and feeds back the final review decision to the administrator;
[0051] The user feedback adjustment module enables the administrator to take corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for the user to submit appeals and feedback to help improve the accuracy of the system and the user experience.
[0052] Embodiment 2, based on the above embodiment, the text analysis module extracts and analyzes the features of the text data, specifically including the following steps:
[0053] Step S1: Global text feature capture, using classification tokens and delimiter tokens to convert the text and aspect words in the text data into the form of classification token - text - delimiter token and token - aspect word - delimiter token, encoding using the ALBERT model, and generating global view features through average pooling operation and self-attention mechanism. The aspect word refers to the words specifically used to describe different attributes, characteristics, and aspects of an object and entity in the sentiment analysis task. The formula used is as follows:
[0054] ;
[0055] ;
[0056] In the formula, represents the text feature, represents the th context text feature, represents the starting position index of the aspect word in the text, is a counting variable, represents the text length of the aspect word, represents the aspect word feature, represents the self-attention mechanism, is the global view feature;
[0057] Step S2: Local feature capture. Use the ALBERT model to extract semantic information from the text, capture the context relationship of the aspect word in the text, and obtain a word embedding matrix containing the context relationship;
[0058] Step S3: Position embedding and distance feature. Define the syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of words. The formula used is as follows:
[0059] ;
[0060] ;
[0061] In the formula, represents the syntactic distance feature, is the defined syntactic distance, represents the maximum possible syntactic distance, represents the position-based weight adjustment coefficient, represents the feature of each word in the text, is the position-aware transformation function;
[0062] Step S4: Use GCN to capture syntactic relationships. Use GCN to capture the syntactic relationships between words in the text and update the word features of different layers of GCN. The formula used is as follows:
[0063] ;
[0064] In the formula, represents the adjacency matrix of symmetrically normalized words, represents the output result of the previous layer of GCN, and represent the weight matrix and bias term of the GCN layer, is function, represents the feature of the word after being updated by GCN;
[0065] Step S5: Deep semantic feature extraction. Add an attention layer after the GCN layer and use the word embedding matrix to extract local features of word interaction information to obtain local interaction features. The formula used is as follows:
[0066] ;
[0067] ;
[0068] ;
[0069] ;
[0070] In the formula, , and denote the query, key, and value matrices in the self-attention mechanism, respectively. is the word embedding matrix, , and are the weights of the query, key, and value matrices, Represents local interaction features;
[0071] Step S6: Concatenate the global view features and local interaction features of the text to form the final text features.
[0072] Through the above operations, the traditional methods based on keyword filtering and single-modal feature extraction are difficult to meet the needs of effective identification of complex and diverse illegal content. There are problems such as insufficient extraction, ignoring the relationship between text and target words, and insufficient interaction between modalities. This solution uses the self-attention mechanism and graph convolutional network to fully realize the interaction between aspect words and text. The self-attention mechanism can dynamically adjust the importance of different words in the text, so that the model pays more attention to the parts related to the target words, thereby enhancing the ability to capture the emotional tendency of the text; the graph convolutional network further captures the contextual relationships and local features in the text, providing richer information for sentiment classification. Through the combination of these two technologies, this solution fully realizes the interaction between aspect words and text, and significantly improves the accuracy of sentiment classification;
[0073] Embodiment 3, based on the above embodiment, the image content analysis module extracts features of image and video data in the training set data, specifically comprising the following steps:
[0074] Step A1: Initialization, dividing the video data into frames and images, dividing the divided images into blocks, into P image blocks;
[0075] Step A2: Initial image embedding: Use the pre-trained VIT model to embed each image block to obtain the initial image embedding;
[0076] Step A3: Primary and secondary embeddings. Randomly select the primary embedding of one image patch and the secondary embeddings of other patches. The primary embedding obtains the Key and Value features through two linear layers, and the secondary embeddings obtain the Query feature through a linear layer connection.
[0077] Step A4: Cross-attention calculation. Use the Softmax function to activate the Key and Query features to obtain the cross-attention scores.
[0078] Step A5: Feature output. Obtain and output the video image features according to the cross-attention scores and the Key, Value, and Query features.
[0079] Example 4. This example is based on the above example. The audio analysis module includes a speech recognition unit and an information processing unit.
[0080] The speech recognition unit uses automatic speech recognition technology to convert the speech information in the video into readable text data, recognize the music in the video, and obtain the song information of the music from the big data through audio fingerprint technology, including music type, emotion, rhythm, and text description of the music, to obtain the music text data.
[0081] The information processing unit collectively refers to the extracted readable text data and music text data as audio text data, and uses the method of the text analysis module to extract the text features of the audio text data to obtain the text features of the audio text data, which are called audio features to distinguish them from the text features obtained by the text analysis module.
[0082] Example 5. This example is based on the above example. The comprehensive decision-making module integrates the information of text features, video image features, and audio features and conducts comprehensive analysis, specifically including the following steps:
[0083] Step B1: Graph structure construction. Use text features, video image features, and audio features to construct a relational graph G.
[0084] Step B2: Multimodal feature alignment. Based on the network with a masked graph self-attention structure, align and fuse the text features, video image features, and audio features according to the relational graph G, and apply max pooling and average pooling operations to obtain the fused features.
[0085] Step B4: Multimodal hidden features. Use a multi-layer perceptron to encode the fused features, obtain the multimodal final encoding, and obtain the predicted emotion scores by calculating the cosine similarity between it and each emotion aspect. The formula used is as follows:
[0086] ;
[0087] In the formula, represents the cosine similarity function, is a multi-layer perceptron, represents the fused feature, represents the th benchmark embedding of the sentiment aspect, represents the th final sentiment score of the sentiment aspect, represents the th feature in the output of the multi-layer perceptron;
[0088] Step B4: Loss function calculation, use the cross-entropy loss function to calculate the difference between the predicted sentiment score and the actual sentiment score, and update the parameters of the entire network through backpropagation.
[0089] Example 6, based on the above example, the comprehensive decision-making module integrates the information of text features, video image features, and audio features, analyzes the fused features, uses a multi-layer perceptron to classify the features into five categories: negative, positive, neutral, violent, and sensitive, calculates the corresponding sentiment scores for each category, generates an audit decision including pass, marked for review, need to modify, and delete based on the sentiment scores, and feeds back the final audit decision to the administrator.
[0090] Through the above operations, for the existing visual-text sentiment analysis method, due to the limited utilization of the correlation between different modalities and the neglect of the heterogeneity and homogeneity between different modalities, this solution constructs a graph structure including text, image, and audio features. By performing feature alignment and splicing fusion operations on these features, it ensures that the information between different modalities can fully interact and complement each other, maximally mines the common information between modalities while retaining the uniqueness of each modality, thereby achieving more efficient sentiment analysis.
[0091] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0092] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0093] The above description of the present invention and its embodiments is not restrictive. What is shown in the drawings is only one of the embodiments of the present invention, and the actual structure is not limited thereto. In general, if those of ordinary skill in the art are inspired by it and, without departing from the gist of the present invention, design similar structural modes and embodiments without creative efforts, they shall fall within the protection scope of the present invention.
Claims
1. A short video content review system based on sentiment analysis, characterized in that: The system includes a data acquisition and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module, and a user feedback and adjustment module; The data acquisition and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual annotation on the short video data to obtain training set data; The text analysis module obtains the text data in the training set data, including the text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes the features of the text data to obtain text features; The image content analysis module extracts the features of the image and video data in the training set data, identifies the objects, scenes, and behaviors in the image and video data, and determines whether there is any illegal content to obtain video image features; The audio analysis module uses automatic speech technology to convert the speech and music in the video into audio text data, and uses the method of the text analysis module to extract the features of the audio text data to obtain audio features; The comprehensive decision-making module integrates the information of text features, video image features, and audio features, and conducts comprehensive analysis to obtain an emotion score, makes a final review decision, and feeds back the final review decision to the administrator; The user feedback and adjustment module enables the administrator to take corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for the user to submit appeals and feedback to help improve the accuracy of the system and the user experience; The comprehensive decision-making module integrates the information of text features, video image features, and audio features, and conducts comprehensive analysis, specifically including the following steps: Step B1: Graph structure construction, using text features, video image features, and audio features to construct a relational graph G; Step B2: Multimodal feature alignment, based on a network with a masked graph self-attention structure, align and fuse the text features, video image features, and audio features according to the relational graph G, and apply max pooling and average pooling operations to obtain the fused features; Step B3: Multimodal hidden features, using a multi-layer perceptron to encode the fused features, obtain the multimodal final encoding, and obtain the predicted emotion score by calculating the cosine similarity between it and each emotion aspect; Step B4: Loss function calculation, using the cross-entropy loss function to calculate the difference between the predicted emotion score and the actual emotion score, and updating the parameters of the entire network through backpropagation.
2. The short video content review system based on sentiment analysis according to claim 1, wherein: The text analysis module extracts and analyzes the features of the text data, specifically including the following steps: Step S1: Global text feature capture, using classification tokens and separator tokens to convert the text and aspect words in the text data into the form of classification token - text - separator token and token - aspect word - separator token, encoding using the ALBERT model, and generating global view features through average pooling operation and self-attention mechanism. The aspect word refers to a word specifically used to describe different attributes, characteristics, and aspects of an object and entity in an emotion analysis task; Step S2: Local feature capture. Use the ALBERT model to extract semantic information from the text, capture the context relationship of the aspect words in the text, and obtain a word embedding matrix containing the context relationship. Step S3: Position embedding and distance feature. Define the syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of the words. The formula used is as follows: ; ; In the formula, represents the syntactic distance feature, is the defined syntactic distance, represents the maximum possible syntactic distance, represents the position-based weight adjustment coefficient, represents the feature of each word in the text, is the position-aware transformation function; Step S4: Use GCN to capture syntactic relationships. Use GCN to capture the syntactic relationships between words in the text and update the word features of different layers of GCN. Step S5: Deep semantic feature extraction. Add an attention layer after the GCN layer, use the word embedding matrix to extract the local features of the aspect word interaction information, and obtain the local interaction features. Step S6: Concatenate the global view features and local interaction features of the text to form the final text features.
Citation Information
Patent Citations
Big data processing method and system based on classification algorithm
CN118585688A