Short video content auditing system based on sentiment analysis
By using a self-attention mechanism and graph convolution network in the short video content review system and combining multimodal features with graph structure, the problem that traditional methods are difficult to identify complex and irregular content is solved, and more efficient sentiment analysis and content review are achieved.
Patent Information
- Application Number
- CN202510484504.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Traditional short video content review methods based on keyword filtering and single modal feature extraction are difficult to effectively identify complex and diverse illegal content, and there are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modals.
The self-attention mechanism and graph convolution network (GCN) are used to achieve the interaction between aspect words and text, and to build a graph structure including text, image and audio features. Through feature alignment and stitching and fusion operations, we ensure that the information between different modes can be fully interacted and complemented.
It significantly improves the accuracy of emotional classification, can more effectively identify complex and diverse illegal content, enhances the ability to capture text emotional tendencies, and maximizes the common information between modals.
Smart Images

Figure CN120014524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent auditing, and specifically refers to a short video content auditing system based on sentiment analysis. Background Art
[0002] With the vigorous development of the Internet, short video platforms have risen rapidly, providing users with a convenient way to obtain information and entertainment. However, the openness and freedom of short video platforms also bring huge challenges to content review. Traditional methods based on keyword filtering are often difficult to effectively identify complex and diverse illegal content. Traditional content review methods mainly rely on technologies based on keyword filtering and single-modal feature extraction. These methods are difficult to meet the needs of effective identification of complex and diverse illegal content. There are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modalities. Existing visual-text sentiment analysis methods usually have poor performance due to limited use of the correlation between different modalities, that is, ignoring the heterogeneity and homogeneity between different modalities. These seriously affect the accuracy and reliability of sentiment analysis. Summary of the invention
[0003] In view of the above situation, in order to overcome the defects of the prior art, the present invention provides a short video content review system based on sentiment analysis. Traditional content review methods mainly rely on technologies based on keyword filtering and single modal feature extraction. These methods are difficult to meet the needs of effective identification of complex and diverse illegal content, and there are problems such as insufficient extraction, ignoring the association between text and target words, and insufficient interaction between modalities. This scheme adopts the self-attention mechanism and graph convolutional network (Graph Convolutional Network, GCN) to fully realize the interaction between aspect words and text to enhance the accuracy of sentiment classification; the existing visual-text sentiment analysis method has limited use of the correlation between different modalities and ignores the heterogeneity and homogeneity between different modalities. This scheme constructs a graph structure including text, image and audio features. By aligning and splicing these features, it ensures that the information between different modalities can fully interact and complement each other. While retaining the uniqueness of each modality, it maximizes the mining of common information between modalities, thereby achieving more efficient sentiment analysis.
[0004] The short video content review system based on sentiment analysis provided by the present invention comprises a data acquisition and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module and a user feedback adjustment module; The data collection and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual labeling on the short video data to obtain training set data; The text analysis module obtains text data in the training set data, including text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes features of the text data to obtain text features; The image content analysis module extracts features of the image and video data in the training set data, identifies objects, scenes, and behaviors in the image and video data, determines whether there is any illegal content, and obtains video image features; The audio analysis module uses automatic speech recognition technology to convert the voice and music in the video into audio text data, and uses the method of the text analysis module to extract features from the audio text data to obtain audio features; The comprehensive decision module integrates information of text features, video image features, and audio features, performs comprehensive analysis, obtains sentiment scores, makes a final review decision, and feeds back the final review decision to the administrator; The administrator of the user feedback adjustment module takes corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for users to submit complaints and feedback to help improve the accuracy of the system and user experience.
[0005] Furthermore, the text analysis module extracts and analyzes the features of the text data, specifically including the following steps: Step S1: Global text feature capture, using classification tokens and separation tokens to convert text and aspect words in text data into the form of classification token-text-separation token and token-aspect word-separation token, using the ALBERT model for encoding, and generating global view features through average pooling operation and self-attention mechanism. The aspect words refer to words specifically used to describe different attributes, characteristics and aspects of an object and entity in the sentiment analysis task. The formula used is as follows: ; ; In the formula, Represents text features, Indicates contextual text features, Indicates the starting position index of the aspect word in the text, is a counting variable, Indicates the text length of the aspect word, Indicates aspect word features, represents the self-attention mechanism, is the global view feature; Step S2: Local feature capture, using the ALBERT model to extract semantic information from the text, capture the contextual relationship of aspect words in the text, and obtain a word embedding matrix containing the contextual relationship; Step S3: Position embedding and distance features, define syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of words. The formula used is as follows: ; ; In the formula, represents the syntactic distance feature, The syntactic distance is defined as represents the maximum possible syntactic distance, represents the weight adjustment coefficient based on position, Represents the characteristics of each word in the text, is the position-aware transformation function; Step S4: Use GCN to capture syntactic relations. Use GCN to capture the syntactic relations between words in the text and update the word features of different layers of GCN. The formula used is as follows: ; In the formula, represents the symmetric normalized adjacency matrix of words, Represents the output result of the previous layer of GCN, and represents the weight matrix and bias term of the GCN layer, for function, Represents the features of the word after GCN update;
[0006] Step S5: Deep semantic feature extraction. Add an attention layer after the GCN layer and use the word embedding matrix to extract local features of word interaction information to obtain local interaction features. The formula used is as follows: ; In the formula, and denote the query, key, and value matrices in the multi-head self-attention mechanism, respectively. Represents local interaction features; Step S6: Concatenate the global view features and local interaction features of the text to form the final text features.
[0007] Furthermore, the image content analysis module extracts features of image and video data in the training set data, specifically comprising the following steps: Step A1: Initialization, dividing the video data into frames and images, dividing the divided images into blocks, into P image blocks; Step A2: Initial image embedding: Use the pre-trained VIT model to embed each image block to obtain the initial image embedding; Step A3: Primary and secondary embeddings, randomly select the primary embedding of an image block and the secondary embeddings of other blocks, the primary embedding is passed through two linear layers to obtain the Key and Value features, and the secondary embedding is connected through a linear layer to obtain the Query feature; Step A4: Cross attention calculation, using the Softmax function to activate the Key and Query features to obtain the cross attention score and Value features; Step A5: Feature output, obtain and output video image features based on the cross-attention score and Key, Value and Query features.
[0008] Furthermore, the comprehensive decision module integrates information of text features, video image features, and audio features, and performs comprehensive analysis, specifically including the following steps: Step B1: Graph structure construction, using text features, video image features, and audio features to construct a relationship graph G; Step B2: Multimodal feature alignment: Based on the network of the mask graph self-attention structure, the text features, video image features, and audio features are aligned and fused according to the relationship graph G, and the maximum pooling and average pooling operations are applied to obtain the fused features; Step B3: Multimodal hidden features, use a multi-layer perceptron to encode the fused features, obtain the multimodal final encoding, and obtain the predicted sentiment score by calculating its cosine similarity with each sentiment aspect; Step B4: Loss function calculation, use the cross entropy loss function to calculate the difference between the predicted sentiment score and the actual sentiment score, and update the parameters of the entire network through back propagation.
[0009] The beneficial effects achieved by the present invention using the above scheme are as follows: (1) Traditional methods based on keyword filtering and single-modal feature extraction are difficult to meet the needs of effectively identifying complex and diverse illegal content. There are problems such as insufficient extraction, ignoring the relationship between text and target words, and insufficient interaction between modalities. This solution uses the self-attention mechanism and graph convolutional network to fully realize the interaction between aspect words and text. The self-attention mechanism can dynamically adjust the importance of different words in the text, so that the model pays more attention to the parts related to the target words, thereby enhancing the ability to capture the emotional tendency of the text; the graph convolutional network further captures the contextual relationships and local features in the text, providing richer information for sentiment classification. Through the combination of these two technologies, this solution fully realizes the interaction between aspect words and text, significantly improving the accuracy of sentiment classification; (2) In view of the fact that existing visual-text sentiment analysis methods have limited use of the correlation between different modalities and ignored the heterogeneity and homogeneity between different modalities, this scheme constructs a graph structure including text, image and audio features. By aligning and concatenating these features, it ensures that the information between different modalities can fully interact and complement each other. While retaining the uniqueness of each modality, it maximizes the mining of common information between modalities, thereby achieving more efficient sentiment analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A schematic diagram of a short video content review system based on sentiment analysis proposed by the present invention; The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention. DETAILED DESCRIPTION
[0011] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments; based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0012] Example 1, see Figure 1 The short video content review system based on sentiment analysis provided by the present invention includes a data acquisition and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module and a user feedback adjustment module; The data collection and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual labeling on the short video data to obtain training set data; The text analysis module obtains text data in the training set data, including text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes features of the text data to obtain text features; The image content analysis module extracts features of the image and video data in the training set data, identifies objects, scenes, and behaviors in the image and video data, determines whether there is any illegal content, and obtains video image features; The audio analysis module uses automatic speech recognition technology to convert the voice and music in the video into audio text data, and uses the method of the text analysis module to extract features from the audio text data to obtain audio features; The comprehensive decision module integrates information of text features, video image features, and audio features, performs comprehensive analysis, obtains sentiment scores, makes a final review decision, and feeds back the final review decision to the administrator; The administrator of the user feedback adjustment module takes corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for users to submit complaints and feedback to help improve the accuracy of the system and user experience.
[0013] Embodiment 2, based on the above embodiment, the text analysis module extracts and analyzes the features of the text data, specifically comprising the following steps: Step S1: Global text feature capture, using classification tokens and separation tokens to convert text and aspect words in text data into the form of classification token-text-separation token and token-aspect word-separation token, using the ALBERT model for encoding, and generating global view features through average pooling operation and self-attention mechanism. The aspect words refer to words specifically used to describe different attributes, characteristics and aspects of an object and entity in the sentiment analysis task. The formula used is as follows: ; ; In the formula, Represents text features, Indicates contextual text features, Indicates the starting position index of the aspect word in the text, is a counting variable, Indicates the text length of the aspect word, Indicates aspect word features, represents the self-attention mechanism, is the global view feature;
[0014] Step S2: Local feature capture, using the ALBERT model to extract semantic information from the text, capture the contextual relationship of aspect words in the text, and obtain a word embedding matrix containing the contextual relationship;
[0015] Step S3: Position embedding and distance features, define syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of words. The formula used is as follows: ; ; In the formula, represents the syntactic distance feature, The syntactic distance is defined as represents the maximum possible syntactic distance, represents the weight adjustment coefficient based on position, Represents the characteristics of each word in the text, is the position-aware transformation function;
[0016] Step S4: Use GCN to capture syntactic relations. Use GCN to capture the syntactic relations between words in the text and update the word features of different layers of GCN. The formula used is as follows: ; In the formula, represents the symmetric normalized adjacency matrix of words, Represents the output result of the previous layer of GCN, and represents the weight matrix and bias term of the GCN layer, for function, Represents the features of the word after GCN update; Step S5: Deep semantic feature extraction. Add an attention layer after the GCN layer and use the word embedding matrix to extract local features of word interaction information to obtain local interaction features. The formula used is as follows: ; ; ; ; In the formula, , and denote the query, key, and value matrices in the self-attention mechanism, respectively. is the word embedding matrix, , and are the weights of the query, key, and value matrices, Represents local interaction features; Step S6: Concatenate the global view features and local interaction features of the text to form the final text features.
[0017] Through the above operations, the traditional methods based on keyword filtering and single-modal feature extraction are difficult to meet the needs of effective identification of complex and diverse illegal content. There are problems such as insufficient extraction, ignoring the relationship between text and target words, and insufficient interaction between modalities. This solution uses the self-attention mechanism and graph convolutional network to fully realize the interaction between aspect words and text. The self-attention mechanism can dynamically adjust the importance of different words in the text, so that the model pays more attention to the parts related to the target words, thereby enhancing the ability to capture the emotional tendency of the text; the graph convolutional network further captures the contextual relationships and local features in the text, providing richer information for sentiment classification. Through the combination of these two technologies, this solution fully realizes the interaction between aspect words and text, and significantly improves the accuracy of sentiment classification; Embodiment 3, based on the above embodiment, the image content analysis module extracts features of image and video data in the training set data, specifically comprising the following steps: Step A1: Initialization, dividing the video data into frames and images, dividing the divided images into blocks, into P image blocks; Step A2: Initial image embedding: Use the pre-trained VIT model to embed each image block to obtain the initial image embedding; Step A3: Primary and secondary embeddings, randomly select the primary embedding of an image block and the secondary embeddings of other blocks, the primary embedding is passed through two linear layers to obtain the Key and Value features, and the secondary embedding is connected through a linear layer to obtain the Query feature; Step A4: Cross attention calculation, using the Softmax function to activate the Key and Query features to obtain the cross attention score; Step A5: Feature output, obtain and output video image features based on the cross-attention score and Key, Value and Query features.
[0018] Embodiment 4, based on the above embodiment, the audio analysis module includes a speech recognition unit and an information processing unit; The speech recognition unit uses automatic speech recognition technology to convert the speech information in the video into readable text data, recognizes the music in the video, and obtains the song information of the music from the big data through audio fingerprint technology, including the music type, emotion, rhythm and text description of the music, to obtain music text data; The information processing unit collectively refers to the extracted readable text data and music text data as audio text data, and uses the method of the text analysis module to extract text features from the audio text data to obtain text features of the audio text data, which are called audio features to distinguish them from the text features obtained by the text analysis module.
[0019] Embodiment 5, based on the above embodiment, the comprehensive decision module integrates information of text features, video image features, and audio features, and performs comprehensive analysis, specifically including the following steps: Step B1: Graph structure construction, using text features, video image features, and audio features to construct a relationship graph G; Step B2: Multimodal feature alignment: Based on the network of the mask graph self-attention structure, the text features, video image features, and audio features are aligned and fused according to the relationship graph G, and the maximum pooling and average pooling operations are applied to obtain the fused features; Step B4: Multimodal hidden features, use a multi-layer perceptron to encode the fused features, obtain the multimodal final encoding, and obtain the predicted sentiment score by calculating its cosine similarity with each sentiment aspect. The formula used is as follows: ; In the formula, represents the cosine similarity function, is a multi-layer perceptron, represents the fused features, Indicates A baseline embedding for sentiment, Statement The final sentiment score for each sentiment aspect, Represents the first Features Step B4: Loss function calculation, use the cross entropy loss function to calculate the difference between the predicted sentiment score and the actual sentiment score, and update the parameters of the entire network through back propagation.
[0020] Embodiment 6. This embodiment is based on the above embodiment. The comprehensive decision module integrates the information of text features, video image features, and audio features, analyzes the fused features, uses a multi-layer perceptron to classify the features into five categories: negative, positive, neutral, violent, and sensitive, and calculates the corresponding sentiment score for each category. According to the sentiment score, an audit decision including pass, mark for review, need to be modified, and delete is generated, and the final audit decision is fed back to the administrator.
[0021] Through the above operations, in view of the fact that the existing visual-text sentiment analysis methods have limited use of the correlation between different modalities and ignore the heterogeneity and homogeneity between different modalities, this scheme constructs a graph structure including text, image and audio features. By aligning and concatenating these features, it ensures that the information between different modalities can fully interact and complement each other. While retaining the uniqueness of each modality, it maximizes the mining of common information between modalities, thereby achieving more efficient sentiment analysis.
[0022] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
[0023] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
[0024] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A short video content review system based on sentiment analysis, characterized by: The system includes a data acquisition and processing module, a text analysis module, an image content analysis module, an audio analysis module, a comprehensive decision-making module, and a user feedback adjustment module; The data collection and processing module collects short video data publicly available on the Internet and social platforms, and performs preliminary processing operations such as format conversion, editing, and manual labeling on the short video data to obtain training set data; The text analysis module obtains text data in the training set data, including text content in video subtitles, titles, descriptions, and comments, and extracts and analyzes features of the text data to obtain text features; The image content analysis module extracts features of the image and video data in the training set data, identifies objects, scenes, and behaviors in the image and video data, determines whether there is any illegal content, and obtains video image features; The audio analysis module uses automatic speech recognition technology to convert the voice and music in the video into audio text data, and uses the method of the text analysis module to extract features from the audio text data to obtain audio features; The comprehensive decision module integrates information of text features, video image features, and audio features, performs comprehensive analysis, obtains sentiment scores, makes a final review decision, and feeds back the final review decision to the administrator; The administrator of the user feedback adjustment module takes corresponding measures according to the final review decision, feeds back the final review decision to the user, and provides a platform for users to submit complaints and feedback to help improve the accuracy of the system and user experience.
2. The short video content review system based on sentiment analysis according to claim 1 is characterized in that: The text analysis module extracts and analyzes the features of the text data, specifically comprising the following steps: Step S1: Global text feature capture, using classification tokens and separation tokens to convert text and aspect words in text data into the form of classification token-text-separation token and token-aspect word-separation token, using the ALBERT model for encoding, and generating global view features through average pooling operations and self-attention mechanisms. The aspect words refer to words specifically used to describe different attributes, characteristics, and aspects of an object and entity in sentiment analysis tasks; Step S2: Local feature capture, using the ALBERT model to extract semantic information from the text, capture the contextual relationship of aspect words in the text, and obtain a word embedding matrix containing the contextual relationship; Step S3: Position embedding and distance features, define syntactic distance, calculate the syntactic distance between words in the text, and adjust the importance of words. The formula used is as follows: ; ; In the formula, represents the syntactic distance feature, The syntactic distance is defined as represents the maximum possible syntactic distance, represents the weight adjustment coefficient based on position, Represents the characteristics of each word in the text, is the position-aware transformation function; Step S4: Use GCN to capture syntactic relations. Use GCN to capture the syntactic relations between words in the text and update the word features of different layers of GCN. Step S5: Deep semantic feature extraction, adding an attention layer after the GCN layer, using the word embedding matrix to extract local features of word interaction information, and obtain local interaction features; Step S6: Concatenate the global view features and local interaction features of the text to form the final text features.
3. The short video content review system based on sentiment analysis according to claim 1 is characterized in that: The comprehensive decision module integrates information of text features, video image features, and audio features, and performs comprehensive analysis, specifically including the following steps: Step B1: Graph structure construction, using text features, video image features, and audio features to construct a relationship graph G; Step B2: Multimodal feature alignment: Based on the network of the mask graph self-attention structure, the text features, video image features, and audio features are aligned and fused according to the relationship graph G, and the maximum pooling and average pooling operations are applied to obtain the fused features; Step B3: Multimodal hidden features, use a multi-layer perceptron to encode the fused features, obtain the multimodal final encoding, and obtain the predicted sentiment score by calculating its cosine similarity with each sentiment aspect; Step B4: Loss function calculation, use the cross entropy loss function to calculate the difference between the predicted sentiment score and the actual sentiment score, and update the parameters of the entire network through back propagation.
Citation Information
Patent Citations
Heterogeneous network construction and distance measurement method fusing multi-modal information
CN109992784A
Audio and video multi-mode sentiment classification method and system
CN113408385A
Multi-modal sentiment analysis method and system based on feedback feature learning
CN116842153A
Emotion visualization analysis method and system for multi-mode short video
CN117892260A
Big data processing method and system based on classification algorithm
CN118585688A