Content compliance detection method and system based on multi-modal large model, and storage medium
By building a large multimodal model, combining text, visual and audio features, and performing feature fusion and weight adjustment, the problem of low accuracy in cross-modal illegal content detection in existing technologies is solved, and efficient content compliance detection is achieved.
Patent Information
- Application Number
- CN202510904850.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-03
AI Technical Summary
Existing content compliance detection methods cannot effectively identify cross-modal illegal content, resulting in low detection accuracy.
A content compliance detection method based on a multimodal large model is adopted. By obtaining the text, visual and audio features of the sample data, a sample heterogeneous graph is constructed, and the feature weights are adjusted according to the scene type, and finally content compliance detection is performed.
It improves the accuracy of content compliance detection, can effectively identify cross-modal illegal content, and is suitable for real-time detection of mobile applications in public scenarios.
Smart Images

Figure CN120751172A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data detection technology, and in particular to a content compliance detection method, system and storage medium based on a multimodal large model. Background Art
[0002] With the rapid development of Internet technology, the online content generated by mobile applications (APPs) has shown an explosive growth trend. All kinds of harmful, illegal, vulgar and other bad content are rampant on the Internet, seriously polluting the network environment, affecting the physical and mental health of netizens, especially minors, and also bringing tremendous pressure to network administrators. Therefore, the content compliance detection of mobile applications playing content in public scenes has received more and more attention.
[0003] In the existing content compliance detection process, text keyword detection is generally used to detect the content of playback data, resulting in the inability to identify cross-modal illegal content (such as the combination of voice and pictures to convey obscure illegal information), thereby reducing the accuracy of content compliance detection. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a content compliance detection method, system and storage medium based on a multimodal large model to solve the problem of low accuracy of content compliance detection in the prior art.
[0005] The embodiment of the present invention is implemented as follows: a content compliance detection method based on a multimodal large model, the method comprising:
[0006] Obtaining sample data and inputting the sample data into a multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features;
[0007] Performing feature fusion on the text features, the visual features, and the audio features to obtain a sample heterogeneous graph, and performing feature weight adjustment on the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph;
[0008] Performing content prediction on the feature heterogeneous graph to obtain a sample prediction result, and determining a model loss based on the sample prediction result;
[0009] updating parameters of the multimodal content compliance detection model according to the model loss until the multimodal content compliance detection model converges;
[0010] The data to be played is input into the converged multimodal content compliance detection model to perform content compliance detection to obtain a content compliance detection result.
[0011] Preferably, the sample data is input into a multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features, including:
[0012] Inputting the text samples in the sample data into the language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments;
[0013] Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words according to the part-of-speech information to obtain the text features;
[0014] Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features;
[0015] Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model to perform noise suppression to obtain denoised samples;
[0016] The frequency spectrum features, voiceprint features and emotion features in the denoised sample are obtained, and the frequency spectrum features, the voiceprint features and the emotion features are combined to obtain the audio features.
[0017] Preferably, feature extraction is performed on the split shot to obtain the visual features;
[0018] Obtaining video key frames in the split shot, and obtaining a histogram of a color space in the video key frames to obtain a color histogram;
[0019] Performing texture recognition on the video key frames to obtain texture features, and performing edge detection on the video key frames;
[0020] Determine an object contour based on an edge detection result, calculate a contour moment of the object contour, and determine a shape descriptor based on the contour moment;
[0021] The color histogram, the texture feature and the shape descriptor are combined to obtain the visual feature.
[0022] Preferably, inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots includes:
[0023] Extracting pixel values of pixel points of the video frames in the video sample respectively according to the visual model, and calculating pixel differences between adjacent video frames according to the pixel values;
[0024] A continuous frame set and a sudden change frame in the video frame are determined according to the pixel difference value, and the continuous frame set and the sudden change frame are segmented respectively to obtain the segmented shots.
[0025] Preferably, the text features, the visual features, and the audio features are subjected to feature fusion to obtain a sample heterogeneous graph, including:
[0026] Performing feature mapping on the text features, the visual features, and the audio features to obtain text nodes, visual nodes, and audio nodes, and performing image description on the visual features according to the text features to obtain description text;
[0027] Accompanying the visual feature according to the audio feature to obtain audio accompaniment information, and matching the audio feature with text according to the text feature to obtain audio corresponding text;
[0028] Generating modal association edges according to the description text, the audio accompaniment information, and the audio corresponding text to obtain a first modal association edge, a second modal association edge, and a third modal association edge;
[0029] The text node and the visual node are connected according to the first modal association edge, the audio node and the visual node are connected according to the second modal association edge, and the text node and the audio node are connected according to the third modal association edge to obtain the sample heterogeneous graph.
[0030] Preferably, adjusting the feature weights of the sample heterogeneous graph according to the scene type of the sample data to obtain the feature heterogeneous graph includes:
[0031] Obtaining a source identifier of the sample data, and matching the source identifier with a scene query table to obtain the scene type;
[0032] The scene type is matched with a weight adjustment table to obtain a target feature and a feature adjustment value, and feature weights of the target features in the sample heterogeneous graph are adjusted according to the feature adjustment value to obtain the feature heterogeneous graph.
[0033] Preferably, after inputting the to-be-played data into the converged multimodal content compliance detection model for content compliance detection and obtaining the content compliance detection result, the method further includes:
[0034] Obtaining historical detection results of the multimodal content compliance detection model, and determining a threshold adjustment value based on a false positive rate and a false negative rate in the historical detection results;
[0035] The compliance threshold of the multimodal content compliance detection model is adjusted according to the threshold adjustment value.
[0036] Another object of an embodiment of the present invention is to provide a content compliance detection system based on a multimodal large model, the system comprising:
[0037] A feature extraction module is used to obtain sample data and input the sample data into the multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features;
[0038] a weight adjustment module, configured to perform feature fusion on the text features, the visual features, and the audio features to obtain a sample heterogeneous graph, and to perform feature weight adjustment on the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph;
[0039] A model training module is used to perform content prediction on the feature heterogeneous graph to obtain sample prediction results, and determine the model loss based on the sample prediction results;
[0040] updating parameters of the multimodal content compliance detection model according to the model loss until the multimodal content compliance detection model converges;
[0041] The compliance detection module is used to input the to-be-played data into the converged multimodal content compliance detection model to perform content compliance detection and obtain a content compliance detection result.
[0042] Preferably, the feature extraction module is further used to:
[0043] Inputting the text samples in the sample data into the language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments;
[0044] Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words according to the part-of-speech information to obtain the text features;
[0045] Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features;
[0046] Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model to perform noise suppression to obtain denoised samples;
[0047] The frequency spectrum features, voiceprint features and emotion features in the denoised sample are obtained, and the frequency spectrum features, the voiceprint features and the emotion features are combined to obtain the audio features.
[0048] The embodiments of the present invention can effectively construct multimodal features by constructing a sample heterogeneous graph, adjust the feature weight of the sample heterogeneous graph according to the scene type of the sample data, effectively achieve the effect of dynamic weight distribution on the sample heterogeneous graph, and improve the accuracy of the feature heterogeneous graph. The multimodal content compliance detection model is updated with parameters through model loss, so that the converged multimodal content compliance detection model can effectively adopt a multimodal feature detection method to perform content compliance detection on cross-modal data to be played, thereby effectively improving the accuracy of content compliance detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flowchart of a content compliance detection method based on a multimodal large model provided by the first embodiment of the present invention;
[0050] Figure 2 2 is a schematic diagram of the structure of a content compliance detection system based on a multimodal large model provided by the second embodiment of the present invention;
[0051] Figure 3 It is a structural diagram of a terminal device provided by the third embodiment of the present invention. DETAILED DESCRIPTION
[0052] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0053] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0054] Example 1
[0055] See also Figure 1 , is a flowchart of a content compliance detection method based on a multimodal large model provided by a first embodiment of the present invention. The content compliance detection method based on a multimodal large model can be applied to any device or system. The content compliance detection method based on a multimodal large model includes the following steps:
[0056] Step S10: acquiring sample data and inputting the sample data into a multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features;
[0057] Sample data includes text samples, video samples, and audio samples. For text samples, the text content is captured in real time through the app's playback interface, bullet screen system, or third-party API interface to obtain text collection data. Text collection data includes subtitles, bullet screen, comments, and voice transcription text. Text preprocessing is performed on the text collection data to obtain sample data.
[0058] In this step, the text preprocessing steps include:
[0059] Text cleaning: Remove noise such as HTML tags, special symbols, emoticons, etc. from text collection data, and use regular expressions to filter invalid characters.
[0060] Word segmentation and part-of-speech tagging: Use Jieba Word Segmenter (Chinese) or Natural Language Toolkit (English) for word segmentation, combined with part-of-speech tagging tools (such as Language Technology Platform) to identify nouns, verbs, adjectives, etc., providing a basis for subsequent semantic analysis.
[0061] Word cloud analysis: Generate word cloud diagrams through word frequency statistics, identify high-frequency words and potentially sensitive words, assist in the formulation of manual review rules, and label the text data collected after word cloud analysis to obtain sample data.
[0062] For video data, we capture video content through screen recording, camera capture, or video stream parsing to obtain video acquisition data, which is then labeled to obtain video samples. For audio samples, we capture audio content through a microphone array or audio stream parsing to obtain audio acquisition data, which is then labeled to obtain audio samples.
[0063] Optionally, the sample data is input into a multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features, including:
[0064] Inputting the text samples in the sample data into the language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments;
[0065] Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words based on the part-of-speech information to obtain the text features; wherein the sample segmented words are matched with a part-of-speech query table to obtain part-of-speech information, adjacent part-of-speech information is combined to obtain a combined part-of-speech, and the combined part-of-speech is matched with a semantic query table to obtain the text features, wherein the part-of-speech query table stores correspondences between different sample segmented words and corresponding part-of-speech information, and the semantic query table stores correspondences between different combined parts of speech and corresponding text features;
[0066] Inputting video samples from the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features; wherein a pixel difference method or a histogram method can be used to detect shot boundaries and distinguish between continuous frames and sudden changes in frames to reduce redundant data; the visual features include color histograms, texture features, and shape descriptors representing contour features;
[0067] Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model for noise suppression to obtain denoised samples; wherein, spectral subtraction or a deep learning model (such as RNNoise) is used to remove background noise in the audio samples to improve speech clarity. In this step, beamforming technology can also be used to focus on the target sound source in the audio sample to suppress sidelobe interference;
[0068] The spectral features, voiceprint features and emotional features in the denoised sample are obtained, and the spectral features, the voiceprint features and the emotional features are combined to obtain the audio features; wherein, the spectral features of the denoised sample are obtained to capture the speech rhythm, and the voiceprint features are obtained by extracting parameters such as the fundamental frequency (F0) and the resonance peak to identify the speaker, and the emotional features are obtained by extracting the emotional labels (such as anger and joy) through a pre-trained emotion recognition model (such as Wav2Vec2.0-Emotion).
[0069] Further, feature extraction is performed on the segmented shots to obtain the visual features;
[0070] Obtaining video key frames from the segmented shots, and obtaining a histogram of a color space in the video key frames to obtain a color histogram; wherein, based on the color histogram, texture features (such as a gray-level co-occurrence matrix), and motion analysis of the segmented shots, extracting 1-3 key frames representing the shot content to obtain the video key frames, and extracting a histogram of an HSV (Hue, Saturation, Value) color space from the video key frames to obtain a color histogram of the video key frames to capture the main color tone of the video key frames;
[0071] Performing texture recognition on the video key frames to obtain texture features, and performing edge detection on the video key frames; wherein, local binary patterns or histograms of directional gradients are used to describe texture details to obtain texture features, and edge detection is performed on the video key frames using a Canny edge detection algorithm;
[0072] Determine the object contour based on the edge detection result, calculate the contour moment of the object contour, and determine the shape descriptor based on the contour moment; wherein the shape descriptor is obtained by calculating the contour moment (such as area, perimeter, center of mass);
[0073] The color histogram, the texture feature and the shape descriptor are combined to obtain the visual feature.
[0074] Furthermore, the video samples in the sample data are input into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, including:
[0075] Extracting pixel values of pixel points of the video frames in the video sample respectively according to the visual model, and calculating pixel differences between adjacent video frames according to the pixel values;
[0076] determining a continuous frame set and a sudden change frame in the video frame according to the pixel difference, and segmenting the continuous frame set and the sudden change frame respectively to obtain the segmented shots;
[0077] Among them, if the pixel difference between adjacent video frames is greater than or equal to the pixel threshold, the adjacent video frames are segmented, and the continuous adjacent video frames are combined to obtain a continuous frame set. If any video frame is not combined with the adjacent video frame, the video frame is determined to be a mutation frame. The pixel threshold can be set according to needs.
[0078] Preferably, in this embodiment, the language model adopts a lightweight BERT variant model (such as TinyBERT or DistilBERT) to balance performance and efficiency. In the pre-training stage of the language model, the language model is pre-trained on a large amount of general corpus to learn language representation. The pre-trained language model is task-adapted on specific domain data (such as political, pornographic, and violence-related texts), and is optimized through Masked Language Model and Next Sentence Prediction tasks to achieve the effect of fine-tuning the language model. During the language model training process, synonym replacement and random insertion / deletion / exchange of words can be used to increase the diversity of training data, and Dropout and L2 regularization can be used to prevent overfitting of the language model. The trained language model is transferred through knowledge distillation to reduce the number of parameters of the language model.
[0079] The visual model is based on the improved YOLOv7-tiny model, adding a temporal attention mechanism (such as the Transformer module) to enhance the dynamic scene detection capability. During the pre-training phase of the visual model, it is trained on large-scale datasets such as ImageNet to learn common visual features. During the fine-tuning phase of the visual model, target detection and classification tasks are optimized on illegal content datasets (such as images / videos containing violence, pornography, and political symbols). The robustness of the model is improved by randomly cropping, rotating, flipping, and adjusting the brightness / contrast of the training data. By freezing the underlying convolutional layers of the visual model, only the high-level fully connected layers are fine-tuned to accelerate convergence. The recall rate is improved by fusing the prediction results of multiple models (such as YOLOv7-tiny and EfficientDet).
[0080] The audio model uses the Wav2Vec2.0 architecture, combined with the Automatic Speech Recognition (ASR) module to achieve speech-to-text conversion and semantic understanding. During the audio model pre-training phase, it is trained on large-scale speech datasets such as LibriSpeech and Common Voice to learn speech representation. During the audio model fine-tuning phase, speech recognition and classification tasks are optimized on domain-specific data (such as speech containing sensitive words and illegal content). Training efficiency is improved by optimizing hyperparameters such as the audio model's batch size and learning rate. The CTC loss function is used to handle the temporal dependencies of speech sequences, and GPUs (such as the NVIDIA RTX 3090) are used for parallel computing to accelerate the inference process.
[0081] Step S20, performing feature fusion on the text features, the visual features, and the audio features to obtain a sample heterogeneous graph, and adjusting feature weights of the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph;
[0082] The multimodal content compliance detection model is equipped with a graph neural network, which is used to fuse and weight text, visual, and audio features. A heterogeneous graph is constructed using text, visual, and audio features as nodes, with nodes connected by modal-related edges (e.g., "text describing an image," "audio accompanying a video"). A graph attention network is used to aggregate information between nodes, dynamically assigning weights through an attention mechanism to highlight key features. Feature fusion involves early, mid-term, and late fusion. Early fusion is used for direct fusion after feature extraction to capture low-level correlation information. Mid-term fusion is used for fusion in the middle layer of the model to balance low-level and high-level features. Late fusion is used for fusion at the model output layer to combine the independent prediction results of each modality.
[0083] Optionally, the text features, the visual features, and the audio features are fused to obtain a sample heterogeneous graph, including:
[0084] Performing feature mapping on the text features, the visual features, and the audio features to obtain text nodes, visual nodes, and audio nodes, and performing image descriptions on the visual features based on the text features to obtain description text; wherein the feature mapping can be configured with mapping relationships as required, and the description text is used to describe the video frame content corresponding to the visual features in text form;
[0085] Accompanying the visual feature according to the audio feature to obtain audio accompaniment information, and assigning text to the audio feature according to the text feature to obtain audio-corresponding text; wherein the audio accompaniment information is used to accompany the video frame content corresponding to the visual feature in an audio manner, and the audio-corresponding text is used to describe the audio content corresponding to the audio feature in a text manner;
[0086] According to the description text, the audio accompaniment information, and the audio corresponding text, modal association edges are generated respectively to obtain a first modal association edge, a second modal association edge, and a third modal association edge; wherein, the description text, the audio accompaniment information, and the audio corresponding text are vector-converted to obtain a first vector, a second vector, and a third vector; the first vector, the second vector, and the third vector are matched with an association edge lookup table to obtain a first modal association edge, a second modal association edge, and a third modal association edge; the association edge lookup table stores correspondences between different vectors and corresponding modal association edges;
[0087] The text node and the visual node are connected according to the first modal association edge, the audio node and the visual node are connected according to the second modal association edge, and the text node and the audio node are connected according to the third modal association edge to obtain the sample heterogeneous graph.
[0088] Furthermore, the feature weight of the sample heterogeneous graph is adjusted according to the scene type of the sample data to obtain a feature heterogeneous graph, including:
[0089] Obtaining a source identifier of the sample data, and matching the source identifier with a scene query table to obtain the scene type; wherein the scene query table stores a correspondence between different source identifiers and corresponding scene types;
[0090] The scene type is matched with the weight adjustment table to obtain the target feature and the feature adjustment value, and the feature weight of the target feature in the sample heterogeneous graph is adjusted according to the feature adjustment value to obtain the feature heterogeneous graph; wherein the weight adjustment table stores the correspondence between different scene types and the corresponding target features and feature adjustment values. In this step, the modal weight is dynamically adjusted according to the scene type (such as live broadcast, short video, long video). For example, the live broadcast scene focuses on audio real-time, and the audio modal weight is increased to 0.6; the short video scene focuses on visual impact, and the visual modal weight is increased to 0.7.
[0091] Step S30, performing content prediction on the feature heterogeneous graph to obtain a sample prediction result, and determining a model loss based on the sample prediction result;
[0092] Among them, the sample labels of the sample data are obtained, and the loss is calculated for the sample labels and sample prediction results to obtain the model loss.
[0093] Step S40: updating parameters of the multimodal content compliance detection model according to the model loss until the multimodal content compliance detection model converges;
[0094] Among them, the parameters of the graph neural network are updated according to the model loss until the number of iterations of the multimodal content compliance detection model is greater than the number threshold or the model loss is less than the loss threshold. The multimodal content compliance detection model is then judged to have converged. The number threshold and loss threshold can be set according to needs.
[0095] Step S50: inputting the to-be-played data into the converged multimodal content compliance detection model to perform content compliance detection and obtain a content compliance detection result;
[0096] Among them, the converged multimodal content compliance detection model extracts multimodal features from the data to be played, fuses the extracted multimodal features to obtain a content heterogeneous graph, adjusts the feature weights of the content heterogeneous graph based on the scene type of the data to be played, obtains a target heterogeneous graph, performs content compliance detection on the target heterogeneous graph, and obtains the content compliance detection results.
[0097] Optionally, after the data to be played is input into the converged multimodal content compliance detection model for content compliance detection, and the content compliance detection result is obtained, it also includes: obtaining the historical detection results of the multimodal content compliance detection model, determining the threshold adjustment value according to the false alarm rate and missed alarm rate in the historical detection results, and adjusting the compliance threshold of the multimodal content compliance detection model according to the threshold adjustment value; wherein, the Q-learning algorithm is used to dynamically adjust the compliance threshold according to the historical detection results (such as the false alarm rate and missed alarm rate). For example, when the false alarm rate is higher than 5%, the threshold is lowered to reduce false alarms; when the missed alarm rate is higher than 3%, the threshold is increased to reduce missed alarms.
[0098] Preferably, in this embodiment, an online learning module is also provided to generate adversarial samples (such as deep fake content) by training the generator (Generator) to enhance the detection capability of the discriminator (Discriminator). For example, images / audio containing obscure illegal information are generated to test whether the multimodal content compliance detection model can accurately identify it. Missed / false detection samples are automatically collected daily, added to the training set after manual review, and the model is iteratively optimized. For example, new detection rules for popular Internet memes (such as "YYDS" and "Juejuezi") are added to reduce the false alarm rate. The Prometheus+Grafana monitoring system is used to track model accuracy, recall rate, F1 value and other indicators in real time. When the indicators drop, an alarm is triggered and the model is automatically rolled back to the historical version.
[0099] In this embodiment, by constructing a sample heterogeneous graph, multimodal features can be effectively constructed. By adjusting the feature weights of the sample heterogeneous graph according to the scene type of the sample data, the sample heterogeneous graph can be effectively dynamically weighted, thereby improving the accuracy of the feature heterogeneous graph. The multimodal content compliance detection model is updated with parameters through model loss, so that the converged multimodal content compliance detection model can effectively use a multimodal feature detection method to perform content compliance detection on cross-modal data to be played, thereby effectively improving the accuracy of content compliance detection. Real-time compliance detection through a multimodal content compliance detection model (text, audio, and video fusion) is suitable for dynamic security review of content played by mobile applications in public scenarios. A "language-visual-audio" trimodal collaborative reasoning engine is constructed, which achieves millisecond-level real-time detection through dynamic feature fusion and lightweight model optimization, and supports adaptive learning of new violation patterns.
[0100] Example 2
[0101] See also Figure 2 , is a schematic diagram of the structure of a content compliance detection system 100 based on a multimodal large model provided in a second embodiment of the present invention, including:
[0102] The feature extraction module 10 is used to obtain sample data and input the sample data into the multimodal content compliance detection model for feature extraction to obtain text features, visual features and audio features.
[0103] Optionally, the feature extraction module 10 is further configured to: input a text sample in the sample data into a language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments;
[0104] Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words according to the part-of-speech information to obtain the text features;
[0105] Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features;
[0106] Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model to perform noise suppression to obtain denoised samples;
[0107] The frequency spectrum features, voiceprint features and emotion features in the denoised sample are obtained, and the frequency spectrum features, the voiceprint features and the emotion features are combined to obtain the audio features.
[0108] Furthermore, the feature extraction module 10 is further configured to: obtain a video key frame in the segmented shot, and obtain a histogram of a color space in the video key frame to obtain a color histogram;
[0109] Performing texture recognition on the video key frames to obtain texture features, and performing edge detection on the video key frames;
[0110] Determine an object contour based on an edge detection result, calculate a contour moment of the object contour, and determine a shape descriptor based on the contour moment;
[0111] The color histogram, the texture feature and the shape descriptor are combined to obtain the visual feature.
[0112] Furthermore, the feature extraction module 10 is further configured to: extract pixel values of pixel points of the video frames in the video sample according to the visual model, and calculate pixel differences between adjacent video frames according to the pixel values;
[0113] A continuous frame set and a sudden change frame in the video frame are determined according to the pixel difference value, and the continuous frame set and the sudden change frame are segmented respectively to obtain the segmented shots.
[0114] The weight adjustment module 11 is used to perform feature fusion on the text features, the visual features and the audio features to obtain a sample heterogeneous graph, and to adjust the feature weights of the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph.
[0115] Optionally, the weight adjustment module 11 is further configured to: perform feature mapping on the text features, the visual features, and the audio features to obtain text nodes, visual nodes, and audio nodes, and perform image description on the visual features based on the text features to obtain description text;
[0116] Accompanying the visual feature according to the audio feature to obtain audio accompaniment information, and matching the audio feature with text according to the text feature to obtain audio corresponding text;
[0117] Generating modal association edges according to the description text, the audio accompaniment information, and the audio corresponding text to obtain a first modal association edge, a second modal association edge, and a third modal association edge;
[0118] The text node and the visual node are connected according to the first modal association edge, the audio node and the visual node are connected according to the second modal association edge, and the text node and the audio node are connected according to the third modal association edge to obtain the sample heterogeneous graph.
[0119] Furthermore, the weight adjustment module 11 is further configured to: obtain a source identifier of the sample data, and match the source identifier with a scene query table to obtain the scene type;
[0120] The scene type is matched with a weight adjustment table to obtain a target feature and a feature adjustment value, and feature weights of the target features in the sample heterogeneous graph are adjusted according to the feature adjustment value to obtain the feature heterogeneous graph.
[0121] A model training module 12 is used to perform content prediction on the feature heterogeneous graph to obtain a sample prediction result, and determine a model loss based on the sample prediction result;
[0122] Parameters of the multimodal content compliance detection model are updated according to the model loss until the multimodal content compliance detection model converges.
[0123] The compliance detection module 13 is configured to input the to-be-played data into the converged multimodal content compliance detection model to perform content compliance detection and obtain a content compliance detection result.
[0124] Optionally, the compliance detection module 13 is further configured to: obtain historical detection results of the multimodal content compliance detection model, and determine a threshold adjustment value according to a false positive rate and a false negative rate in the historical detection results;
[0125] The compliance threshold of the multimodal content compliance detection model is adjusted according to the threshold adjustment value.
[0126] In this embodiment, by constructing a sample heterogeneous graph, multimodal features can be effectively constructed, and the feature weight of the sample heterogeneous graph can be adjusted according to the scene type of the sample data, which can effectively achieve the effect of dynamic weight distribution on the sample heterogeneous graph, thereby improving the accuracy of the feature heterogeneous graph. The parameters of the multimodal content compliance detection model are updated through model loss, so that the converged multimodal content compliance detection model can effectively adopt a multimodal feature detection method to perform content compliance detection on cross-modal data to be played, thereby effectively improving the accuracy of content compliance detection.
[0127] Example 3
[0128] Figure 3 This is a block diagram of a terminal device 2 provided in the third embodiment of the present application. Figure 3 As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for a content compliance detection method based on a multimodal large model. When the processor 20 executes the computer program 22, the steps of each embodiment of the content compliance detection method based on a multimodal large model are implemented.
[0129] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to implement the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0130] The processor 20 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0131] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Furthermore, the memory 21 may include both an internal storage unit of the terminal device 2 and an external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is about to be output.
[0132] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0133] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable storage medium may include: any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunications signals.
[0134] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A content compliance detection method based on a multimodal large model, characterized in that: The method comprises: Obtaining sample data and inputting the sample data into a multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features; Performing feature fusion on the text features, the visual features, and the audio features to obtain a sample heterogeneous graph, and adjusting feature weights of the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph; Performing content prediction on the feature heterogeneous graph to obtain a sample prediction result, and determining a model loss based on the sample prediction result; updating parameters of the multimodal content compliance detection model according to the model loss until the multimodal content compliance detection model converges; The data to be played is input into the converged multimodal content compliance detection model to perform content compliance detection to obtain a content compliance detection result.
2. The content compliance detection method based on a multimodal large model according to claim 1 is characterized in that: The sample data is input into the multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features, including: Inputting the text samples in the sample data into the language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments; Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words according to the part-of-speech information to obtain the text features; Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features; Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model to perform noise suppression to obtain denoised samples; The frequency spectrum features, voiceprint features and emotion features in the denoised sample are obtained, and the frequency spectrum features, the voiceprint features and the emotion features are combined to obtain the audio features.
3. The content compliance detection method based on a multimodal large model according to claim 2 is characterized in that: performing feature extraction on the segmented shots to obtain the visual features; Obtaining video key frames in the split shot, and obtaining a histogram of a color space in the video key frames to obtain a color histogram; Performing texture recognition on the video key frames to obtain texture features, and performing edge detection on the video key frames; Determine an object contour based on an edge detection result, calculate a contour moment of the object contour, and determine a shape descriptor based on the contour moment; The color histogram, the texture feature and the shape descriptor are combined to obtain the visual feature.
4. The content compliance detection method based on a multimodal large model according to claim 2 is characterized in that: Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, including: Extracting pixel values of pixel points of the video frames in the video sample respectively according to the visual model, and calculating pixel differences between adjacent video frames according to the pixel values; A continuous frame set and a sudden change frame in the video frame are determined according to the pixel difference value, and the continuous frame set and the sudden change frame are segmented respectively to obtain the segmented shots.
5. The content compliance detection method based on a multimodal large model according to claim 1 is characterized in that: The text features, the visual features, and the audio features are subjected to feature fusion to obtain a sample heterogeneous graph, including: Performing feature mapping on the text features, the visual features, and the audio features to obtain text nodes, visual nodes, and audio nodes, and performing image description on the visual features according to the text features to obtain description text; Accompanying the visual feature according to the audio feature to obtain audio accompaniment information, and matching the audio feature with text according to the text feature to obtain audio corresponding text; Generating modal association edges according to the description text, the audio accompaniment information, and the audio corresponding text to obtain a first modal association edge, a second modal association edge, and a third modal association edge; The text node and the visual node are connected according to the first modal association edge, the audio node and the visual node are connected according to the second modal association edge, and the text node and the audio node are connected according to the third modal association edge to obtain the sample heterogeneous graph.
6. The content compliance detection method based on a multimodal large model according to claim 1 is characterized in that: Adjusting the feature weights of the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph includes: Obtaining a source identifier of the sample data, and matching the source identifier with a scene query table to obtain the scene type; The scene type is matched with a weight adjustment table to obtain a target feature and a feature adjustment value, and feature weights of the target features in the sample heterogeneous graph are adjusted according to the feature adjustment value to obtain the feature heterogeneous graph.
7. The content compliance detection method based on a multimodal large model according to claim 1 is characterized in that: After inputting the to-be-played data into the converged multimodal content compliance detection model for content compliance detection and obtaining the content compliance detection result, the method further includes: Obtaining historical detection results of the multimodal content compliance detection model, and determining a threshold adjustment value based on a false positive rate and a false negative rate in the historical detection results; The compliance threshold of the multimodal content compliance detection model is adjusted according to the threshold adjustment value.
8. A content compliance detection system based on a multimodal large model, characterized by: The system comprises: A feature extraction module is used to obtain sample data and input the sample data into the multimodal content compliance detection model for feature extraction to obtain text features, visual features, and audio features; a weight adjustment module, configured to perform feature fusion on the text features, the visual features, and the audio features to obtain a sample heterogeneous graph, and to perform feature weight adjustment on the sample heterogeneous graph according to the scene type of the sample data to obtain a feature heterogeneous graph; A model training module is used to perform content prediction on the feature heterogeneous graph to obtain sample prediction results, and determine the model loss based on the sample prediction results; updating parameters of the multimodal content compliance detection model according to the model loss until the multimodal content compliance detection model converges; The compliance detection module is used to input the to-be-played data into the converged multimodal content compliance detection model to perform content compliance detection and obtain a content compliance detection result.
9. The content compliance detection system based on a multimodal large model according to claim 8, characterized in that: The feature extraction module is also used to: Inputting the text samples in the sample data into the language model in the multimodal content compliance detection model for word segmentation to obtain sample word segments; Obtaining part-of-speech information of the sample segmented words, and performing semantic recognition on the sample segmented words according to the part-of-speech information to obtain the text features; Inputting the video samples in the sample data into the visual model in the multimodal content compliance detection model to perform shot segmentation to obtain segmented shots, and performing feature extraction on the segmented shots to obtain the visual features; Inputting the audio samples in the sample data into the audio model in the multimodal content compliance detection model to perform noise suppression to obtain denoised samples; The frequency spectrum features, voiceprint features and emotion features in the denoised sample are obtained, and the frequency spectrum features, the voiceprint features and the emotion features are combined to obtain the audio features.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Large model violation semantic detection method, system and device and medium
CN121579695A
A large model violation semantic detection method, system, device and medium
CN121579695B
Intelligent voice interaction content security compliance test system based on large model
CN121687114A