Multi-modal content intelligent auditing and violation detection method and system
By separating audio and video and constructing a multimodal resource target dictionary, combined with a spatiotemporal encoder and knowledge graph, the problem of low efficiency in reviewing fake short videos was solved, and efficient detection of fake information and generation of visual reports were achieved.
Patent Information
- Application Number
- CN202511001186.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-04
AI Technical Summary
Existing methods for reviewing fake short videos are inefficient and struggle to effectively detect and identify multimodal trust differences in user-manipulated content, leading to the spread of information risks.
By separating audio and video in the video stream, extracting audio-to-text and visual keyframes in parallel, constructing a multimodal resource target dictionary, using intermodal confidence difference indicators to determine violations, combining spatiotemporal encoders and knowledge graphs for cross-modal evidence verification, and dynamically adjusting modal weights.
It improves the efficiency of false information review and detection, ensures the accuracy of cross-modal feature time alignment, enables minute-level adjustments and visualization report generation, improves preprocessing efficiency, and enhances adversarial detection capabilities.
Smart Images

Figure CN120892602A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of information processing, and particularly relates to a multi-modal content intelligent auditing and violation detection method and system. BACKGROUND
[0002] The proportion of short video users among netizens reaches 94.8%, and the scale of short video users has broken through 1 billion for the first time. Due to the strong subjectivity, easy processing and entertainment of short video production, some users manipulate short video content to cheat traffic. That is, by exaggerating, distorting, misleading or even directly fabricating short video content, a false short video is generated to obtain user trust and cause potential information risks. The false short video refers to a short video with factual errors in content produced by editing, special effects, sound mixing and other means, which misleads the audience and spreads false content by tampering with real scenes, fabricating events and fabricating facts.
[0003] The spread of false short videos may cause information risks and harm the network ecology. By revealing the influence of different multi-modal semantic features of false short videos on user trust, it is helpful to suppress the spread of false short videos from the perspective of users.
[0004] It is found through research that the perceived trust of users for false short videos manipulated by text fabrication, concealment, ambiguity and incitement, audio fabrication, concealment and incitement, image fabrication and incitement classes are significantly different:
[0005] ① At the text level, the perceived trust of users for text fabrication is higher, and the perceived trust of users for text concealment, ambiguity and incitement is lower;
[0006] ② At the audio level, the perceived trust of users for audio fabrication and concealment is higher, and the perceived trust of users for audio incitement is lower;
[0007] ③ At the image level, the perceived trust of users for image fabrication is higher, and the perceived trust of users for image incitement is lower.
[0008] The research can clarify the formation mechanism of user trust in the context of false short videos and provide support for optimizing false short video content governance policies. SUMMARY
[0009] The purpose of the present application is to provide a multi-modal content intelligent auditing and violation detection method and system, which separates audio and video from video stream, processes audio-to-text and visual key frame extraction in parallel, constructs a multi-modal resource target dictionary, and makes violation judgment according to the confidence difference index between modalities, solving the problems of existing false information auditing difficulty and low detection efficiency.
[0010] To solve the above technical problems, the present application is realized by the following technical scheme:
[0011] This invention relates to a multimodal content intelligent review and violation detection method, comprising the following steps:
[0012] Step S1: Perform audio-visual separation on the video stream, and process audio-to-text and visual keyframe extraction in parallel;
[0013] Step S2: Construct a spatiotemporal encoder to record the correspondence between timestamps of each mode;
[0014] Step S3: Construct a multimodal resource target dictionary and use relational categories to form a knowledge graph between entities;
[0015] Step S4: Train the shared semantic space through comparative learning and calculate the confidence difference index between modalities;
[0016] Step S5: Perform violation determination, trigger the sensitive feature threshold for any modality, and initiate multimodal evidence cross-validation;
[0017] Step S6: Establish a feature library of violation cases to support incremental learning and dynamically adjust the weight coefficients of each modality.
[0018] As a preferred technical solution, the specific process of separating audio and video in step S1 is as follows:
[0019] Step S11: Adaptive spectrum analysis is used to dynamically separate the environment and speech based on the frequency characteristics of human voice;
[0020] Step S12: Develop a video frame semantic clustering algorithm to automatically identify shot switching points as keyframe extraction nodes;
[0021] Step S13: Establish audio and video energy matrices and detect abnormal silence or black screen segments;
[0022] The specific process of parallel processing of audio-to-text and visual keyframe extraction is as follows:
[0023] Step S14: Real-time speech-to-text conversion (ASR + voiceprint recognition) is performed using a dual audio verification mechanism for the audio channel;
[0024] Step S15: Keyframe extraction is used to introduce motion saliency detection for the visual channel to avoid missing dynamic violations in fixed-interval sampling, and a spatiotemporal interest point detector (STIP) is developed to capture sensitive actions.
[0025] As a preferred technical solution, the specific steps of adaptive spectrum analysis for speech separation in step S11 are as follows:
[0026] Step S111, dynamic frequency band division: a mel-scale filter bank is constructed, the filter center frequency is dynamically adjusted according to the data audio distribution, the 80Hz-8000Hz human voice feature frequency band is extracted, the double-threshold endpoint detection technology is adopted, and the speech / non-speech segment is preliminarily segmented by combining the short-time energy and the zero-crossing rate;
[0027] Step S112, voiceprint feature enhancement: the MFCC cepstrum coefficient is extracted to construct a speaker voiceprint template library, the target human voice and the environment sound are distinguished by a Gaussian mixture model, a frequency domain masking matrix is designed, and adaptive attenuation is performed on the non-human voice frequency band (attenuation coefficient α = 1-speech probability);
[0028] Step S113, multi-modal verification: the lip movement feature and the speech activity detection result are synchronously analyzed, and the re-separation is triggered when the audio human voice segment and the video lip movement do not match;
[0029] Step S114, joint analysis of space-time features: the HSV histogram difference of the continuous frames is calculated, the lens boundary is detected by combining the optical flow motion vector (threshold θ = 0.35), the space-time interest points are extracted by using a 3D convolutional neural network, and sensitive action modes such as violence and nudity are identified;
[0030] Step S115, semantic clustering optimization: the frame-level deep features are extracted by using ResNet-152, similar scenes are merged by DBSCAN clustering (ε = 0.4, minPts = 5), the key frame priority score is established, and the specific score calculation formula is as follows: Score = 0.6*visual saliency + 0.3*motion intensity + 0.1*audio event correlation;
[0031] Step S116, abnormal segment detection: an audio-video energy matrix is constructed, the horizontal dimension is the frame number, and the vertical dimension is [audio RMS energy, picture brightness, motion intensity]; when 10 consecutive frames satisfy brightness < 15lux and audio energy < 0.1, it is determined as a black screen silent segment;
[0032] The working process of the cross-modal correspondence engine is as follows:
[0033] Hierarchical time coding is performed; wherein, a first time axis: synchronizing each modal data stream based on the video PTS clock; a second time axis: labeling sensitive content trigger points (such as the moment when a dirty word appears, the starting frame of a violent action); a third correlation axis: establishing a cross-modal causal chain (such as the appearance of an induced picture after the speech instruction “click here”);
[0034] The audio-video synchronization deviation degree is calculated, and when the audio event timestamp-corresponding visual event timestamp is greater than 200ms, it is determined as a potential adversarial sample.
[0035] As a preferred technical solution, in step S2, the space-time encoder is designed as a three-level hierarchical timestamp system:
[0036] The first anchor point is based on the video PTS time reference, the second anchor point is based on the trigger point of each modal feature event (such as the moment when a sensitive word appears), and the third anchor point is based on the cross-modal event causal chain (such as the appearance of a violation picture after a voice instruction).
[0037] As a preferred technical solution, in the step S3, a multi-modal resource target dictionary is constructed, and a knowledge graph between entities is formed by using a relationship category. The specific process is as follows:
[0038] Step S31: Determine the modal type of the resource, which is used for specialized processing of features of different modalities;
[0039] Step S32: Obtain resource labels, and convert the content of unstructured data and features without labeled labels;
[0040] Step S33: Match the obtained labels against the multi-modal resource directory dictionary;
[0041] Step S34: Generate a resource directory according to the cataloging rules;
[0042] Step S35: Match the resource directory corresponding to the data content according to the data label in the knowledge representation;
[0043] Step S36: Extract the data content to form a complete expression and complete the knowledge representation process.
[0044] As a preferred technical solution, in the step S4, the specific steps of calculating the inter-modal confidence difference index by comparing and learning the shared semantic space are as follows:
[0045] Step S41: Use a dual-flow Transformer architecture to process text and visual features respectively, design a modal-specific adapter layer, and project the original features to a 256-dimensional public subspace;
[0046] Step S42: Construct a triple sample, adopt a cross-modal InfoNCE loss function, and add an inter-modal attention penalty term;
[0047] Step S43: Fuse the multi-modal content detection model of the consistency of text and image semantics;
[0048] Step S44: Use a multilayer perceptron containing a three-layer fully connected network as a classifier.
[0049] As a preferred technical solution, in the step S41, the text content T and the visual content V of the collected information are obtained, wherein the text is represented as a set T = {w1, w2, …, wn} composed of n word groups, each word is converted into a vector E n} and the visual content V is represented as a set V = {v1, v2, …, vn} composed of n visual features, each visual feature is converted into a vector E i = [ei1 ,e i2 ,...,e ij In the formula, i represents the i-th word, and j represents the dimension of the word vector; for visual content V, it is mapped to a label set V = {t1, t2, ..., tm} consisting of m labels through a pre-trained CNN. m}, which are then converted into word vectors using the Word2Vec method;
[0050] The word vectors are summed, and the mean of each dimension is calculated to obtain a vector with the same dimension as each word, which is used as the semantic vector Em of the text content or image content. The specific calculation formula is as follows:
[0051]
[0052] The semantic consistency between text and image is measured by the cosine distance between the text semantic vector Em(T) and the image semantic vector Em(V), as shown in the following formula:
[0053]
[0054] As a preferred technical solution, in step S43, the multimodal content detection model includes text feature extraction, image feature extraction, and text-image semantic consistency feature extraction;
[0055] The text feature extraction is performed using a Bi-LSTM model. The input is the text word vector W transformed from word vectors, and the output is the set of hidden states H at each step. The result after mean pooling is used as the text feature f. T The formula for its calculation is as follows: f T =AveragePooling(Bi-LSTM([W]));
[0056] The image feature sports model is fused, and a ResNet network is used to extract image features. The data output from the ResNet network is then input into a fully connected layer to reduce the dimensionality of the image features, resulting in image features f. I Its calculation process can be expressed as: f I =σ(w×Resnet+b); where w and b represent the weights and biases of the fully connected layer, respectively, and σ represents the activation function;
[0057] The text and image semantic consistency feature extraction is performed on the image content, five pre-trained CNN models are used for semantic labeling of the image, and labels with a probability greater than 0.01 output by each model are selected as image semantic labels, different convolution kernel numbers and convolution kernel sizes are used for the five CNN models to realize abstract representation of image data, and vectorization representation is performed on the image semantic labels, words are mapped to a vector space, vector cosine is used to calculate the text and image semantic consistency feature, similarity calculation is performed between the text vector and the label vector of the five models, and finally the semantic consistency feature f containing five elements is obtained s , and the calculation formula is as follows: f S =[Consistency1, Consistency2, Consistency3, Consistency4, Consistency5].
[0058] As a preferred technical solution, in the step S44, each piece of information is represented by f T , f I represents the image feature, f S represents the semantic consistency feature, and the fused tweet feature can be represented as F, and the specific formula is as follows: A three-layer fully connected network multilayer perceptron is used as a classifier, and a Softmax function is used to output a classification result.
[0059] As a preferred technical solution, in the step S5, the violation determination includes primary alarm, deep verification, and confrontation detection.
[0060] The primary alarm adopts a dynamic threshold decision mechanism, a multi-modal sensitive feature library is preset, a feature confidence score is calculated in real time, a hierarchical threshold strategy is adopted, and the parameters are automatically adjusted according to the time period / scene to realize single-mode rapid response.
[0061] The feature confidence score is calculated in real time, and a hierarchical threshold strategy is adopted:
[0062] Emergency threshold (> 90%): immediate blocking (such as detecting violent keywords);
[0063] Warning threshold (70%-90%): triggering deep verification (such as suspected voice change);
[0064] Monitoring threshold (< 70%): only record logs;
[0065] At the same time, single-mode rapid response can also be performed, such as:
[0066] Text channel: metaphor violation identification based on syntax tree analysis (such as “apple → prohibited goods alias”);
[0067] Audio channel: 3D voiceprint fingerprint matching abnormal events (such as sudden explosion + continuous buzzing);
[0068] Visual channel: spatiotemporal interest point detection dynamic violation (such as continuous body conflict action);
[0069] The deep verification includes cross-modal consistency checking and knowledge graph assisted decision making; the cross-modal consistency checking is used for detecting timestamp deviation of events of each modality (such as voice "praising the scenery" but violent scene appearing in the picture -> similarity <0.3 triggering an alarm), calculating audio-video semantic similarity and identifying caption shielding behavior (OCR text coverage <80% and voiceprint anomaly); the knowledge graph assisted decision making dynamically allocates according to a scene by constructing evidence chain confidence; the evidence chain confidence formula is:
[0070] Comprehensive confidence = text weight × C_text + audio weight × C_audio + visual weight × C_visual;
[0071] The adversarial detection includes voice transformation attack defense, visual adversarial sample cracking and OCR adversarial protection; the voice transformation attack defense performs live detection on a voiceprint, such as analyzing fundamental frequency stability (normal speaking fundamental frequency fluctuation <50Hz); the visual adversarial sample cracking strengthens sensitive area features through an attention mechanism, and identifies image disturbance in combination with high-frequency or low-frequency information; the OCR adversarial protection performs font disturbance detection to compare glyph structure similarity.
[0072] The application is a kind of multi-modal content intelligent auditing and violation detection system, comprising a multi-modal acquisition layer, a feature extraction engine, a cross-modal correlation system and a dynamic decision center;
[0073] The multi-modal acquisition layer comprises a distributed crawler module and a streaming processing module; the distributed crawler mode is used for multi-channel content grabbing from web pages, live broadcast and APP; the streaming processing module is used for real-time disassembly of video into text, audio and time-frequency;
[0074] The feature extraction engine comprises a text analysis unit, an audio processing unit and a visual analysis unit; the text analysis unit fuses syntax tree analysis and metaphor recognition to detect metamorphic sensitive words; the audio processing unit identifies voice transformation, human voice and background anomalies by constructing a voiceprint fingerprint library; the visual analysis unit is used for capturing dynamic sensitive scenes through a spatiotemporal attention model;
[0075] The cross-modal correlation system comprises a semantic alignment module, a contradiction detector and a knowledge graph interface; the semantic alignment module is used for establishing a shared embedding space of multi-modal feature vectors; the contradiction detector is used for identifying adversarial behavior of audio-video asynchrony; the knowledge graph interface is used for accessing a violation expression feature library to realize semantic extension;
[0076] The dynamic decision center comprises a multi-stage filtering pipeline and an interpretable output; the multi-stage filtering pipeline is used for screening to deep evidence chain analysis; and the interpretable output is used for generating a visual report containing violation segment positioning.
[0077] The present application has the following advantages:
[0078] (1) The present application improves the efficiency of false information review and detection by separating audio and video, processing audio-to-text and visual key frame extraction in parallel, constructing a multi-modal resource target dictionary, and making violation judgment according to the confidence difference index between modalities.
[0079] (2) The present application establishes an audio-picture energy matrix through intelligent audio-video separation technology to detect abnormal silence or black screen segments, and adopts a double-engine verification mechanism for real-time speech-to-text, and introduces motion saliency detection for key frame extraction to avoid missing dynamic violation innovations in time and space coding structure and parallel processing architecture, which improves the preprocessing efficiency by about 60% compared with traditional methods, and ensures that the time alignment accuracy of cross-modal features reaches milliseconds.
[0080] (3) The present application innovatively realizes the process based on the violation judgment stage of the multi-modal content review system, comprehensively determines the dynamic threshold, constructs the cross-modal evidence chain, and realizes the robust defense mechanism, trains the LSTM prediction threshold curve based on the historical violation data, realizes the minute-level adjustment, labels the violation segment timing positioning and modal contradiction heat map, and generates a visual report.
[0081] Of course, implementing any product of the present application does not necessarily need to achieve all the advantages described above at the same time. BRIEF DESCRIPTION OF DRAWINGS
[0082] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0083] Figure 1 A multi-modal content intelligent review and violation detection method flow chart of the present application;
[0084] Figure 2 A multi-modal content intelligent review and violation detection system structure schematic diagram of the present application. DETAILED DESCRIPTION
[0085] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of the present application.
[0086] In addition, the technical features involved in each of the embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0087] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Figures 1-2
[0088] Embodiment one
[0089] Please refer to Figure 1 The present application is a multi-modal content intelligent auditing and violation detection method, which comprises the following steps:
[0090] Step S1: audio and picture separation is performed on the video stream, and audio-to-text and visual key frame extraction are processed in parallel;
[0091] Step S2: a space-time encoder is constructed to record the correspondence between the timestamps of each modality;
[0092] Step S3: a multi-modal resource target dictionary is constructed, and a knowledge graph between entities is formed using relationship categories;
[0093] Step S4: a shared semantic space is trained by comparison and learning, and a confidence difference index between modalities is calculated;
[0094] Step S5: violation determination is performed, sensitive feature thresholds for any modality are triggered, and multi-modal evidence cross-validation is started;
[0095] Step S6: a violation case feature library is established to support incremental learning, and the weight coefficients of each modality are dynamically adjusted.
[0096] In step S1, the specific process of audio and picture separation of the video stream is as follows:
[0097] Step S11: adaptive spectrum analysis is used to dynamically separate the environment and speech according to the frequency characteristics of human voice;
[0098] Step S12: a video frame semantic clustering algorithm is developed to automatically identify shot cut points as key frame extraction nodes;
[0099] Step S13: Establish an audio, picture energy matrix to detect abnormal silence or black screen segments;
[0100] The specific process of parallel processing of audio-to-text and visual key frame extraction is as follows:
[0101] Step S14: Real-time speech-to-text (ASR+voiceprint recognition) is performed using a dual-audio verification mechanism for the audio channel;
[0102] Step S15: Key frame extraction is introduced for the visual channel using motion saliency detection to avoid missing dynamic violations due to fixed interval sampling, and a spatiotemporal interest point detector (STIP) is developed to capture sensitive actions.
[0103] In step S11, the specific steps of speech separation using adaptive spectral analysis are as follows:
[0104] Step S111, dynamic frequency band division: Construct a Mel-scale filter bank, dynamically adjust the filter center frequency according to the data audio distribution, extract the 80Hz-8000Hz human voice feature frequency band, use a double-threshold endpoint detection technique, and combine short-time energy and zero-crossing rate to achieve preliminary segmentation of speech / non-speech segments;
[0105] Step S112, voiceprint feature enhancement: Extract MFCC cepstrum coefficients to construct a speaker voiceprint template library, distinguish target human voice from environmental sound using a Gaussian mixture model, and design a frequency domain masking matrix to implement adaptive attenuation (attenuation coefficient α = 1 - speech probability) on non-human voice frequency bands;
[0106] Step S113, multi-modal verification: Simultaneously analyze lip movement features and speech activity detection results, and trigger re-separation when the audio human voice segment does not match the video lip movement;
[0107] Step S114, joint analysis of spatiotemporal features: Calculate the HSV histogram difference of consecutive frames, detect shot boundaries using optical flow motion vectors (threshold θ = 0.35), extract spatiotemporal interest points using a 3D convolutional neural network, and identify sensitive action patterns such as violence and nudity;
[0108] Step S115, semantic clustering optimization: Extract frame-level deep features using ResNet-152, merge similar scenes using DBSCAN clustering (ε = 0.4, minPts = 5), establish a key frame priority score, and the specific score calculation formula is as follows: Score = 0.6 * visual saliency + 0.3 * motion intensity + 0.1 * audio event correlation;
[0109] Step S116, Abnormal segment detection: construct a sound and picture energy matrix, the horizontal dimension is the frame number, and the vertical dimension is [audio RMS energy, picture brightness, motion intensity]; when 10 consecutive frames meet the condition of brightness < 15 lux and audio energy < 0.1, it is determined as a black screen and mute segment;
[0110] The cross-modal correspondence engine works as follows:
[0111] Hierarchical time coding is performed; wherein, the first time axis: synchronizing each modality data stream based on the video PTS clock; the second time axis: labeling sensitive content trigger points (such as the moment when a dirty word appears, the starting frame of a violent action); the third correlation axis: establishing a cross-modal causal chain (such as the appearance of an induced picture after the voice instruction "click here");
[0112] Calculate the sound and picture synchronization deviation, and determine that it is a potential adversarial sample when the audio event timestamp - corresponding visual event timestamp is greater than 200 ms.
[0113] Voice content review helps content producers detect risky or rule-breaking content in audio files or voice streams (such as live streams), such as spam, advertising, political, terrorist, abusive, pornographic, spam, prohibited, meaningless content. The risk scenarios that should be supported for identification are shown in the following table:
[0114]
[0115] In step S2, the space-time encoder is designed as a three-level hierarchical timestamp system:
[0116] The first anchor point is based on the video PTS time reference, the second marker is based on the trigger point of each modality feature event (such as the moment when a sensitive word appears), and the third correlation is based on the cross-modal event causal chain (such as the appearance of a rule-breaking picture after a voice instruction).
[0117] In step S3, a multi-modal resource target dictionary is constructed, and a knowledge graph is formed between entities using relationship categories, and the specific process is as follows:
[0118] Step S31: Determine the modality type of the resource, which is used to specialize the features of different modalities;
[0119] Step S32: Obtain resource label process, convert the content of unstructured data and features without labeled labels;
[0120] Step S33: Match the obtained labels against the multi-modal resource directory dictionary;
[0121] Step S34: Generate a resource directory according to the cataloging rules;
[0122] Step S35: Match the resource directory corresponding data content according to the data label in the knowledge representation;
[0123] Step S36: Extract the data content, form a complete expression, and complete the knowledge representation process.
[0124] In step S4, the specific steps for calculating the intermodal confidence difference index by training the shared semantic space through comparative learning are as follows:
[0125] Step S41: Using a dual-stream Transformer architecture, design modality-specific adapter layers to process text and visual features separately, projecting the original features onto a 256-dimensional common subspace;
[0126] Step S42: Construct triplet samples, use the cross-modal InfoNCE loss function, and add an inter-modal attention penalty term;
[0127] Step S43: Integrate a multimodal content detection model that combines text and image semantic consistency;
[0128] Step S44: Use a multilayer perceptron containing a three-layer fully connected network as the classifier.
[0129] In step S41, the text content T and visual content V of the collected information are obtained, where the text is represented as a set T = {w1, w2, ..., w...} consisting of n word groups. n}, convert each word into a vector E i =[e i1 ,e i2 ,...,e ij In the formula, i represents the i-th word, and j represents the dimension of the word vector; for visual content V, it is mapped to a label set V = {t1, t2, ..., tm} consisting of m labels through a pre-trained CNN. m}, which are converted into word vectors using the Word2Vec method;
[0130] The word vectors are summed, and the mean of each dimension is calculated to obtain a vector with the same dimension as each word, which is used as the semantic vector Em of the text content or image content. The specific calculation formula is as follows:
[0131]
[0132] The semantic consistency between text and image is measured by the cosine distance between the text semantic vector Em(T) and the image semantic vector Em(V), as shown in the following formula:
[0133]
[0134] In step S43, the multimodal content detection model includes text feature extraction, image feature extraction, and text-image semantic consistency feature extraction;
[0135] The text feature extraction extracts a text word vector W by inputting a Bi-LSTM model, and outputs a hidden state set H at each step. The average pooling result is used as the text feature f T The calculation process is as follows: f T = AveragePooling(Bi-LSTM([W]));
[0136] The text detection is based on a large amount of text feature library, rule library, keyword library, and NLP algorithm text filtering analysis, which helps the content producer to detect whether the formulated text contains illegal information, such as pornography, terrorism, politics, advertising, illegal, abuse, low-quality water, negative comments, ideological risk warning, and supports custom text black library;
[0137] The specific text detection typical risk scenarios are as follows:
[0138]
[0139]
[0140] The image feature sports model is fused, and the Resnet network is used to extract the image feature, and the data output by the Resnet network is input into a fully connected layer to reduce the image feature dimension, and the image feature f I The calculation process can be represented as: f I = σ(w x Resnet + b); In the formula, w and b represent the weight and bias of the fully connected layer, respectively, and σ represents the activation function;
[0141] The picture detection applies an artificial intelligence active learning algorithm to quickly detect illegal content such as pornography, politics, terrorism, junk advertising, graphic violation, and picture logo through a deep learning model. The typical risk scenarios that should be supported for identification are as follows:
[0142]
[0143]
[0144] The text and image semantic consistency feature extraction uses five pre-trained CNN models to perform semantic labeling on image content, and selects labels with a probability greater than 0.01 output by each model as image semantic labels. Different convolution kernel numbers and convolution kernel sizes are used for five CNN models to realize abstract representation of image data, and image semantic labels are vectorized to map words to vector space. The vector cosine is used to calculate the text and image semantic consistency feature. The text vector is calculated with the label vector of the five models, and the semantic consistency feature f containing five elements is finally obtained.s , the calculation formula is as follows: f S =[Consistency1, Consistency2, Consistency3, Consistency4, Consistency5].
[0145] In step S44, f T represents the text features, f I represents the image features, and f S represents the semantic consistency features. The fused tweet features can be represented as F, and the specific formula is as follows: A multilayer perceptron with a three-layer fully connected network is used as a classifier, and the classification result is output through a Softmax function.
[0146] In step S5, the violation determination includes primary alarm, deep verification, and confrontation detection.
[0147] The primary alarm adopts a dynamic threshold decision mechanism, predefines a multi-modal sensitive feature library, calculates the feature confidence score in real time, adopts a hierarchical threshold strategy, and automatically adjusts the parameters according to the time period / scene to realize single-mode rapid response.
[0148] The feature confidence score is calculated in real time, and a hierarchical threshold strategy is adopted:
[0149] Emergency threshold (> 90%): immediate blocking (such as detecting violent keywords);
[0150] Warning threshold (70%-90%): trigger deep verification (such as suspected voice change);
[0151] Monitoring threshold (< 70%): only record logs;
[0152] At the same time, single-mode rapid response can also be performed, such as:
[0153] Text channel: metaphor violation identification based on syntax tree analysis (such as "apple → contraband code name");
[0154] Audio channel: three-dimensional voiceprint fingerprint matching of abnormal events (such as sudden explosion + continuous humming);
[0155] Visual channel: spatiotemporal interest point detection of dynamic violations (such as continuous body conflict actions);
[0156] The deep verification includes cross-modal consistency inspection and knowledge graph assisted decision; the cross-modal consistency inspection is used for detecting timestamp deviation of events of each mode (such as voice “praising the scenery” but violent scene appears in the picture → similarity <0.3 triggers an alarm), calculating audio-visual semantic similarity and identifying caption shielding behavior (OCR text coverage <80% and voiceprint anomaly); the knowledge graph assisted decision dynamically allocates according to the scene by constructing evidence chain confidence; the evidence chain confidence formula is:
[0157] Comprehensive confidence = text weight × C_text + audio weight × C_audio + visual weight × C_visual;
[0158] The adversarial detection includes voice conversion attack defense, visual adversarial sample cracking and OCR adversarial protection; the voice conversion attack defense performs live detection on the voiceprint, such as analyzing the stability of the fundamental frequency (normal speaking fundamental frequency fluctuation <50Hz); the visual adversarial sample cracking strengthens the feature of the sensitive area through the attention mechanism, and identifies the image disturbance in combination with high-frequency or low-frequency information; the OCR adversarial protection detects font disturbance to compare the structural similarity of the characters.
[0159] Embodiment two
[0160] Referring to Figure 2 The application is a multi-modal content intelligent auditing and violation detection system, which can be used to perform the method content of embodiment 1 of the application, including: a multi-modal acquisition layer, a feature extraction engine, a cross-modal correlation system and a dynamic decision center;
[0161] The multi-modal acquisition layer includes a distributed crawler module and a streaming processing module; the distributed crawler mode is used for multi-channel content grabbing from web pages, live broadcasts and APPs; the streaming processing module is used for real-time disassembly of videos into text, audio and time-frequency;
[0162] The feature extraction engine includes a text analysis unit, an audio processing unit and a visual analysis unit; the text analysis unit fuses syntax tree analysis and metaphor recognition to detect deformation sensitive words; the audio processing unit identifies voice conversion, human voice and background anomalies by constructing a voiceprint fingerprint library; the visual analysis unit is used to capture dynamic sensitive scenes through a spatiotemporal attention model;
[0163] The cross-modal correlation system includes a semantic alignment module, a contradiction detector and a knowledge graph interface; the semantic alignment module is used to establish a shared embedding space of multi-modal feature vectors; the contradiction detector is used to identify adversarial behavior of audio-visual asynchrony; the knowledge graph interface is used to access a violation expression feature library to realize semantic extension;
[0164] The dynamic decision center includes a multi-level filtering pipeline and an interpretable output; the multi-level filtering pipeline is used for analysis from preliminary screening to deep evidence chain; the interpretable output is used to generate a visual report containing the positioning of the violation segment.
[0165] It is worth noting that the above system embodiments, including each unit is only divided according to the functional logic, but not limited to the above division, as long as the corresponding function can be realized; in addition, the specific name of each functional unit is only for the convenience of mutual distinction, and is not used to limit the protection scope of the present application.
[0166] In addition, those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by a program instructing related hardware, and the corresponding program can be stored in a computer readable storage medium.
[0167] The preferred embodiments of the application disclosed above are only used to help explain the application. The preferred embodiments do not describe all the details and limit the application to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the present application. The present application selects and describes these embodiments in order to better explain the principles and practical applications of the present application, so that those skilled in the art can well understand and utilize the present application. The present application is limited by the claims and their full scope and equivalents.
Claims
1. A method for intelligent review and violation detection of multimodal content, characterized in that, Includes the following steps: Step S1: Perform audio-visual separation on the video stream, and process audio-to-text and visual keyframe extraction in parallel; Step S2: Construct a spatiotemporal encoder to record the correspondence between timestamps of each mode; Step S3: Construct a multimodal resource target dictionary and use relational categories to form a knowledge graph between entities; Step S4: Train the shared semantic space through comparative learning and calculate the confidence difference index between modalities; Step S5: Perform violation determination, trigger the sensitive feature threshold for any modality, and initiate multimodal evidence cross-validation; Step S6: Establish a feature library of violation cases to support incremental learning and dynamically adjust the weight coefficients of each modality.
2. The method for intelligent review and violation detection of multimodal content according to claim 1, characterized in that, In step S1, the specific process for separating audio and video in the video stream is as follows: Step S11: Adaptive spectrum analysis is used to dynamically separate the environment and speech based on the frequency characteristics of human voice; Step S12: Automatically identify camera switching points as keyframe extraction nodes; Step S13: Establish audio and video energy matrices and detect abnormal silence or black screen segments; The specific process of parallel processing of audio-to-text and visual keyframe extraction is as follows: Step S14: Use a dual audio verification mechanism for real-time speech-to-text conversion for the audio channel; Step S15: Use keyframe extraction for the visual channel to introduce motion saliency detection and capture sensitive actions.
3. The method for intelligent review and violation detection of multimodal content according to claim 2, characterized in that, In step S11, the specific steps for speech separation using adaptive spectrum analysis are as follows: Step S111, Dynamic frequency band division: Construct a Mel-scale filter bank, dynamically adjust the center frequency of the filter according to the audio data distribution, and extract the human voice characteristic frequency band of 80Hz-8000Hz. Step S112, Voiceprint Feature Enhancement: Extract MFCC cepstral coefficients to construct a speaker voiceprint template library, and use a Gaussian mixture model to distinguish the target human voice from ambient sound; Step S113, Multimodal Verification: Simultaneously analyze lip movement features and speech activity detection results. When the audio voice segment and video lip movement do not match, trigger re-separation. Step S114, Spatiotemporal Feature Joint Analysis: Calculate the HSV histogram difference of consecutive frames, detect the lens boundary by combining optical flow motion vectors, and extract spatiotemporal interest points using a 3D convolutional neural network; Step S115, Semantic Clustering Optimization: Use ResNet-152 to extract frame-level deep features, merge similar scenes through DBSCAN clustering, and establish key frame priority scores; Step S116, Abnormal segment detection: Construct an audio-visual energy matrix. When 10 consecutive frames meet the conditions of brightness <15 lux and audio energy <0.1, it is determined to be a black screen silent segment.
4. The method for intelligent review and violation detection of multimodal content according to claim 1, characterized in that, In step S3, a multimodal resource target dictionary is constructed, and a knowledge graph between entities is formed using relational categories. The specific process is as follows: Step S31: Determine the modal type of the resource for specialization of features of different modalities; Step S32: Obtain resource tags and perform content transformation on unstructured data and features without tags; Step S33: Match the acquired tags against the multimodal resource directory dictionary; Step S34: Generate a resource catalog according to the cataloging rules; Step S35: Match the corresponding data content in the resource directory based on the data tags in the knowledge representation; Step S36: Extract the data content, form a complete expression, and complete the knowledge representation process.
5. The method for intelligent review and violation detection of multimodal content according to claim 1, characterized in that, In step S4, the specific steps for calculating the intermodal confidence difference index by training the shared semantic space through comparative learning are as follows: Step S41: Using a dual-stream Transformer architecture, design modality-specific adapter layers to process text and visual features separately, projecting the original features onto a 256-dimensional common subspace; Step S42: Construct triplet samples, use the cross-modal InfoNCE loss function, and add an inter-modal attention penalty term; Step S43: Integrate a multimodal content detection model that combines text and image semantic consistency; Step S44: Use a multilayer perceptron containing a three-layer fully connected network as the classifier.
6. The method for intelligent review and violation detection of multimodal content according to claim 5, characterized in that, In step S41, the text content T and visual content V of the collected information are acquired, wherein the text is represented as a set T = {w1, w2, ..., w...} consisting of n word groups. n }, convert each word into a vector E i =[e i1 ,e i2 ,...,e ij In the formula, i represents the i-th word, and j represents the dimension of the word vector; for visual content V, it is mapped to a label set V = {t1, t2, ..., tm} consisting of m labels through a pre-trained CNN. m }, which are converted into word vectors using the Word2Vec method; The word vectors are summed, and the mean of each dimension is calculated to obtain a vector with the same dimension as each word, which is used as the semantic vector Em of the text content or image content. The specific calculation formula is as follows: The semantic consistency between text and image is measured by the cosine distance between the text semantic vector Em(T) and the image semantic vector Em(V), as shown in the following formula:
7. The method for intelligent review and violation detection of multimodal content according to claim 5, characterized in that, In step S43, the multimodal content detection model includes text feature extraction, image feature extraction, and text-image semantic consistency feature extraction. The text feature extraction is performed using a Bi-LSTM model. The input is the text word vector W transformed from word vectors, and the output is the set of hidden states H at each step. The result after mean pooling is used as the text feature f. T The formula for its calculation is as follows: f T =AveragePooling(Bi-LSTM([W])); The image feature sports model is fused, and a ResNet network is used to extract image features. The data output from the ResNet network is then input into a fully connected layer to reduce the dimensionality of the image features, resulting in image features f. I Its calculation process can be expressed as: f I =σ(w×Resnet+b); where w and b represent the weights and biases of the fully connected layer, respectively, and σ represents the activation function; The text-image semantic consistency feature extraction process employs five pre-trained CNN models to semantically annotate the image content. Labels with a probability greater than 0.01 from each model's output are selected as image semantic labels. Different numbers and sizes of convolutional kernels are used for each of the five CNN models to abstractly represent the image data. Simultaneously, the image semantic labels are vectorized, mapping words to a vector space. Vector cosine is used to calculate the text-image semantic consistency features. The similarity between the text vector and the label vectors of the five models is calculated, ultimately yielding a semantic consistency feature f containing five elements. s The calculation formula is as follows: f S =[Consistency1,Consistency2,Consistency3,Consistency4,Consistency5].
8. The method for intelligent review and violation detection of multimodal content according to claim 5, characterized in that, In step S44, each piece of information is processed using f. T Representing text features, f I f represents image features S The semantic consistency feature, the fused tweet feature, can be represented by F, with the following formula: A three-layer fully connected network multilayer perceptron is used as the classifier, and the classification result is output through the Softmax function.
9. The method for intelligent review and violation detection of multimodal content according to claim 1, characterized in that, In step S5, the violation determination includes primary alarm, deep verification, and adversarial detection; The primary alarm adopts a dynamic threshold decision mechanism, presets a multimodal sensitive feature library, calculates feature confidence scores in real time, adopts a hierarchical threshold strategy, and automatically adjusts parameters according to time period / scenario to achieve rapid response of single mode; The deep verification includes cross-modal consistency testing and knowledge graph-assisted decision-making; the cross-modal consistency testing is used to detect timestamp deviations in each modality of events, calculate audio-visual semantic similarity, and identify subtitle occlusion behavior. The knowledge graph-assisted decision-making system dynamically allocates confidence levels based on the scenario by constructing evidence chains. The adversarial detection includes voice-changing attack defense, visual adversarial sample cracking, and OCR adversarial protection; the voice-changing attack defense is achieved through liveness detection of voiceprints; the visual adversarial sample cracking enhances the features of sensitive areas through an attention mechanism and combines high-frequency or low-frequency information to identify image disturbances. The OCR anti-problem measures use font perturbation detection to compare the similarity of character structures.
10. A multimodal content intelligent review and violation detection system, comprising a multimodal acquisition layer, a feature extraction engine, a cross-modal association system, and a dynamic decision center, characterized in that: The multimodal acquisition layer includes a distributed crawler module and a streaming processing module; the distributed crawler mode is used to crawl content from multiple channels such as web pages, live broadcasts, and apps; the streaming processing module is used to decompose video into text, audio, and time-frequency data in real time. The feature extraction engine includes a text analysis unit, an audio processing unit, and a visual analysis unit; the text analysis unit integrates syntax tree analysis and metaphor recognition to detect distorted sensitive words; the audio processing unit identifies voice changes, human voices, and background anomalies by constructing a voiceprint fingerprint database; the visual analysis unit is used to capture dynamic sensitive scenes through a spatiotemporal attention model. The cross-modal association system includes a semantic alignment module, a contradiction detector, and a knowledge graph interface; the semantic alignment module is used to establish a shared embedding space for multimodal feature vectors; the contradiction detector is used to identify adversarial behaviors caused by audio-visual asynchrony; and the knowledge graph interface is used to access a violation expression feature library to achieve semantic expansion. The dynamic decision center includes a multi-level filtering pipeline and interpretable output; the multi-level filtering pipeline is used for analysis from initial screening to in-depth evidence chain analysis; the interpretable output is used to generate a visual report containing the location of the violation fragment.
Citation Information
Cited By
Deep learning-fused exploration scene monitoring illegal behavior automatic identification method and system
CN121659076A
Operation instruction conversion method and device, storage medium, electronic device and computer program product
CN121708932A
Block chain and multi-modal learning-based creative traceability and infringement detection method
CN122174216A
Video processing method and device, storage medium and program product
CN122200197A
Incremental learning live broadcast multi-type violation early warning method and system
CN122340285A