Multimedia content AI detection method, device, equipment and storage medium
By preprocessing and feature extraction of multimedia content, combining emotion and behavioral analysis and zero-sample learning, a new type of violation feature template is generated, and cross-modal semantic alignment is used to detect it, the problem of insufficient recognition ability of new types of violation content and implicit violation behavior in the existing technology is solved, and efficient and accurate multimedia content detection is achieved.
Patent Information
- Application Number
- CN202510052926.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-14
AI Technical Summary
Existing multimedia content AI detection technology is difficult to effectively identify new types of illegal content, implicit violations, and comprehensive inspection of multimodal content, resulting in the limitation of the reliability and efficiency of the content audit system.
By preprocessing and feature extraction of multimodal data, global dynamic features, global audio weighted features and behavioral statistical features are generated. Combined with the results of emotion and behavioral analysis, zero-sample learning is used to generate new violation feature templates, and through cross-modal semantic alignment and feature matching, the detection and scoring of new violation content is achieved.
It significantly improves the ability to identify implicit violations in the multimedia content detection system, adaptability to new types of violations, and the efficiency of multimodal semantic fusion, providing an efficient, accurate and widely applicable content review solution.
Smart Images

Figure CN119484890B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of AI detection technology, and in particular to a multimedia content AI detection method, device, equipment and storage medium. Background Art
[0002] With the explosive growth of multimedia content, content supervision and review have become one of the core needs of major platforms. Especially in application scenarios such as social media, live broadcast platforms, and streaming platforms, the spread of illegal content may bring serious social and legal risks. To address this problem, existing technologies mainly rely on rule-based detection methods and artificial intelligence classification models.
[0003] The rule-based method reviews content by constructing predefined violation feature templates, such as keyword matching, bad image detection rules, etc. This method has certain real-time advantages, but its limitations are also very obvious: template rules are difficult to adapt to diverse content and cannot effectively deal with hidden violations (such as tone attacks, complex semantic disguises, etc.). Artificial intelligence technology automatically learns features from data through deep learning models to detect illegal content in images, videos, and audio. This type of technology overcomes some of the shortcomings of rule-based methods and has certain generalization and automation capabilities. However, existing artificial intelligence detection technologies still have the following problems:
[0004] 1. Training methods that rely on labeled data are difficult to deal with new types of illegal content: Traditional deep learning methods rely on large-scale labeled samples, and the detection capabilities of the model are often limited to the scope covered by the training data. When new forms of violations (such as deep fakes and unseen semantic scenes) appear, the model cannot effectively detect them.
[0005] 2. Insufficient ability to identify implicit violations: Implicit violations are often reflected through indirect information such as emotional expressions (such as aggressive tone or expressions) and user interaction behaviors (such as repeated clicks and abnormal stays). Existing technologies fail to fully utilize this information for analysis, resulting in low accuracy in identifying implicit violations.
[0006] 3. Insufficient integration of multimodal content detection: In actual scenarios, illegal content may span multiple modalities such as vision, audio, and text. For example, a video may not be obviously illegal in the picture, but it may express inflammatory information through audio and subtitles. Existing technologies have limited capabilities in multimodal semantic alignment and association analysis, making it difficult to achieve comprehensive and accurate detection.
[0007] The above technical defects restrict the reliability and efficiency of the content review system, especially in terms of real-time processing and generalization capabilities. Summary of the invention
[0008] The purpose of the present invention is to propose a multimedia content AI detection method, device, equipment and storage medium, which can significantly improve the multimedia content detection system's ability to identify hidden illegal behaviors, its adaptability to new illegal content and the efficiency of multimodal semantic fusion, and provide a highly efficient, accurate and widely applicable solution for content review.
[0009] In order to achieve the above object, the present invention provides a multimedia content AI detection method, the method comprising:
[0010] S1. Preprocessing and feature extraction are performed on the input multimodal data, and a multimodal feature representation is generated through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction log; and the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features;
[0011] S2, extracting the emotion and behavior risk features related to implicit violation behaviors from the multimodal feature representation, and generating emotion and behavior analysis results;
[0012] S3. When new multimodal data is input, the results of sentiment and behavior analysis are combined to generate a new violation template through zero-shot learning and match it with the newly input multimodal features to obtain a new violation feature template and matching score;
[0013] S4. Semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then modify the preliminary violation risk score according to the template feature distribution to obtain the final violation risk score;
[0014] S5. The final violation risk score is graded, and a dynamic adjustment mechanism is designed to adapt to different types of violation risk distributions, generate violation levels, and conduct retrospective analysis on high-risk content fragments to locate the specific source of the illegal content and generate a high-risk fragment index. At the same time, the degree of correlation between high-risk fragments and new violation feature templates is calculated through feature contribution. Finally, a structured report is generated by integrating the grading results, high-risk fragment index and feature contribution.
[0015] Furthermore, the preprocessing includes:
[0016] The video is decomposed into a frame sequence by frame decomposition technology ,in Indicates Frame; the audio is divided into multiple fixed time windows , Indicates The user interaction log records the time-related behavior features and constructs a discrete behavior sequence , Indicated in Behavioral characteristics at a point in time;
[0017] Use a custom convolutional neural network to extract Extract visual features and generate frame-by-frame representations , Indicates Deep features of frames; for audio clips , extracting multi-dimensional audio features through spectral analysis ,in Indicates Audio features of time windows; based on user behavior sequence Constructing time series feature matrix , each row of the matrix is the feature of a discrete time point.
[0018] Furthermore, in S1, feature extraction includes:
[0019] Construct time-related features and use the feature change rate between frames to capture the global dynamic relationship of the frame sequence:
[0020] ;
[0021] in, Represents the global dynamic characteristics of the video, For the Frame and The rate of change of features between frames;
[0022] Weighted calculation is introduced to generate weighted features using the emotional frequency band weights of the audio signal:
[0023] ;
[0024] in, represents the global audio weighted feature, For the The weight of the time window, Used to control the sensitivity of weight distribution;
[0025] Calculate the statistical representation of behavioral characteristics and generate user behavior feature vectors :
[0026] ;
[0027] Among them, each Represents the statistical characteristics of behavior in a time series, such as the mean and variance of click frequency or dwell time;
[0028] The global dynamic features of the video , global audio weighted features and behavioral statistical characteristics Combine to generate a unified multimodal feature representation:
[0029] ;
[0030] in, It is a post-multimodal feature representation, which includes the global dynamic features of the video, the emotion-weighted spectral features of the audio, and the statistical representation of the user behavior.
[0031] Furthermore, the S2 specifically includes:
[0032] An emotion fusion regularized decoding model is designed to decode the emotion features and extract multimodal emotion risk features; the emotion features are calculated through global audio weighted features and global correlation features, and are expressed as:
[0033] ;
[0034] in, and Respectively represent the features of video and audio in the emotion decoding embedding space, and is the projection matrix, which is used to capture a specific emotional dimension;
[0035] Fusion and , and introduce the regularization term Optimize the expression of emotional features:
[0036] ;
[0037] in, express Activation function, represents the positive correlation fusion of video and audio features, is a regularization term used to capture the conflict or complementary relationship between video and audio features. represents the final multimodal affective risk signature;
[0038] The behavioral risk adaptive modeling method is used to decode the user behavior feature vector to form a behavioral risk score, which is expressed as:
[0039] ;
[0040] in, is the projection matrix of the behavioral features, is the risk offset, is the regularization bias of abnormal behavior, which is used to amplify the behavioral characteristics that deviate from the normal interaction pattern. represents the behavioral risk score;
[0041] Integrate emotional and behavioral characteristics to generate emotional and behavioral analysis results , and a weighted fusion method is used to ensure the synergy between the two:
[0042] ;
[0043] in, is a multimodal sentiment feature, is the behavioral risk score, and It is the fusion weight, which controls the contribution ratio of emotional and behavioral characteristics to the comprehensive analysis results.
[0044] Furthermore, the novel violation template is generated by zero-sample learning and matched with the newly input multimodal features to obtain the novel violation feature template and the matching score, specifically including:
[0045] A feature generation model is constructed by combining known violation feature templates and sentiment and behavior analysis results to generate potential new violation feature templates; wherein, the feature generation model adopts adaptive feature expansion to simulate potential violation features, and at the same time applies complexity regularization terms to the generated results to ensure the diversity of generated features, which is expressed as:
[0046] ;
[0047] in, It is a known violation feature template, which comes from the labeled data in the database; It is a feature generation model; It is the result of emotion and behavior analysis; It is the basic feature generation matrix, which is used to expand the template features; It combines the results of emotion and behavior analysis The nonlinear activation part of is used to generate features related to actual emotions and behavioral risks. is a complexity regularization term that constrains the deviation of the generated features and ensures the reasonable correlation between the generated feature template and the existing template; It is a new violation feature template generated to detect possible unseen violation features;
[0048] The new violation feature template is matched with the input multimodal feature, and the feature matching score is calculated to quantify the relevance between the input data and the potential violation feature; wherein the matching score adopts the content relevance attention mechanism to capture the interaction between different modal features, which is expressed as:
[0049] ;
[0050] in, is the feature matching score, which quantifies the matching degree between the input data and the generated new violation feature template, is the first Item features; is a feature matching function that calculates the attention-weighted similarity between the features of the novel violation feature template and the input multimodal features; is the matching matrix, which is used to perform multimodal alignment between the new violation feature template and the input features;
[0051] Introduce emotional and behavioral features to dynamically weight the matching score and optimize the matching score's ability to represent potential illegal content, expressed as:
[0052] ;
[0053] in, is the dynamically adjusted matching score, which indicates the final violation risk after emotional behavior correction; and To adjust the parameters, control the feature matching score and combining sentiment and behavioral analysis results Impact on final results; is the distribution variance of the generated features, which serves as a correction term to improve the diversity of the generated templates;
[0054] The S4 specifically includes:
[0055] Designing a cross-modal alignment model , through the semantic alignment attention mechanism, the new violation feature template and multimodal feature representation Fusion into a unified feature space to generate aligned semantic representations , expressed as:
[0056] ;
[0057] in, The function normalizes the weight distribution between the template and the input features; is the alignment weight matrix used to learn the association between template features and multimodal features; It is an asymmetric regularization term that amplifies the influence of salient features in fusion by constraining the absolute difference between the template and the input features; It is the fused aligned semantic representation, which contains the association information between the template and the multimodal input features;
[0058] Using the fused aligned semantic representation and dynamically adjusted matching scores Generate a preliminary breach risk score , expressed as:
[0059] ;
[0060] in, is the Frobenius norm of the aligned features, which is used to measure the complexity of the aligned features; and is the weighting coefficient, which controls the contribution ratio of the matching score to the alignment feature;
[0061] For preliminary scoring Design a correction formula based on feature distribution differences and introduce optimization items for template feature distribution , to ensure the robustness and accuracy of the scoring results, the correction formula is expressed as:
[0062] ;
[0063] in, It is the variance term of the aligned feature distribution, which is used to smooth the inconsistency in the feature fusion process; It is the final violation risk score that integrates template matching, feature complexity, and distribution optimization information.
[0064] Furthermore, the S5 specifically includes:
[0065] Final violation risk score Classify and generate violation levels by designing a dynamic adjustment mechanism to adapt to different types of violation risk distributions , where the classification formula is:
[0066] ;
[0067] in, and are weight parameters, respectively controlling the risk score and semantic feature distribution Impact on classification; It is the alignment semantics representation The distribution variance is used to measure the distribution complexity of semantic features and assist in risk classification; The function normalizes the weighted scores to generate a probability distribution; is the violation level, which is divided into high risk (1), medium risk (2) and low risk (3);
[0068] Conduct retrospective analysis of high-risk content segments to locate the specific source of illegal content:
[0069] By semantically aligning features and templates Interaction relationship, generate an index set of high-risk fragments , expressed as:
[0070] ;
[0071] in, Indicates the alignment semantics Fragment features; Indicates the first Features Used to measure the similarity between template and fragment features; Before selection High-similarity segments generate high-risk segment indexes ;
[0072] Design feature contribution function , calculate the degree of association between the illegal fragment and the template feature, and quantify the role of specific template features in high-risk fragments, expressed as:
[0073] ;
[0074] in, Indicates the first The contribution of the features; the denominator It is the normalization of all template features;
[0075] Construct the final inspection report.
[0076] Furthermore, the final test report includes risk grading results, high-risk fragment markings, template matching interpretation and feature complexity analysis.
[0077] In addition, to achieve the above-mentioned purpose, the present invention also provides a multimedia content AI detection device, characterized in that the multimedia content AI detection device comprises:
[0078] A first acquisition unit is used to preprocess and extract features from input multimodal data, and generate a multimodal feature representation through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction logs; and the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features;
[0079] An analysis unit, used to extract the emotion and behavior risk features related to the implicit violation from the multimodal feature representation, and generate emotion and behavior analysis results;
[0080] The second acquisition unit is used to generate a new violation template by zero-shot learning when new multimodal data is input, combining the emotion and behavior analysis results, and matching it with the newly input multimodal features to obtain a new violation feature template and a matching score;
[0081] A determination unit is used to semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then correct the preliminary violation risk score according to the template feature distribution to obtain the final violation risk score;
[0082] The detection output unit is used to grade the final violation risk score. By designing a dynamic adjustment mechanism to adapt to different types of violation risk distributions, it generates a violation level, performs retrospective analysis on high-risk content fragments, locates the specific source of the illegal content, generates a high-risk fragment index, and calculates the degree of correlation between the high-risk fragment and the new violation feature template through feature contribution. Finally, a structured report is generated by integrating the grading results, high-risk fragment index and feature contribution.
[0083] In addition, to achieve the above-mentioned purpose, the present invention also proposes a multimedia content AI detection device, which includes: a memory, a processor, and a multimedia content AI detection program stored in the memory and executable on the processor, and the multimedia content AI detection is configured to implement the steps of the multimedia content AI detection method described above.
[0084] In addition, to achieve the above-mentioned purpose, the present invention also proposes a storage medium, on which a multimedia content AI detection program is stored. When the multimedia content AI detection program is executed by a processor, the steps of the multimedia content AI detection method described above are implemented.
[0085] The beneficial technical effects of the present invention are at least as follows:
[0086] (1) The present invention constructs a detection mechanism for hidden violations by decoding emotional expressions in multimedia content (such as facial expressions in videos and tones in audio) and user interaction behaviors (such as click frequency and abnormal dwell time when watching). Through emotional feature analysis, it is possible to capture hidden violations such as aggressive tones and misleading expressions that are difficult to identify with traditional methods; through behavioral pattern analysis, it is possible to locate abnormal segments that users are concerned about and assist in judging potential risks. This innovation addresses the problem of difficulty in identifying hidden violations and makes up for the shortcomings of existing technologies.
[0087] (2) The present invention uses zero-sample learning technology to expand the system's ability to detect unseen illegal content. The feature space of existing illegal samples is expanded by generating a model to simulate possible new illegal features; then, through cross-modal feature matching, it is detected whether the content has illegal tendencies consistent with the expanded features. This innovation overcomes the lag of existing methods in detecting new illegal content and significantly improves the system's generalization ability.
[0088] (3) The present invention designs a cross-modal semantic alignment and fusion mechanism, which integrates the features of multi-modal data such as vision, audio, text, and user behavior to achieve a comprehensive analysis of illegal content from a semantic level. Through the semantic alignment model, it is possible to associate illegal features reflected in different modalities to form a unified detection result, thereby improving the coverage and accuracy of detection. This innovation is particularly aimed at the detection difficulties of cross-modal illegal content and improves the comprehensiveness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.
[0090] Figure 1 This is a flow chart of the multimedia content AI detection method of the present invention.
[0091] Figure 2 This is a framework diagram of the multimedia content AI detection device of the present invention. DETAILED DESCRIPTION
[0092] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0093] like Figure 1 As shown, the multimedia content AI detection method provided by the embodiment of the present invention includes the following steps S1-S5:
[0094] S1. Preprocess and extract features from input multimodal data, and generate multimodal feature representation through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction log; the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features.
[0095] Specifically, the first step is to preprocess and extract features from the input multimodal data (video, audio, user interaction logs), and generate high-quality multimodal feature representations through a unified modeling method. , providing the core foundation for subsequent emotion and behavior decoding analysis. The solution design is targeted at the complex multimodal data characteristics of this patent scenario, emphasizing the temporal correlation modeling and semantic aggregation of features.
[0096] Decomposition and preprocessing of input data:
[0097] Receive multimedia content, including video sequences, audio signals, and user behavior interaction logs.
[0098] The video part generates a frame sequence through frame decomposition technology ,in Indicates frame.
[0099] The audio signal is divided into multiple fixed time windows , Indicates An audio clip with a time window.
[0100] User interaction logs record time-related behavior features, such as clicks and dwell time, to construct discrete behavior sequences , Indicated in Behavioral characteristics at a certain point in time.
[0101] Furthermore, the temporal correlation of video frame features is modeled:
[0102] Use a custom convolutional neural network (CNN) to extract Extract visual features and generate frame-by-frame representations , Indicates Deep features of frames.
[0103] Construct time-related features and use the feature change rate between frames to capture the global dynamic relationship of the frame sequence:
[0104] ;
[0105] in, represents the global correlation features of the frame sequence, For the Frame and The rate of change of features between frames.
[0106] Furthermore, the emotion-weighted feature extraction of the audio signal:
[0107] For audio clips , extracting multi-dimensional audio features through spectral analysis ,in Indicates The audio features of a time window.
[0108] Weighted calculation is introduced to generate weighted features using the emotional frequency band weights of the audio signal:
[0109] ;
[0110] in, represents the global audio weighted feature, For the The weight of the time window, Used to control the sensitivity of the weight distribution.
[0111] Furthermore, statistical extraction of user behavior characteristics:
[0112] According to user behavior sequence Constructing time series feature matrix , each row of the matrix is the feature of a discrete time point.
[0113] Calculate the statistical representation of the behavior characteristics and generate the user behavior feature vector:
[0114] ;
[0115] Among them, each Represents the statistical characteristics of behavior in a time series, such as the mean and variance of click frequency or dwell time.
[0116] Furthermore, the global features of the video , audio weighted features and user behavior characteristics Combine to generate a unified multimodal feature representation:
[0117] ;
[0118] in, It is the input feature of the subsequent steps, including the global dynamic features of the video, the emotion-weighted spectral features of the audio, and the statistical representation of user behavior.
[0119] S2. Extract the emotion and behavior risk features related to implicit violation behaviors from the multimodal feature representation and generate emotion and behavior analysis results.
[0120] Specifically, receive the input features from step 1 .in, represents the global dynamic characteristics of the video frame, is the weighted spectral feature of the audio, It is the behavioral characteristics of user interaction.
[0121] Furthermore, the emotional features are decoded and an emotional fusion regularized decoding model is designed to extract multimodal emotional risk features:
[0122] Compute sentiment embeddings for video and audio via feature projection:
[0123] ;
[0124] in, and Respectively represent the features of video and audio in the emotion decoding embedding space, and is the projection matrix, which is used to capture a specific emotion dimension.
[0125] Fusion and , and introduce regularization terms to optimize the expression of emotional features:
[0126] ;
[0127] in, represents the positive correlation fusion of video and audio features, Used to capture the conflicting or complementary relationship between video and audio features. Represents the final multimodal sentiment risk feature.
[0128] Furthermore, the behavioral characteristics Perform risk decoding and use behavioral risk adaptive modeling to generate behavioral risk scores:
[0129] User behavior characteristics , the risk score is calculated by projection and anomaly deviation modeling:
[0130] ;
[0131] in, is the projection matrix of the behavioral features, is the risk offset, is the regularization bias of abnormal behavior, which is used to amplify the behavioral characteristics that deviate from the normal interaction pattern. Represents the behavioral risk score.
[0132] Furthermore, the emotional characteristics and behavioral characteristics are integrated to generate emotional and behavioral analysis results. , and a weighted fusion method is used to ensure the synergy between the two:
[0133] ;
[0134] in, is a multimodal sentiment feature, is the behavioral risk score, and It is the fusion weight, which controls the contribution ratio of emotional and behavioral characteristics to the comprehensive analysis results.
[0135] Understandable, output comprehensive analysis results , which contains comprehensive information on emotional characteristics and behavioral risks, providing high-quality input for subsequent zero-sample feature generation and matching. Through this result, the system can more effectively capture hidden violations and potential risks, laying the foundation for subsequent detection modules.
[0136] S3. When new multimodal data is input, the results of sentiment and behavior analysis are combined to generate a new violation template through zero-shot learning and match it with the newly input multimodal features to obtain a new violation feature template and matching score.
[0137] Specifically, the input received is the comprehensive emotion and behavior analysis results output from step 2 And the multimodal feature representation of step 1 . are the emotional and behavioral risk features generated after multimodal analysis, Contains global dynamic features of the video , audio weighted features and user behavior characteristics The goal of this step is to generate possible new violation feature templates , and matches it with the input data to quantify the potential violation risk.
[0138] Furthermore, we construct a feature generation model , combined with known violation feature templates and analysis results , generate potential new violation feature templates The model design uses adaptive feature expansion to simulate potential violation features, while applying complexity regularization terms to the generated results to ensure the diversity of generated features:
[0139] ;
[0140] in, It is a known violation feature template, which comes from the labeled data in the database. It is the basic feature generation matrix, which is used to expand the template features. The combined analysis results The nonlinear activation part of is used to generate features related to actual emotions and behavioral risks. is a complexity regularization term that constrains the deviation of the generated features and ensures the reasonable correlation between the generated feature template and the existing template. It is a new violation feature template generated to detect possible unseen violation features.
[0141] Furthermore, the generated new violation feature template Multimodal features of the input Perform matching and calculate feature matching scores , to quantify the relevance of input data to potential violation features. The matching function design introduces a content-related attention mechanism to capture the interactive relationship between different modal features:
[0142] ;
[0143] in, is the first Item features. is a feature matching function that calculates the attention-weighted similarity between template features and input multimodal features. is the matching matrix used to perform multimodal alignment between template and input features. is the feature matching score, which quantifies the matching degree between the input data and the generated template. Dynamic adjustments are made to introduce weighted corrections of sentiment and behavioral characteristics to further optimize the matching score's ability to characterize potential illegal content:
[0144] ;
[0145] in, is the dynamically adjusted matching score, which indicates the final violation risk after emotional behavior correction; and To adjust the parameters, control the feature matching score and combining sentiment and behavioral analysis results Impact on final results; is the distribution variance of the generated features, which serves as a correction term to improve the diversity of the generated templates.
[0146] Understandable, output the final result, including the matching score And the generated new violation feature template :
[0147] It provides a potential violation risk score for the input data, providing a basis for subsequent cross-modal semantic alignment.
[0148] As a new type of feature template generated, it expands the existing feature space and provides an important basis for detecting unseen illegal content.
[0149] This step solves the problem of insufficient adaptability of the existing model to unprecedented illegal content by generating and matching new illegal feature templates, further improving the generalization and adaptability of the detection system, and laying a solid foundation for the multimodal fusion and final violation judgment in subsequent steps.
[0150] S4. Semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then correct the preliminary violation risk score according to the template feature distribution to obtain the final violation risk score.
[0151] Specifically, through the semantic alignment attention mechanism, the template features and multimodal features Fusion into a unified feature space to generate aligned semantic representations .
[0152] Furthermore, we design a cross-modal alignment model , an asymmetric regularization term of content relevance is introduced to emphasize the contribution of important features:
[0153] ;
[0154] in, The function normalizes the weight distribution between the template and the input features. is the alignment weight matrix used to learn the association between template features and multimodal features. It is an asymmetric regularization term that amplifies the influence of salient features in fusion by constraining the absolute difference between the template and the input features. It is the fused aligned semantic representation, which contains the association information between the template and the multimodal input features.
[0155] Furthermore, using the alignment results and matching score Generate a preliminary breach risk score , highlighting the importance of different inputs through a weighted mechanism. The weighted fusion model is defined as:
[0156] ;
[0157] in, is the Frobenius norm of the aligned features, which is used to measure the complexity of the aligned features; and is the weighting coefficient that controls the contribution of the matching score to the alignment feature
[0158] in, is the matching score calculated in step S3, reflecting the preliminary possibility of violation. is the Frobenius norm of the aligned features, which is used to measure the complexity of the aligned features. and is a weighting factor that controls the contribution ratio of the matching score to the alignment feature. is a preliminary violation risk score that combines template matching and alignment features.
[0159] For preliminary scoring , introduce the optimization term of template feature distribution to ensure the robustness and accuracy of the scoring results. Design a correction formula based on feature distribution differences:
[0160] ;
[0161] in, It is the variance term of the aligned feature distribution, which is used to smooth the inconsistency in the feature fusion process. It is the final violation risk score, which combines template matching, feature complexity and distribution optimization information. Output the final violation risk score , as the core result of multimodal fusion detection. It is directly used for subsequent risk classification and report generation, and provides accurate illegal content labeling for the detection system. This step solves the semantic disconnection problem between multimodal features through cross-modal alignment attention mechanism, asymmetric regularization term and distribution optimization correction, and improves the detection system's ability to identify potential illegal content.
[0162] S5. The final violation risk score is graded, and a dynamic adjustment mechanism is designed to adapt to different types of violation risk distributions, generate violation levels, and conduct retrospective analysis on high-risk content fragments to locate the specific source of the illegal content and generate a high-risk fragment index. At the same time, the degree of correlation between high-risk fragments and new violation feature templates is calculated through feature contribution. Finally, a structured report is generated by integrating the grading results, high-risk fragment index and feature contribution.
[0163] Specifically, risk scoring Classify and generate violation levels by designing a dynamic adjustment mechanism to adapt to different types of violation risk distributions The grading formula is:
[0164] ;
[0165] in, and are weight parameters, respectively controlling the risk score and semantic feature distribution Impact on classification. It is the alignment semantics representation The distribution variance of is used to measure the distribution complexity of semantic features and assist in risk classification. The function normalizes the weighted scores to generate a probability distribution. is the violation level, which is divided into high risk (1), medium risk (2) and low risk (3).
[0166] Furthermore, we conduct retrospective analysis of high-risk content segments to locate the specific source of illegal content. and templates Interaction relationship, generate an index set of high-risk fragments :
[0167] ;
[0168] in, Indicates the alignment semantics The features of each segment (such as corresponding video frames, audio segments, etc.) Indicates the first Features. Used to measure the similarity between template and fragment features. Before selection High-similarity segments generate high-risk segment indexes .
[0169] Furthermore, feature contribution analysis is introduced to generate reports and calculate the degree of correlation between the illegal fragments and the template features. Design feature contribution function , used to quantify the role of specific template features in high-risk fragments:
[0170] ;
[0171] in, Representing Template Features The contribution of is the normalization of all template features. This formula emphasizes the relative contribution of individual template features in high-risk fragments and supports report generation.
[0172] Furthermore, the final test report is constructed by integrating the grading results , High-risk segment index and feature contribution Generate structured reports. Report Generation Model Includes the following:
[0173] Risk grading result: indicates the overall risk level of the content (High, Medium, Low risk).
[0174] High-risk segment marker: contains the index of high-risk segments (such as video frame number, audio time period).
[0175] Template matching explanation: Provides illegal content and template features Specific related information, combined with contribution Describe the importance of the feature.
[0176] Feature Complexity Analysis: Analyzing Alignment Features The distribution characteristics (such as ), illustrating the relationship between content complexity and violation risk.
[0177] Output the final detection report, providing clear and actionable results. The report supports multiple formats (such as structured tables and charts) to enable users to quickly identify and review risky content. This step not only helps content reviewers accurately locate problems, but also provides feedback support for potential violation template expansion. This step solves the problems of interpretability and operability of detection results through dynamic risk grading, retrospective analysis, and structured report generation.
[0178] like Figure 2 As shown, the multimedia content AI detection device proposed in the embodiment of the present invention includes:
[0179] The first acquisition unit 101 is used to preprocess and extract features from the input multimodal data, and generate a multimodal feature representation through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction logs; and the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features;
[0180] An analysis unit 102, used to extract emotional and behavioral risk features related to implicit illegal behaviors from the multimodal feature representation, and generate emotional and behavioral analysis results;
[0181] The second acquisition unit 103 is used to generate a new violation template by zero-sample learning when new multimodal data is input, combining the emotion and behavior analysis results, and matching it with the newly input multimodal features to obtain a new violation feature template and a matching score;
[0182] The determination unit 104 is used to semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then correct the preliminary violation risk score according to the template feature distribution to obtain a final violation risk score;
[0183] The detection output unit 105 is used to grade the final violation risk score, adapt to different types of violation risk distributions by designing a dynamic adjustment mechanism, generate a violation level, and perform a retrospective analysis on high-risk content segments to locate the specific source of the illegal content and generate a high-risk segment index. At the same time, the correlation between the high-risk segment and the new violation feature template is calculated through the feature contribution, and finally a structured report is generated by integrating the grading results, the high-risk segment index and the feature contribution.
[0184] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of the present invention. In practical applications, technicians in this field can select part or all of them according to actual needs to achieve the purpose of the present embodiment, and no limitation is made here.
[0185] In addition, for technical details that are not described in detail in this embodiment, reference can be made to the parameter operation method provided in any embodiment of the present invention, and will not be repeated here.
[0186] Other embodiments or specific implementations of the multimedia content AI detection device of the present invention may refer to the above-mentioned method embodiments, which will not be described in detail here.
[0187] In addition, an embodiment of the present invention further proposes a storage medium, on which a multimedia content AI detection program is stored. When the multimedia content AI detection program is executed by a processor, the steps of the multimedia content AI detection method described above are implemented.
[0188] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.
[0189] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0191] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A multimedia content AI detection method, characterized in that: The method comprises: S1. Preprocessing and feature extraction are performed on the input multimodal data, and a multimodal feature representation is generated through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction log; and the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features; S2, extracting the emotion and behavior risk features related to implicit violation behaviors from the multimodal feature representation, and generating emotion and behavior analysis results; S3. When new multimodal data is input, the results of sentiment and behavior analysis are combined to generate a new violation template through zero-shot learning and match it with the newly input multimodal features to obtain a new violation feature template and matching score; S4. Semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then modify the preliminary violation risk score according to the template feature distribution to obtain the final violation risk score; S5. The final violation risk score is graded, and a dynamic adjustment mechanism is designed to adapt to different types of violation risk distributions, generate violation levels, and conduct retrospective analysis on high-risk content fragments to locate the specific source of the illegal content and generate a high-risk fragment index. At the same time, the degree of correlation between high-risk fragments and new violation feature templates is calculated through feature contribution. Finally, a structured report is generated by integrating the grading results, high-risk fragment index and feature contribution.
2. The multimedia content AI detection method according to claim 1, characterized in that: The pre-processing comprises: The video is decomposed into a frame sequence by frame decomposition technology ,in Indicates Frame; the audio is divided into multiple fixed time windows , Indicates The user interaction log records the time-related behavior features and constructs a discrete behavior sequence , Indicated in Behavioral characteristics at a point in time; Use a custom convolutional neural network to extract Extract visual features and generate frame-by-frame representations , Indicates Deep features of frames; for audio clips , extracting multi-dimensional audio features through spectral analysis ,in Indicates Audio features of time windows; based on discrete behavior sequences Constructing time series feature matrix , each row of the matrix is the feature of a discrete time point.
3. The multimedia content AI detection method according to claim 2, characterized in that: In S1, feature extraction includes: Construct time-related features and use the feature change rate between frames to capture the global dynamic relationship of the frame sequence: ; in, Represents the global dynamic characteristics of the video, For the frame and frame The feature change rate between them, n represents the total number of frames of i frame, i represents i frame; Introduce weighted calculation and use the emotional frequency band weights of the audio signal to generate weighted features: ; in, represents the global audio weighted feature, For the The weight of the time window, Used to control the sensitivity of the weight distribution, represents the j-th frame, j is the j-th frame, and m is the total number of j-th frames; Calculate the statistical representation of behavioral characteristics and generate user behavior feature vectors : ; Among them, each Represents the statistical characteristics of behavior in time series, Represents the first behavioral statistical feature, Represents the second behavioral statistical feature, represents the statistical characteristics of the i-th behavior; The global dynamic features of the video , global audio weighted features and behavioral statistical characteristics Combine to generate a unified multimodal feature representation: ; in, It is a multimodal feature representation.
4. The multimedia content AI detection method according to claim 3, characterized in that: The S2 specifically includes: An emotion fusion regularized decoding model is designed to decode the emotion features and extract multimodal emotion risk features; the emotion features are calculated through global audio weighted features and global correlation features, and are expressed as: ; in, and Respectively represent the features of video and audio in the emotion decoding embedding space, and is the projection matrix, which is used to capture a specific emotional dimension; Fusion and , and introduce the regularization term Optimize the expression of emotional features: ; in, express Activation function, represents the positive correlation fusion of video and audio features, is a regularization term used to capture the conflict or complementary relationship between video and audio features. represents the final multimodal affective risk signature; The behavioral risk adaptive modeling method is used to decode the user behavior feature vector to form a behavioral risk score, which is expressed as: ; in, is the projection matrix of the behavioral features, is the risk offset, is the regularization bias of abnormal behavior, which is used to amplify the behavioral characteristics that deviate from the normal interaction pattern. represents the behavioral risk score, k is the total time series, is the statistical feature of the i-th behavior, for abnormal behavior; Integrate emotional and behavioral characteristics to generate emotional and behavioral analysis results , and a weighted fusion method is used to ensure the synergy between the two: ; in, is a multimodal affective risk feature, is the behavioral risk score, and It is the fusion weight, which controls the contribution ratio of emotional and behavioral characteristics to the comprehensive analysis results.
5. The multimedia content AI detection method according to claim 4, characterized in that: The method generates a new violation template through zero-sample learning and matches it with the newly input multimodal features to obtain a new violation feature template and a matching score, specifically including: A feature generation model is constructed by combining known violation feature templates and sentiment and behavior analysis results to generate potential new violation feature templates; wherein, the feature generation model adopts adaptive feature expansion to simulate potential violation features, and at the same time applies complexity regularization terms to the generated results to ensure the diversity of generated features, which is expressed as: ; in, It is a known violation feature template, which comes from the labeled data in the database; It is a feature generation model; It is the result of emotion and behavior analysis; It is the basic feature generation matrix, which is used to expand the template features; It combines the results of emotion and behavior analysis The nonlinear activation part of is used to generate features related to actual emotional and behavioral risks; is a complexity regularization term that constrains the deviation of the generated features and ensures the reasonable correlation between the generated feature template and the existing template; is a new violation feature template generated to detect possible unseen violation features. is the activation function, For the results of emotion and behavior analysis, Representation and sentiment and behavior analysis results The associated weight matrix is used to Transformed into activation information for generating new violation features; The new violation feature template is matched with the input multimodal feature, and the feature matching score is calculated to quantify the relevance between the input data and the potential violation feature; wherein the matching score adopts the content relevance attention mechanism to capture the interaction between different modal features, which is expressed as: ; in, is the feature matching score, which quantifies the matching degree between the input data and the generated new violation feature template, is the first Item features; is a feature matching function that calculates the attention-weighted similarity between the features of the novel violation feature template and the input multimodal features; is the matching matrix used to perform multimodal alignment between the new violation feature template and the input features, represents multimodal feature representation, express The transposed matrix of Introduce emotional and behavioral features to dynamically weight the matching score and optimize the matching score's ability to represent potential illegal content, expressed as: ; in, is the dynamically adjusted matching score, which indicates the final violation risk after emotional behavior correction; and To adjust the parameters, the feature matching scores are controlled respectively; and combining sentiment and behavioral analysis results Impact on final results; is the distribution variance of the generated features, which serves as a correction term to improve the diversity of the generated templates; The S4 specifically includes: Designing a cross-modal alignment model , through the semantic alignment attention mechanism, the new violation feature template and multimodal feature representation Fusion into a unified feature space to generate aligned semantic representations , expressed as: ; in, The function normalizes the weight distribution between the template and the input features; is the alignment weight matrix used to learn the association between template features and multimodal features; It is an asymmetric regularization term that amplifies the influence of salient features in fusion by constraining the absolute difference between the template and the input features; It is the fused aligned semantic representation, which contains the association information between the template and the multimodal input features; Using the fused aligned semantic representation and dynamically adjusted matching scores Generate a preliminary breach risk score , expressed as: ; in, is the Frobenius norm of the aligned features, which is used to measure the complexity of the aligned features; and is the weighting coefficient, which controls the contribution ratio of the matching score to the alignment feature; For preliminary scoring Design a correction formula based on feature distribution differences and introduce optimization items for template feature distribution , to ensure the robustness and accuracy of the scoring results, the correction formula is expressed as: ; in, It is the variance term of the aligned feature distribution, which is used to smooth the inconsistency in the feature fusion process; It is the final violation risk score that integrates template matching, feature complexity, and distribution optimization information.
6. The multimedia content AI detection method according to claim 5, characterized in that: The S5 specifically includes: Final violation risk score Classify and generate violation levels by designing a dynamic adjustment mechanism to adapt to different types of violation risk distributions , where the classification formula is: ; in, and are weight parameters, respectively controlling the risk score and semantic feature distribution Impact on classification; It is the alignment semantics representation The distribution variance is used to measure the distribution complexity of semantic features and assist in risk classification; The function normalizes the weighted scores to generate a probability distribution; is the violation level, which is divided into high risk, medium risk and low risk; Conduct retrospective analysis of high-risk content segments to locate the specific source of illegal content: By aligning semantics and templates Interaction relationship, generate an index set of high-risk fragments , expressed as: ; in, Indicates the alignment semantics Fragment features; Indicates the first Features Used to measure the similarity between template and fragment features; Before selection High-similarity segments generate high-risk segment indexes ; Design feature contribution function , calculate the degree of association between the illegal fragment and the template feature, and quantify the role of specific template features in high-risk fragments, expressed as: ; in, Indicates the first The contribution of the features; the denominator It is the normalization of all template features; Construct the final inspection report.
7. The multimedia content AI detection method according to claim 6, characterized in that: The final test report includes risk grading results, high-risk fragment markings, template matching interpretation and feature complexity analysis.
8. A multimedia content AI detection device, characterized in that: The multimedia content AI detection device comprises: A first acquisition unit is used to preprocess and extract features from input multimodal data, and generate a multimodal feature representation through a unified modeling method; wherein the multimodal data includes video, audio, and user interaction logs; and the multimodal feature representation includes global dynamic features, global audio weighted features, and behavioral statistical features; An analysis unit, used to extract the emotion and behavior risk features related to the implicit violation from the multimodal feature representation, and generate emotion and behavior analysis results; The second acquisition unit is used to generate a new violation template by zero-shot learning when new multimodal data is input, combining the emotion and behavior analysis results, and matching it with the newly input multimodal features to obtain a new violation feature template and a matching score; A determination unit is used to semantically align the new violation feature template with the multimodal feature, generate a preliminary violation risk score based on the alignment result and the matching score, and then correct the preliminary violation risk score according to the template feature distribution to obtain the final violation risk score; The detection output unit is used to grade the final violation risk score. By designing a dynamic adjustment mechanism to adapt to different types of violation risk distributions, it generates a violation level, performs retrospective analysis on high-risk content fragments, locates the specific source of the illegal content, generates a high-risk fragment index, and calculates the degree of correlation between the high-risk fragment and the new violation feature template through feature contribution. Finally, a structured report is generated by integrating the grading results, high-risk fragment index and feature contribution.
9. Multimedia content AI detection device, characterized in that: The device comprises: a memory, a processor, and a multimedia content AI detection program stored in the memory and executable on the processor, wherein the multimedia content AI detection program is configured to implement the steps of the multimedia content AI detection method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium stores a multimedia content AI detection program, and when the multimedia content AI detection program is executed by the processor, the steps of the multimedia content AI detection method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for improving voice call quality inspection effect
CN112885376A
Financial live broadcast violation detection method, device and equipment and readable storage medium
CN113038153A