User generated content detection method and device, equipment and medium
The user-generated content detection model trained by the cross-modal-adversarial domain invariant learning mechanism solves the problem of susceptibility to artifact interference in single-modal analysis, and achieves higher detection accuracy and generalization ability, making it suitable for user-generated content detection in the financial and healthcare fields.
Patent Information
- Application Number
- CN202511059157.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, user-generated content detection methods rely on unimodal feature analysis, which is susceptible to interference from specific artifacts in the dataset, resulting in insufficient cross-scene generalization ability and low detection accuracy.
A cross-modal-adversarial domain invariant learning mechanism is used to train a user-generated content detection model. Visual and audio inputs are extracted using a feature extraction module, encoded using a cross-modal decoupling module, fused using a multi-head cross-attention mechanism, and detected using a classifier.
It enhances the model's ability to generalize to unknown distributions, improves the accuracy of anomaly detection for user-generated content, and can adaptively aggregate multimodal features to highlight consistent anomaly patterns across modalities.
Smart Images

Figure CN120980274A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of finance, healthcare and artificial intelligence, and in particular to a method, apparatus, device and medium for detecting user-generated content. Background Technology
[0002] With the rapid development of deepfake technology, the proliferation of deepfake content in user-generated content (UGC) has become a major challenge in the field of digital media security. This type of fake content may be used to maliciously spread false information or to forge identities.
[0003] For example, in the financial sector, the forgery of user-generated content may bypass the liveness detection system of financial institutions, creating security risks; in the medical and health field, the forgery of user-generated content may lead to the undesirable spread of fake science videos.
[0004] In existing technologies, the detection of user-generated content mainly relies on single-modal (such as visual) feature analysis. However, such methods are easily affected by specific artifacts in the dataset (such as compression marks or lighting conditions), resulting in insufficient cross-scene generalization ability. Summary of the Invention
[0005] In view of the above, it is necessary to provide a method, apparatus, device and medium for detecting user-generated content, in order to solve the problem of low accuracy of user-generated content detection results.
[0006] A method for detecting user-generated content, the method comprising:
[0007] In response to the detection command for the target user-generated video, the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism is invoked;
[0008] The feature extraction module in the user-generated content detection model is used to extract features from the target user-generated video to obtain the target visual input and the target audio input.
[0009] The target visual input and the target audio input are encoded using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features;
[0010] The target decoupling features are fused based on a multi-head cross-attention mechanism to obtain target fused features;
[0011] The target fusion features are detected using the classifier in the user-generated content detection model to obtain the detection results.
[0012] A user-generated content detection device, the user-generated content detection device comprising:
[0013] The retrieval unit is used to retrieve the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism in response to the detection command for the target user-generated video;
[0014] The extraction unit is used to extract features from the target user-generated video using the feature extraction module in the user-generated content detection model to obtain the target visual input and the target audio input.
[0015] The encoding unit is used to encode the target visual input and the target audio input using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features;
[0016] The fusion unit is used to fuse the target decoupled features based on a multi-head cross-attention mechanism to obtain the target fused features;
[0017] The detection unit is used to detect the target fusion features using the classifier in the user-generated content detection model to obtain the detection result.
[0018] A computer device, the computer device comprising:
[0019] Memory, storing at least one instruction; and
[0020] The processor executes instructions stored in the memory to implement the user-generated content detection method.
[0021] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the user-generated content detection method.
[0022] As can be seen from the above technical solutions, this invention can train a user-generated content detection model based on a cross-modal-adversarial domain invariant learning mechanism, ensuring cross-modal feature alignment, and eliminating dataset bias based on adversarial domain invariant learning, thereby enhancing the model's generalization ability to unknown distributions. It utilizes a feature extraction module to extract multimodal features from visual and audio inputs, providing basic features for subsequent cross-modal decoupling. The cross-modal decoupling module encodes visual and audio inputs to obtain decoupling features, enabling feature splitting, avoiding interference from irrelevant features, and providing a purer data foundation for subsequent detection. Based on a multi-head cross-attention mechanism, the decoupling features are fused, and a classifier is used to perform anomaly detection on the fused features. This adaptively aggregates multimodal features, highlighting consistent anomaly patterns across modalities, improving the ability to identify anomalous content, and thus increasing the accuracy of anomaly detection for user-generated content. Attached Figure Description
[0023] Figure 1 This is a flowchart of a preferred embodiment of the user-generated content detection method of the present invention.
[0024] Figure 2 This is a functional block diagram of a preferred embodiment of the user-generated content detection device of the present invention.
[0025] Figure 3 This is a schematic diagram of the structure of a computer device that implements the user-generated content detection method of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0027] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the user-generated content detection method of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0028] The user-generated content detection method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0029] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.
[0030] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0031] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0032] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0033] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0034] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0035] S10, in response to the detection command for the target user-generated video, invokes the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism.
[0036] In this embodiment, the detection command can be triggered by relevant maintenance personnel according to actual anomaly detection needs.
[0037] In this embodiment, the target user-generated video can be a user authentication video in the financial field or a medical knowledge popularization video in the medical and health field.
[0038] In this embodiment, before retrieving the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism, the method further includes:
[0039] Constructing the reconstruction loss: in, Denotes the reconstruction loss; x v This represents the visual input extracted by the feature extraction module; Indicates with x v Correspondingly, the reconstructed visual features obtained after reconstruction by the decoder in the cross-modal decoupling module. D vThis represents the visual decoder function. z represents the vector concatenation operation. s The z represents the shared feature output by the shared encoder in the cross-modal decoupling module. v This represents the visual features output by the visual encoder in the cross-modal decoupling module; x a This represents the audio input extracted by the feature extraction module; Indicates with x a Correspondingly, the reconstructed audio features obtained after reconstruction by the decoder in the cross-modal decoupling module. D a z represents the audio decoder function. a This represents the audio features output by the audio encoder in the cross-modal decoupling module;
[0040] Constructing the contrast alignment loss: in, The comparison alignment loss represents the comparison alignment loss; s represents the cosine similarity function. This represents the shared features corresponding to the visual input; This represents the shared feature corresponding to the audio input; τ represents the temperature coefficient; This represents the audio shared features of the k-th negative sample; K represents the number of negative samples.
[0041] Based on the gradient inversion layer and domain discriminator in the cross-modal decoupling module, an adversarial domain-invariant loss is constructed: in, This represents the adversarial domain invariant loss; D represents the operation of finding the expected value; dom Representation domain discriminator function; z s This represents the shared features in the decoupling features output by the cross-modal decoupling module;
[0042] The user-generated content detection model is trained based on the reconstruction loss, the contrast alignment loss, and the adversarial domain invariance loss.
[0043] The shared encoder employs a parameter isolation design, ensuring its independence from the modal encoder through orthogonal constraints.
[0044] Among them, the domain discriminator combined with the modality discrimination loss can simultaneously optimize feature decoupling and modality alignment.
[0045] The training of the temporal discriminator attempts to distinguish the feature source dataset, while the shared encoder optimizes in reverse through a gradient inversion layer to confuse the domain discriminator, ultimately enabling the shared features to ignore the unique distribution of the dataset.
[0046] In the above embodiments, the reconstruction loss and the contrast alignment loss can enhance the model's ability to distinguish similar but different categories, making the differentiation of fake data more effective; adversarial training can remove dataset bias and enhance the model's ability to generalize to unknown distributions.
[0047] S11, the feature extraction module in the user-generated content detection model is used to extract features from the target user-generated video to obtain the target visual input and the target audio input.
[0048] In this embodiment, before extracting features from the target user-generated video using the feature extraction module in the user-generated content detection model, the method further includes:
[0049] The target user-generated video is sampled and processed to obtain a visual frame sequence and synchronized MFCC (Mel-Frequency Cepstral Coefficients) audio features.
[0050] The visual frame sequence is subjected to frame alignment and normalization processing.
[0051] The MFCC audio features are subjected to audio resampling processing.
[0052] The above embodiments enable the standardization of the original video, providing input data in a unified format for subsequent feature extraction.
[0053] In this embodiment, the step of using the feature extraction module in the user-generated content detection model to extract features from the target user-generated video to obtain the target visual input and target audio input includes:
[0054] Based on the ViT (Vision Transformer) architecture, visual features are extracted from the target user-generated video to obtain the target visual input;
[0055] An unsupervised speech pre-training model is used to extract acoustic features from the target user-generated video to obtain the target audio input.
[0056] For example, the unsupervised speech pre-training model can be a Wav2Vec 2.0 model.
[0057] In the above embodiments, the ViT architecture is used as a visual feature extractor and the unsupervised speech pre-trained model is used as an audio feature extractor, which can extract multimodal features and provide a data foundation for subsequent feature decoupling.
[0058] S12, the target visual input and the target audio input are encoded using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features.
[0059] In this embodiment, the cross-modal decoupling module includes a shared encoder, a visual encoder, and an audio encoder and decoder.
[0060] The shared encoder is used to output identity-related features.
[0061] The visual encoder is used to capture traces of forgery.
[0062] The audio encoder is used to extract audio authenticity features.
[0063] The decoder employs a symmetric encoder structure to constrain feature integrity through reconstruction loss. The decoder includes a visual decoder and an audio decoder.
[0064] In this embodiment, the step of encoding the target visual input and the target audio input using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features includes:
[0065] The target visual input and the target audio input are encoded using the shared encoder to obtain target shared features;
[0066] The visual encoder is used to encode the target visual input to obtain the target visual features;
[0067] The target audio input is encoded using the audio encoder to obtain target audio features;
[0068] The target shared features, the target visual features, and the target audio features are determined as the target decoupling features;
[0069] The shared encoder includes 4 Transformer layers, the visual encoder includes 2 convolutional neural network layers and 1 Transformer layer, and the audio encoder includes 2 convolutional neural network layers and 1 Transformer layer.
[0070] The cross-modal decoupling module processes visual and audio inputs through a dual-branch VAE (Variational Autoencoder) structure.
[0071] Through the above embodiments, the input features can be decomposed into shared features and modality-specific features, avoiding interference from irrelevant features and providing purer features related to forgery for subsequent detection.
[0072] S13, the target decoupling features are fused based on the multi-head cross-attention mechanism to obtain the target fused features.
[0073] In this embodiment, the fusion of the target decoupling features based on the multi-head cross-attention mechanism to obtain the target fused features includes:
[0074] The target fusion features are calculated using the following formula:
[0075] Where Q = W q z v K = W k z a V = W v z a ;
[0076] Where h represents the target fusion feature; Q represents the query vector; K represents the key vector; V represents the value vector; d represents the feature dimension; W q W k W v This represents the projection matrix.
[0077] For example, an 8-head attention mechanism can be used to generate fused features based on a multi-head cross-attention mechanism, and dynamic feature aggregation can be achieved by calculating the correlation between visual and audio features.
[0078] Among them, the cross-attention mechanism introduces learnable relative position encoding, which can effectively capture the lip-speech temporal relationship.
[0079] The target fusion feature can focus on forgery traces that are inconsistent between visual and audio (such as lip-syncing and speech asynchrony).
[0080] Through the above embodiments, multimodal features can be adaptively aggregated to highlight consistent forgery patterns across modalities and improve the ability to identify forged content.
[0081] S14, the classifier in the user-generated content detection model is used to detect the target fusion features to obtain the detection result.
[0082] In this embodiment, the step of using the classifier in the user-generated content detection model to detect the target fusion features and obtaining the detection result includes:
[0083] The target fusion features are input into a deep residual network to obtain the anomaly probability;
[0084] Obtain the probability threshold;
[0085] When the anomaly probability is greater than the probability threshold, the detection result is determined to be an anomaly in the target user's generated video; or
[0086] When the anomaly probability is less than or equal to the probability threshold, the detection result is determined to be that the target user's generated video has no anomalies.
[0087] For example, the deep residual network can be a lightweight ResNet-18 (Residual Network 18) network.
[0088] Specifically, the deep residual network can be used to classify fused features and output the forgery probability.
[0089] The probability threshold can be selected as the optimal value based on the experimental results.
[0090] The above embodiments enable the detection decision of whether the input content is a deepfake.
[0091] This embodiment improves the robustness and accuracy of detection by integrating a multimodal learning framework for visual and audio signals. It employs cross-modal decoupling to decompose input features into shared features (such as speaker identity) and modality-specific features (such as visual and audio features), utilizes a variational autoencoder framework for feature separation, and ensures cross-modal feature alignment through contrastive loss. It combines an adversarial domain-invariant representation learning mechanism to eliminate dataset bias and enhance the model's generalization ability to unknown distributions. A cross-attention feature fusion mechanism aggregates multimodal information, highlighting forgery-related features, and a lightweight classifier performs the detection. This embodiment uses a Transformer-based encoder (such as ViT and Wav2Vec 2.0) and a multi-head attention mechanism, significantly outperforming traditional single-modal or simple fusion methods, and is suitable for deep forgery detection in complex user-generated content.
[0092] For example, in the financial field, consider a short video (5 seconds, 25 frames) suspected of being a forged identity verification profile, where the person's face has been replaced but the original audio is retained. Each frame of the video is used to extract 768-dimensional features using ViT, resulting in a 25×768 matrix x. v MFCC audio input extracts 40-dimensional features, which are then encoded into a 25×512 matrix x using Wav2Vec 2.0. a Furthermore, through feature decoupling, the shared feature z is achieved. s Capture the real speaker's identity features (consistent with the audio), visual features z v Highlighting the lighting anomalies caused by facial replacement, and the audio features z aIt preserves pure, authentic voiceprint features. The cross-attention weight matrix shows that the visual-audio correlation in frames 12-15 (during speech) is significantly lower than the threshold, and the classifier outputs a forgery probability of 98.7%, thus accurately identifying deepfakes and avoiding unnecessary losses due to fake identities.
[0093] For example, in the medical field, consider a suspected fake medical science video (6 seconds, 30 frames), in which the doctor's face has been replaced but the original audio is retained. Each frame of the video is used to extract 768-dimensional features using ViT, resulting in a 30×768 matrix x. v MFCC audio input extracts 40-dimensional features, which are then encoded into a 30×512 matrix x using Wav2Vec2.0. a Furthermore, through feature decoupling, the shared feature z is achieved. s Capture the real doctor's identity features (consistent with audio), visual features z v Highlighting the lighting and shadow distortion caused by face replacement, audio features z a Maintain a clear and professional explanation of voiceprint characteristics. The cross-attention weight matrix shows that the visual-audio correlation in frames 18-22 (during image analysis) is significantly lower than the threshold, and the classifier outputs a forgery probability of 97.6%, thus accurately identifying deepfakes and avoiding adverse effects caused by the false dissemination of medical science knowledge.
[0094] In this embodiment, after obtaining the detection result, the method further includes:
[0095] When the detection result indicates that the target user's video generation is abnormal, the generation time of the detection result is obtained as the target timestamp;
[0096] The detection results are converted into a configuration format to obtain the conversion result;
[0097] The conversion result is marked with the target timestamp to obtain the marking result;
[0098] The marking results are sent to the triggerer of the detection command according to the high warning level.
[0099] The configuration format can be suitable for the corresponding system, such as JSON (JavaScript Object Notation) format.
[0100] Through the above embodiments, when an anomaly (such as forgery) is detected, the triggerer of the detection command can be notified in a timely manner based on the detection time, thereby prompting relevant personnel to deal with the anomaly in a timely manner and avoid further losses.
[0101] As can be seen from the above technical solutions, this invention can train a user-generated content detection model based on a cross-modal-adversarial domain invariant learning mechanism, ensuring cross-modal feature alignment, and eliminating dataset bias based on adversarial domain invariant learning, thereby enhancing the model's generalization ability to unknown distributions. It utilizes a feature extraction module to extract multimodal features from visual and audio inputs, providing basic features for subsequent cross-modal decoupling. The cross-modal decoupling module encodes visual and audio inputs to obtain decoupling features, enabling feature splitting, avoiding interference from irrelevant features, and providing a purer data foundation for subsequent detection. Based on a multi-head cross-attention mechanism, the decoupling features are fused, and a classifier is used to perform anomaly detection on the fused features. This adaptively aggregates multimodal features, highlighting consistent anomaly patterns across modalities, improving the ability to identify anomalous content, and thus increasing the accuracy of anomaly detection for user-generated content.
[0102] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the user-generated content detection device of the present invention. The user-generated content detection device 11 includes a retrieval unit 110, an extraction unit 111, an encoding unit 112, a fusion unit 113, and a detection unit 114. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0103] The retrieval unit 110 is used to retrieve a user-generated content detection model trained based on a cross-modal-adversarial domain invariant learning mechanism in response to a detection command for a target user-generated video.
[0104] In this embodiment, the detection command can be triggered by relevant maintenance personnel according to actual anomaly detection needs.
[0105] In this embodiment, the target user-generated video can be a user authentication video in the financial field or a medical knowledge popularization video in the medical and health field.
[0106] In this embodiment, before invoking the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism:
[0107] Constructing the reconstruction loss: in, Denotes the reconstruction loss; x v This represents the visual input extracted by the feature extraction module; Indicates with x v Correspondingly, the reconstructed visual features obtained after reconstruction by the decoder in the cross-modal decoupling module. D vThis represents the visual decoder function. z represents the vector concatenation operation. s The z represents the shared feature output by the shared encoder in the cross-modal decoupling module. v This represents the visual features output by the visual encoder in the cross-modal decoupling module; x a This represents the audio input extracted by the feature extraction module; Indicates with x a Correspondingly, the reconstructed audio features obtained after reconstruction by the decoder in the cross-modal decoupling module. D a z represents the audio decoder function. a This represents the audio features output by the audio encoder in the cross-modal decoupling module;
[0108] Constructing the contrast alignment loss: in, The comparison alignment loss represents the comparison alignment loss; s represents the cosine similarity function. This represents the shared features corresponding to the visual input; This represents the shared feature corresponding to the audio input; τ represents the temperature coefficient; This represents the audio shared features of the k-th negative sample; K represents the number of negative samples.
[0109] Based on the gradient inversion layer and domain discriminator in the cross-modal decoupling module, an adversarial domain-invariant loss is constructed: in, This represents the adversarial domain invariant loss; D represents the operation of finding the expected value; dom Representation domain discriminator function; z s This represents the shared features in the decoupling features output by the cross-modal decoupling module;
[0110] The user-generated content detection model is trained based on the reconstruction loss, the contrast alignment loss, and the adversarial domain invariance loss.
[0111] The shared encoder employs a parameter isolation design, ensuring its independence from the modal encoder through orthogonal constraints.
[0112] Among them, the domain discriminator combined with the modality discrimination loss can simultaneously optimize feature decoupling and modality alignment.
[0113] The training of the temporal discriminator attempts to distinguish the feature source dataset, while the shared encoder optimizes in reverse through a gradient inversion layer to confuse the domain discriminator, ultimately enabling the shared features to ignore the unique distribution of the dataset.
[0114] In the above embodiments, the reconstruction loss and the contrast alignment loss can enhance the model's ability to distinguish similar but different categories, making the differentiation of fake data more effective; adversarial training can remove dataset bias and enhance the model's ability to generalize to unknown distributions.
[0115] The extraction unit 111 is used to extract features from the target user-generated video using the feature extraction module in the user-generated content detection model to obtain the target visual input and the target audio input.
[0116] In this embodiment, before the extraction unit 111 uses the feature extraction module in the user-generated content detection model to extract features from the target user-generated video, it performs frame sampling and processing on the target user-generated video to obtain a visual frame sequence and synchronized MFCC (Mel-Frequency Cepstral Coefficients) audio features.
[0117] The visual frame sequence is subjected to frame alignment and normalization processing.
[0118] The MFCC audio features are subjected to audio resampling processing.
[0119] The above embodiments enable the standardization of the original video, providing input data in a unified format for subsequent feature extraction.
[0120] In this embodiment, the extraction unit 111 uses the feature extraction module in the user-generated content detection model to extract features from the target user-generated video, obtaining the target visual input and target audio input, including:
[0121] Based on the ViT (Vision Transformer) architecture, visual features are extracted from the target user-generated video to obtain the target visual input;
[0122] An unsupervised speech pre-training model is used to extract acoustic features from the target user-generated video to obtain the target audio input.
[0123] For example, the unsupervised speech pre-training model can be a Wav2Vec 2.0 model.
[0124] In the above embodiments, the ViT architecture is used as a visual feature extractor and the unsupervised speech pre-trained model is used as an audio feature extractor, which can extract multimodal features and provide a data foundation for subsequent feature decoupling.
[0125] The encoding unit 112 is used to encode the target visual input and the target audio input using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features.
[0126] In this embodiment, the cross-modal decoupling module includes a shared encoder, a visual encoder, and an audio encoder and decoder.
[0127] The shared encoder is used to output identity-related features.
[0128] The visual encoder is used to capture traces of forgery.
[0129] The audio encoder is used to extract audio authenticity features.
[0130] The decoder employs a symmetric encoder structure to constrain feature integrity through reconstruction loss. The decoder includes a visual decoder and an audio decoder.
[0131] In this embodiment, the encoding unit 112 uses the cross-modal decoupling module in the user-generated content detection model to encode the target visual input and the target audio input, obtaining target decoupling features including:
[0132] The target visual input and the target audio input are encoded using the shared encoder to obtain target shared features;
[0133] The visual encoder is used to encode the target visual input to obtain the target visual features;
[0134] The target audio input is encoded using the audio encoder to obtain target audio features;
[0135] The target shared features, the target visual features, and the target audio features are determined as the target decoupling features;
[0136] The shared encoder includes 4 Transformer layers, the visual encoder includes 2 convolutional neural network layers and 1 Transformer layer, and the audio encoder includes 2 convolutional neural network layers and 1 Transformer layer.
[0137] The cross-modal decoupling module processes visual and audio inputs through a dual-branch VAE (Variational Autoencoder) structure.
[0138] Through the above embodiments, the input features can be decomposed into shared features and modality-specific features, avoiding interference from irrelevant features and providing purer features related to forgery for subsequent detection.
[0139] The fusion unit 113 is used to fuse the target decoupled features based on a multi-head cross-attention mechanism to obtain target fused features.
[0140] In this embodiment, the fusion unit 113 fuses the target decoupling features based on a multi-head cross-attention mechanism to obtain target fused features, including:
[0141] The target fusion features are calculated using the following formula:
[0142] Where Q = W q z v K = W k z a V = W v z a ;
[0143] Where h represents the target fusion feature; Q represents the query vector; K represents the key vector; V represents the value vector; d represents the feature dimension; W q W k W v This represents the projection matrix.
[0144] For example, an 8-head attention mechanism can be used to generate fused features based on a multi-head cross-attention mechanism, and dynamic feature aggregation can be achieved by calculating the correlation between visual and audio features.
[0145] Among them, the cross-attention mechanism introduces learnable relative position encoding, which can effectively capture the lip-speech temporal relationship.
[0146] The target fusion feature can focus on forgery traces that are inconsistent between visual and audio (such as lip-syncing and speech asynchrony).
[0147] Through the above embodiments, multimodal features can be adaptively aggregated to highlight consistent forgery patterns across modalities and improve the ability to identify forged content.
[0148] The detection unit 114 is used to detect the target fusion features using the classifier in the user-generated content detection model to obtain the detection result.
[0149] In this embodiment, the detection unit 114 uses the classifier in the user-generated content detection model to detect the target fusion features, and the detection results include:
[0150] The target fusion features are input into a deep residual network to obtain the anomaly probability;
[0151] Obtain the probability threshold;
[0152] When the anomaly probability is greater than the probability threshold, the detection result is determined to be an anomaly in the target user's generated video; or
[0153] When the anomaly probability is less than or equal to the probability threshold, the detection result is determined to be that the target user's generated video has no anomalies.
[0154] For example, the deep residual network can be a lightweight ResNet-18 (Residual Network 18) network.
[0155] Specifically, the deep residual network can be used to classify fused features and output the forgery probability.
[0156] The probability threshold can be selected as the optimal value based on the experimental results.
[0157] The above embodiments enable the detection decision of whether the input content is a deepfake.
[0158] This embodiment improves the robustness and accuracy of detection by integrating a multimodal learning framework for visual and audio signals. It employs cross-modal decoupling to decompose input features into shared features (such as speaker identity) and modality-specific features (such as visual and audio features), utilizes a variational autoencoder framework for feature separation, and ensures cross-modal feature alignment through contrastive loss. It combines an adversarial domain-invariant representation learning mechanism to eliminate dataset bias and enhance the model's generalization ability to unknown distributions. A cross-attention feature fusion mechanism aggregates multimodal information, highlighting forgery-related features, and a lightweight classifier performs the detection. This embodiment uses a Transformer-based encoder (such as ViT and Wav2Vec 2.0) and a multi-head attention mechanism, significantly outperforming traditional single-modal or simple fusion methods, and is suitable for deep forgery detection in complex user-generated content.
[0159] For example, in the financial field, consider a short video (5 seconds, 25 frames) suspected of being a forged identity verification profile, where the person's face has been replaced but the original audio is retained. Each frame of the video is used to extract 768-dimensional features using ViT, resulting in a 25×768 matrix x. v MFCC audio input extracts 40-dimensional features, which are then encoded into a 25×512 matrix x using Wav2Vec 2.0. a Furthermore, through feature decoupling, the shared feature z is achieved. s Capture the real speaker's identity features (consistent with the audio), visual features z v Highlighting the lighting anomalies caused by facial replacement, and the audio features z aIt preserves pure, authentic voiceprint features. The cross-attention weight matrix shows that the visual-audio correlation in frames 12-15 (during speech) is significantly lower than the threshold, and the classifier outputs a forgery probability of 98.7%, thus accurately identifying deepfakes and avoiding unnecessary losses due to fake identities.
[0160] For example, in the medical field, consider a suspected fake medical science video (6 seconds, 30 frames), in which the doctor's face has been replaced but the original audio is retained. Each frame of the video is used to extract 768-dimensional features using ViT, resulting in a 30×768 matrix x. v MFCC audio input extracts 40-dimensional features, which are then encoded into a 30×512 matrix x using Wav2Vec2.0. a Furthermore, through feature decoupling, the shared feature z is achieved. s Capture the real doctor's identity features (consistent with audio), visual features z v Highlighting the lighting and shadow distortion caused by face replacement, audio features z a Maintain a clear and professional explanation of voiceprint characteristics. The cross-attention weight matrix shows that the visual-audio correlation in frames 18-22 (during image analysis) is significantly lower than the threshold, and the classifier outputs a forgery probability of 97.6%, thus accurately identifying deepfakes and avoiding adverse effects caused by the false dissemination of medical science knowledge.
[0161] In this embodiment, after obtaining the detection result, when the detection result indicates that the target user's video generation is abnormal, the generation time of the detection result is obtained as the target timestamp;
[0162] The detection results are converted into a configuration format to obtain the conversion result;
[0163] The conversion result is marked with the target timestamp to obtain the marking result;
[0164] The marking results are sent to the triggerer of the detection command according to the high warning level.
[0165] The configuration format can be suitable for the corresponding system, such as JSON (JavaScript Object Notation) format.
[0166] Through the above embodiments, when an anomaly (such as forgery) is detected, the triggerer of the detection command can be notified in a timely manner based on the detection time, thereby prompting relevant personnel to deal with the anomaly in a timely manner and avoid further losses.
[0167] As can be seen from the above technical solutions, this invention can train a user-generated content detection model based on a cross-modal-adversarial domain invariant learning mechanism, ensuring cross-modal feature alignment, and eliminating dataset bias based on adversarial domain invariant learning, thereby enhancing the model's generalization ability to unknown distributions. It utilizes a feature extraction module to extract multimodal features from visual and audio inputs, providing basic features for subsequent cross-modal decoupling. The cross-modal decoupling module encodes visual and audio inputs to obtain decoupling features, enabling feature splitting, avoiding interference from irrelevant features, and providing a purer data foundation for subsequent detection. Based on a multi-head cross-attention mechanism, the decoupling features are fused, and a classifier is used to perform anomaly detection on the fused features. This adaptively aggregates multimodal features, highlighting consistent anomaly patterns across modalities, improving the ability to identify anomalous content, and thus increasing the accuracy of anomaly detection for user-generated content.
[0168] like Figure 3 The diagram shown is a schematic representation of the structure of a computer device that implements the user-generated content detection method of the present invention.
[0169] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a user-generated content detection program.
[0170] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0171] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0172] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as code for user-generated content detection programs, but also to temporarily store data that has been output or will be output.
[0173] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing user-generated content detection programs) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0174] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the various user-generated content detection method embodiments described above, for example... Figure 1 The steps are shown.
[0175] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a retrieval unit 110, an extraction unit 111, an encoding unit 112, a fusion unit 113, and a detection unit 114.
[0176] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the user-generated content detection method described in the various embodiments of this invention.
[0177] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0178] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0179] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0180] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0181] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0182] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0183] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.
[0184] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0185] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0186] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0187] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a user-generated content detection method, and the processor 13 can execute the multiple instructions to achieve the following:
[0188] In response to the detection command for the target user-generated video, the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism is invoked;
[0189] The feature extraction module in the user-generated content detection model is used to extract features from the target user-generated video to obtain the target visual input and the target audio input.
[0190] The target visual input and the target audio input are encoded using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features;
[0191] The target decoupling features are fused based on a multi-head cross-attention mechanism to obtain target fused features;
[0192] The target fusion features are detected using the classifier in the user-generated content detection model to obtain the detection results.
[0193] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0194] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0195] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0196] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0197] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0198] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0199] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0200] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0201] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for detecting user-generated content, characterized in that, The user-generated content detection method includes: In response to the detection command for the target user-generated video, the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism is invoked; The feature extraction module in the user-generated content detection model is used to extract features from the target user-generated video to obtain the target visual input and the target audio input. The target visual input and the target audio input are encoded using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features; The target decoupling features are fused based on a multi-head cross-attention mechanism to obtain target fused features; The target fusion features are detected using the classifier in the user-generated content detection model to obtain the detection results.
2. The user-generated content detection method as described in claim 1, characterized in that, Before invoking the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism, the method further includes: Constructing the reconstruction loss: in, Denotes the reconstruction loss; x v This represents the visual input extracted by the feature extraction module; Indicates with x v Correspondingly, the reconstructed visual features obtained after reconstruction by the decoder in the cross-modal decoupling module. D v This represents the visual decoder function. z represents the vector concatenation operation. s The z represents the shared feature output by the shared encoder in the cross-modal decoupling module. v This represents the visual features output by the visual encoder in the cross-modal decoupling module; x a This represents the audio input extracted by the feature extraction module; Indicates with x a Correspondingly, the reconstructed audio features obtained after reconstruction by the decoder in the cross-modal decoupling module. D a z represents the audio decoder function. a This represents the audio features output by the audio encoder in the cross-modal decoupling module; Constructing the contrast alignment loss: in, The comparison alignment loss represents the comparison alignment loss; s represents the cosine similarity function. This represents the shared features corresponding to the visual input; This represents the shared feature corresponding to the audio input; τ represents the temperature coefficient; This represents the audio shared features of the k-th negative sample; K represents the number of negative samples. Based on the gradient inversion layer and domain discriminator in the cross-modal decoupling module, an adversarial domain-invariant loss is constructed: in, This represents the adversarial domain invariant loss; D represents the operation of finding the expected value; dom Representation domain discriminator function; z s This represents the shared features in the decoupling features output by the cross-modal decoupling module; The user-generated content detection model is trained based on the reconstruction loss, the contrast alignment loss, and the adversarial domain invariance loss.
3. The user-generated content detection method as described in claim 1, characterized in that, The step of using the feature extraction module in the user-generated content detection model to extract features from the target user-generated video to obtain the target visual input and target audio input includes: Based on the ViT architecture, visual features are extracted from the target user-generated video to obtain the target visual input; An unsupervised speech pre-training model is used to extract acoustic features from the target user-generated video to obtain the target audio input.
4. The user-generated content detection method as described in claim 2, characterized in that, The step of encoding the target visual input and the target audio input using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features includes: The target visual input and the target audio input are encoded using the shared encoder to obtain target shared features; The visual encoder is used to encode the target visual input to obtain the target visual features; The target audio input is encoded using the audio encoder to obtain target audio features; The target shared features, the target visual features, and the target audio features are determined as the target decoupling features; The shared encoder includes 4 Transformer layers, the visual encoder includes 2 convolutional neural network layers and 1 Transformer layer, and the audio encoder includes 2 convolutional neural network layers and 1 Transformer layer.
5. The user-generated content detection method as described in claim 2, characterized in that, The method of fusing the target decoupling features based on the multi-head cross-attention mechanism to obtain the target fused features includes: The target fusion features are calculated using the following formula: where Q = W q z v , K = W k z a , V = W v z a ; Where h represents the target fusion feature; Q represents the query vector; K represents the key vector; V represents the value vector; d represents the feature dimension; W q W k W v This represents the projection matrix.
6. The user-generated content detection method as described in claim 1, characterized in that, The detection of the target fusion features using the classifier in the user-generated content detection model yields the following results: The target fusion features are input into a deep residual network to obtain the anomaly probability; Obtain the probability threshold; When the anomaly probability is greater than the probability threshold, the detection result is determined to be an anomaly in the target user's generated video; or When the anomaly probability is less than or equal to the probability threshold, the detection result is determined to be that the target user's generated video has no anomalies.
7. The user-generated content detection method as described in claim 6, characterized in that, After obtaining the detection result, the method further includes: When the detection result indicates that the target user's video generation is abnormal, the generation time of the detection result is obtained as the target timestamp; The detection results are converted into a configuration format to obtain the conversion result; The conversion result is marked with the target timestamp to obtain the marking result; The marking results are sent to the triggerer of the detection command according to the high warning level.
8. A user-generated content detection device, characterized in that, The user-generated content detection device includes: The retrieval unit is used to retrieve the user-generated content detection model trained based on the cross-modal-adversarial domain invariant learning mechanism in response to the detection command for the target user-generated video; The extraction unit is used to extract features from the target user-generated video using the feature extraction module in the user-generated content detection model to obtain the target visual input and the target audio input. The encoding unit is used to encode the target visual input and the target audio input using the cross-modal decoupling module in the user-generated content detection model to obtain target decoupling features; The fusion unit is used to fuse the target decoupled features based on a multi-head cross-attention mechanism to obtain the target fused features; The detection unit is used to detect the target fusion features using the classifier in the user-generated content detection model to obtain the detection result.
9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the user-generated content detection method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the user-generated content detection method as described in any one of claims 1 to 7.