Deep counterfeit multimedia identification method and system based on artificial intelligence

By extracting multimodal features and generating link reasoning, combined with knowledge graph analysis, an adaptive deep fake multimedia recognition system is constructed, which solves the problems of single modality detection and insufficient generalization capabilities of existing technologies, and realizes efficient and accurate detection of fake content.

CN120670950APending Publication Date: 2025-09-19TIANJIN NAT CYBERNET SECURITY CO LTD

Patent Information

Application Number
CN202510776610.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing deep fake detection technologies have problems such as single-modal detection, insufficient generalization capabilities, low computational efficiency, and lack of generation link analysis, making it difficult to effectively identify multi-modal collaborative fake content.

Method used

By adopting multimodal feature extraction and generation link reasoning, combined with dynamic analysis of knowledge graphs, a multimodal classification model is constructed, which adaptively adjusts thresholds and generates countermeasure strategies to achieve high-precision counterfeit content detection.

Benefits of technology

Through multimodal collaborative detection, the detection accuracy is improved, new counterfeiting techniques are quickly identified, the false positive rate is reduced, real-time detection is supported, and large-scale content review needs are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670950A_ABST
    Figure CN120670950A_ABST
Patent Text Reader

Abstract

The invention provides a deep counterfeit multimedia identification method and system based on artificial intelligence, and relates to the technical field of artificial intelligence security. According to the method, high-precision forged content detection is realized through multi-modal feature extraction and link generation reasoning. The method specifically comprises the steps of multi-modal data acquisition and feature extraction, generated link reasoning and knowledge graph adaptation, multi-modal correlation analysis and dynamic threshold setting, forged path rule construction and decision model training, and self-adaptive identification strategy generation. According to the method, multi-modal features are fused, links are generated based on dynamic analysis of the knowledge graph, adaptive threshold judgment and targeted countering are achieved, the problems that a traditional method is insufficient in generalization ability, low in calculation efficiency and the like are solved, and the method can be widely applied to content authenticity identification in the fields of social media, public opinion and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to an artificial intelligence-based deep fake multimedia identification method and system. Background Art

[0002] With the rapid development of artificial intelligence technologies such as generative adversarial networks (GANs) and diffusion models, the threshold for deepfake technology has been significantly lowered. Malicious users use these technologies to forge images, audio, and video, which are widely disseminated on social media and in public opinion, posing a serious threat to personal privacy, social security, and public trust.

[0003] Existing deepfake detection technologies have the following main limitations:

[0004] Single modality detection: Most methods only analyze a single modality in images, audio, or video, and are unable to cope with complex scenarios of multi-modal collaborative forgery.

[0005] Insufficient generalization ability: Traditional models rely on fixed features and thresholds, have poor adaptability to new counterfeiting techniques, and are prone to misjudgment.

[0006] Low computational efficiency: The computational complexity of multimodal data processing is high, making it difficult to meet real-time detection requirements.

[0007] Lack of generation chain analysis: Existing methods do not model the generation path of forged content and cannot effectively identify forged content generated using new tool chains.

[0008] Therefore, there is an urgent need for a targeted deep fake multimedia identification method and system based on artificial intelligence. Summary of the Invention

[0009] The purpose of this invention is to provide an artificial intelligence-based deep fake multimedia identification method and system, which can achieve high-precision forged content detection through multimodal feature extraction and generative link reasoning, and dynamically generate identification strategies for different forgery techniques, avoiding misjudgments caused by traditional methods due to insufficient generalization capabilities.

[0010] In a first aspect, the present application provides an artificial intelligence-based deepfake multimedia identification method, the method comprising:

[0011] Multimodal data collection and feature extraction steps, including extracting image features, audio features, and video timing features, and adaptively selecting feature extraction algorithms based on the data source;

[0012] Generate link reasoning and knowledge graph adaptation steps, including building a forgery technology knowledge graph and assigning dynamic importance values ​​to different modal features based on the graph;

[0013] Multimodal correlation analysis and dynamic threshold setting steps, including calculating the mean of multimodal feature consistency and setting the dynamic threshold based on the deviation of real-time data from the mean;

[0014] Forged path rule construction and decision model training steps, including designing inference rules, building and training a multimodal classification model;

[0015] Adaptive identification strategy generation step, including generating visual evidence reports and targeted countermeasures.

[0016] In a second aspect, the present application provides an artificial intelligence-based deepfake multimedia identification system, the system comprising:

[0017] Input module, used for accessing multimedia data stream;

[0018] Feature extraction module, used to extract image features, audio features, and video timing features, and adaptively select feature extraction algorithms based on the data source;

[0019] A knowledge graph adaptation module, which is used to construct a knowledge graph of counterfeiting technology and assign dynamic importance values ​​to different modal features based on the graph;

[0020] The analysis and threshold setting module is used to calculate the mean of multimodal feature consistency and set dynamic thresholds based on the deviation between real-time data and the mean;

[0021] Model training module, used to design inference rules, build and train multimodal classification models;

[0022] Adaptive strategy module for generating visual evidence reports and targeted countermeasures.

[0023] In a third aspect, the present application provides an artificial intelligence-based deepfake multimedia identification system, the system comprising a processor and a memory:

[0024] The memory is used to store program code and transmit the program code to the processor;

[0025] The processor is configured to execute any one of the possible methods of the first aspect according to instructions in the program code.

[0026] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement any one of the methods described in the first aspect.

[0027] Beneficial effects

[0028] The present invention provides a deep fake multimedia identification method and system based on artificial intelligence, which realizes high-precision fake content detection through multimodal feature extraction and generation link reasoning. Specifically including: multimodal data collection and feature extraction, generation link reasoning and knowledge graph adaptation, multimodal correlation analysis and dynamic threshold setting, fake path rule construction and decision model training, and adaptive identification strategy generation. The present invention integrates multimodal features, generates links based on knowledge graph dynamic analysis, realizes adaptive threshold judgment and targeted countermeasures, solves the problems of insufficient generalization ability and low computational efficiency of traditional methods, and can be widely used in content authenticity identification in social media, public opinion and other fields.

[0029] The method and system of the present invention have the following advantages and effects:

[0030] 1. Multimodal collaborative detection: By fusing image, audio, and video features, it overcomes the limitations of single-modality detection and significantly improves detection accuracy.

[0031] 2. Dynamically generated link reasoning: Based on the knowledge graph, the generation technology association analysis can quickly identify new counterfeiting techniques and improve the system's generalization capabilities.

[0032] 3. Adaptive Thresholds and Strategies: Dynamically adjust the judgment threshold through real-time data to reduce the false positive rate and generate personalized countermeasures for different counterfeiting techniques.

[0033] 4. Efficient computing optimization: Through depth and breadth algorithm optimization, the computational overhead of multimodal data processing is reduced, supporting real-time detection and meeting the needs of large-scale content review. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0035] Figure 1 is a flow chart of the present invention;

[0036] Figure 2 This is a system architecture diagram of the present invention. DETAILED DESCRIPTION

[0037] The preferred embodiments of the present invention are described in detail below with reference to the accompanying drawings so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making a clearer and more precise definition of the protection scope of the present invention.

[0038] The present application provides an artificial intelligence-based deep fake multimedia identification method, which includes:

[0039] Multimodal data collection and feature extraction steps, including extracting image features, audio features, and video timing features, and adaptively selecting feature extraction algorithms based on the data source;

[0040] Specifically, the image, audio, and video data to be detected are preprocessed, including format conversion, noise filtering, and spatiotemporal synchronization.

[0041] According to the source of multimodal data generation (such as social media, live streaming), adaptively select deep algorithms (such as ResNet, ViT model) and wide algorithms (such as facial area priority or global background analysis) for feature extraction.

[0042] Generate link reasoning and knowledge graph adaptation steps, including building a forgery technology knowledge graph and assigning dynamic importance values ​​to different modal features based on the graph;

[0043] Specifically, based on the forgery scenarios (such as face-changing, voice cloning, and video synthesis), a generation technology knowledge graph is constructed, including: forgery tool chain: typical features of GAN, Autoencoder, and Diffusion models; generation link dependencies: for example, face-changing requires facial key point alignment, and voice cloning requires voiceprint feature extraction and synthesis model iteration.

[0044] Based on the knowledge graph, dynamic importance values ​​are assigned to different modal features. For example, in the face-changing scenario, the weight of facial edge artifacts is higher than that of audio features; in the voice cloning scenario, the weight of voiceprint spectrum continuity is higher than that of video features.

[0045] Multimodal correlation analysis and dynamic threshold setting steps, including calculating the mean of multimodal feature consistency and setting the dynamic threshold based on the deviation of real-time data from the mean;

[0046] Specifically, timestamps are used to analyze the duration of forgery traces in the video and calculate the mean consistency of multimodal features. Based on the deviation between real-time data and the historical mean, a dynamic threshold is set to determine the probability of forgery. For example, if the duration of an abnormal facial micro-expression exceeds the threshold and is inconsistent with the emotional characteristics of the audio, the content is judged as high-risk forgery.

[0047] Forged path rule construction and decision model training steps, including designing inference rules, building and training a multimodal classification model;

[0048] The designed inference rules include: if GAN-generated artifacts (such as unnatural high-frequency noise) are detected in the image and there is a synthetic break in the audio spectrum, it is judged as multimodal collaborative forgery; if the motion trajectory between video frames does not conform to physical laws (such as abnormal head rotation acceleration), an association with the generation tool chain (such as Autoencoder iteration defects) is established.

[0049] Model construction and training include: label encoding of multimodal features (such as mapping the type of counterfeit tool to a digital label); vectorizing the feature description text through the BERT or CLIP model and calculating cross-modal similarity; assigning weights to different modalities based on important values ​​and dynamic thresholds, building a random forest or gradient boosting decision tree, and outputting the counterfeit probability and generation technology classification.

[0050] Adaptive identification strategy generation step, including generating visual evidence reports and targeted countermeasures.

[0051] Specifically, evidence visualization includes: embedding detection models and inference rules into the content review system, and automatically generating visual evidence reports for high-risk content (such as highlighting forged areas and playing abnormal audio clips).

[0052] Recommended countermeasure strategies include: recommending targeted countermeasures based on the type of generation technology (such as adding digital watermarks to face-swapped content and triggering voiceprint secondary verification for voice cloned content).

[0053] In some preferred embodiments, the image features include facial micro-expressions, lighting consistency and edge artifact features; the audio features include spectrum continuity and speech emotion consistency features; and the video timing features include inter-frame motion trajectory and lip synchronization features.

[0054] Specifically, CNN is used to extract facial micro-expressions, lighting consistency, and edge artifact features; MFCC (Mel-frequency cepstral coefficients) and voiceprint models are used to extract spectral continuity and speech emotion consistency features; and 3D convolutional networks are used to extract inter-frame motion trajectories and lip synchronization features.

[0055] In some preferred embodiments, in the step of generating link reasoning and adapting the knowledge graph, the knowledge graph includes a forging tool chain and generating link dependencies, and the forging tool chain includes typical features of GAN, Autoencoder and Diffusion models.

[0056] In some preferred embodiments, in the multimodal correlation analysis and dynamic threshold setting steps, if the abnormal duration of facial micro-expressions exceeds a threshold and is inconsistent with the audio emotional characteristics, it is determined to be high-risk forged content.

[0057] In some preferred embodiments, in the forged path rule construction and decision model training steps, the feature description text is vectorized using a BERT or CLIP model to calculate cross-modal similarity.

[0058] Example:

[0059] Taking face-changing video detection as an example, the specific implementation process of the present invention is as follows:

[0060] 1. Data preprocessing: Obtain the video to be detected, extract the video frames and audio tracks, and perform spatiotemporal synchronization.

[0061] 2. Feature extraction:

[0062] Extract micro-expression features of facial areas (such as abnormal blinking frequency) through CNN;

[0063] Analyze head motion trajectory between frames through 3D convolutional network;

[0064] Detect voiceprint and lip synchronization deviation in audio through MFCC and voiceprint models.

[0065] 3. Generative Link Reasoning: The knowledge graph matches the association rules between "GAN generation artifacts" and "Autoencoder motion model defects," assigning high weights to facial edge artifacts and motion trajectory features.

[0066] 4. Multimodal correlation analysis: By calculating the mean consistency of multimodal features, it was found that the duration of abnormal facial micro-expressions exceeded the threshold and was inconsistent with the audio emotional characteristics.

[0067] 5. Decision output: Dynamically determined to be high-risk forged content, with a forgery probability of 92%. The generation technology is classified as "GAN+Autoencoder face-swapping."

[0068] 6. Countermeasures: Generate a report containing abnormal area markers, trigger the manual review queue for priority processing, and add an anti-proliferation watermark to the video.

[0069] Figure 2 This is an architectural diagram of the AI-based deepfake multimedia identification system provided in this application. The system includes:

[0070] Input module, used for accessing multimedia data stream;

[0071] Feature extraction module, used to extract image features, audio features, and video timing features, and adaptively select feature extraction algorithms based on the data source;

[0072] A knowledge graph adaptation module, which is used to construct a knowledge graph of counterfeiting technology and assign dynamic importance values ​​to different modal features based on the graph;

[0073] The analysis and threshold setting module is used to calculate the mean of multimodal feature consistency and set dynamic thresholds based on the deviation between real-time data and the mean;

[0074] Model training module, used to design inference rules, build and train multimodal classification models;

[0075] Adaptive strategy module for generating visual evidence reports and targeted countermeasures.

[0076] The present application provides an artificial intelligence-based deep fake multimedia identification system, the system comprising: the system comprising a processor and a memory:

[0077] The memory is used to store program code and transmit the program code to the processor;

[0078] The processor is configured to execute the method described in any one of all embodiments of the first aspect according to instructions in the program code.

[0079] The present application provides a computer-readable storage medium, which is used to store program code, and the program code is used to be executed by a processor to implement any one of the methods in all embodiments of the first aspect.

[0080] In a specific implementation, the present invention further provides a computer storage medium, wherein the computer storage medium may store a program that, when executed, may include some or all of the steps of various embodiments of the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0081] Those skilled in the art will clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus the necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention or certain portions of the embodiments.

[0082] In particular, for the embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.

[0083] The above-described embodiments of the present invention do not limit the protection scope of the present invention.

Claims

1. A deep fake multimedia identification method based on artificial intelligence, characterized by: The method comprises: Multimodal data collection and feature extraction steps, including extracting image features, audio features, and video timing features, and adaptively selecting feature extraction algorithms based on the data source; Generate link reasoning and knowledge graph adaptation steps, including building a forgery technology knowledge graph and assigning dynamic importance values ​​to different modal features based on the graph; Multimodal correlation analysis and dynamic threshold setting steps, including calculating the mean of multimodal feature consistency and setting the dynamic threshold based on the deviation of real-time data from the mean; Forged path rule construction and decision model training steps, including designing inference rules, building and training a multimodal classification model; Adaptive identification strategy generation step, including generating visual evidence reports and targeted countermeasures.

2. The method according to claim 1, wherein: The image features include facial micro-expressions, lighting consistency and edge artifact features; the audio features include spectrum continuity and speech emotion consistency features; and the video timing features include inter-frame motion trajectory and lip synchronization features.

3. The method according to claim 1, wherein: In the step of generating link reasoning and adapting the knowledge graph, the knowledge graph includes a forgery tool chain and generation link dependencies, and the forgery tool chain includes typical features of GAN, Autoencoder and Diffusion models.

4. The method according to claim 1, wherein: In the multimodal correlation analysis and dynamic threshold setting steps, if the duration of abnormal facial micro-expressions exceeds the threshold and is inconsistent with the audio emotional characteristics, it is determined to be high-risk forged content.

5. The method according to claim 1, wherein: In the forged path rule construction and decision model training steps, the feature description text is vectorized using the BERT or CLIP model to calculate cross-modal similarity.

6. An artificial intelligence-based deep fake multimedia identification system, characterized by: The system comprises: Input module, used for accessing multimedia data stream; Feature extraction module, used to extract image features, audio features, and video timing features, and adaptively select feature extraction algorithms based on the data source; A knowledge graph adaptation module, which is used to construct a knowledge graph of counterfeiting technology and assign dynamic importance values ​​to different modal features based on the graph; The analysis and threshold setting module is used to calculate the mean of multimodal feature consistency and set dynamic thresholds based on the deviation between real-time data and the mean; Model training module, used to design inference rules, build and train multimodal classification models; Adaptive strategy module for generating visual evidence reports and targeted countermeasures.

7. An artificial intelligence-based deep fake multimedia identification system, characterized by: The system includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to implement the method according to any one of claims 1 to 5 according to the instructions in the program code.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to be executed by a processor to implement the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Deep false detection method and system based on AI

    CN117972610A

  • Multi-mode depth forgery detection method and device

    CN119904773A

Cited By

  • Intelligent generated image detection method based on multi-granularity artifact feature fusion

    CN121147726A

  • Digital multimedia evidence obtaining method and device

    CN121660104A