Deep forgery detection method based on multi-modal features and electronic equipment
Through the deep forgery detection method of multimodal features, the problem of poor single modal detection effect is solved, and the comprehensive identification and traceability of forgery content of multiple data types is achieved, which improves the accuracy and efficiency of detection.
Patent Information
- Application Number
- CN202510902331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
AI Technical Summary
The existing deep forgery detection methods are mainly based on a single mode, making it difficult to fully identify and track complex deep forgery content generated by multiple technologies, especially when multiple sensor inputs or data types are involved, the detection effect is greatly reduced.
The deep forgery detection method based on multimodal features is adopted to obtain the data to be detected from the public network, extract the multimodal feature vector, and search using the vector index library. Combining metadata and classification networks, forgery relationship information is constructed to realize the identification and traceability of multimodal forgery content.
Through the cross-recognition and integration of multimodal features, the ability to identify complex forged content is significantly improved, the comprehensiveness and accuracy of detection is enhanced, and the propagation links and sources of forged data can be effectively tracked.
Smart Images

Figure CN120407883A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of artificial intelligence, and more specifically, to a deepfake detection method and an electronic device based on multi-modal features. Background Art
[0002] Deepfake detection technology is a rapidly emerging field in recent years with the rapid development of artificial intelligence and AIGC (Artificial Intelligence Generated Content) technology. With the development of AIGC technology, people can generate more realistic virtual entities such as pictures, audio, text, videos, etc. through various algorithm models. Deepfake technology originated from the rapid development of the fields of deep learning and artificial intelligence. Initially, these technologies were mainly used in academic research and the entertainment industry, such as the production of movie special effects. With the development of Generative Adversarial Networks (GANs) and other deep network structures, it has become increasingly easy to create realistic image, audio, and video content. In recent years, deepfake technology has evolved to the point where high-quality forged content can be produced through simple software applications without the need for a deep technical background. For example, deepfake face-swapping applications allow users to realistically replace the facial features of one person into the video of another person, while voice cloning technology can replicate an individual's voice characteristics to generate brand-new voice content. In the entertainment industry, deepfake technology is used in fields such as movie production, game design, and virtual reality to provide users with a richer and more realistic experience. However, this also brings a serious problem, that is, the abuse of deepfake technology, which in turn has some negative impacts on individuals and society. Moreover, forgery technology also poses a direct threat to personal privacy, such as blackmailing and slandering by forging inappropriate videos or audio of an individual.
[0003] With the continuous progress of technology, forged content has become increasingly difficult to distinguish with the naked eye or simple technical means. At the same time, deepfake technology has triggered a series of legal and ethical issues, including conflicts between copyright, freedom of speech, and individual rights. This requires relevant technologies to also keep progressing to detect and prevent these advanced forgery methods. Existing deepfake detection methods mainly conduct deepfake detection based on a single modality, such as image single-modal detection, video single-modal detection, voice single-modal detection, and text single-modal detection. They are usually limited to a single data modality and may not be able to comprehensively identify and track complex deepfake content generated by multiple technologies. When forgery technology involves multiple sensor inputs or data types, the detection effect of single-modal methods is greatly reduced. Summary of the Invention
[0004] The present disclosure provides a deepfake detection method based on multimodal features to solve at least one of the above problems.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a deepfake detection method, the deepfake detection method comprising: obtaining data to be detected from a public network; determining a multimodal feature vector extracted from the data to be detected as a vector to be detected; performing a retrieval process on the data to be detected based on a vector index library and the vector to be detected, wherein the vector index library is constructed based on fake data that meets a first preset forgery condition collected from a public network, and each data entry stored in the vector index library includes corresponding fake data and a multimodal feature vector; and determining the data to be detected as target fake data when a data entry matching the data to be detected is retrieved from the vector index library.
[0006] Optionally, the deepfake detection method further comprises: adding the target fake data and its multimodal feature vector as a new data entry to the vector index library.
[0007] Optionally, each data entry in the vector index library is associated with metadata, the metadata including at least one of the following metadata of the corresponding fake data: release platform, release time, release account, affiliated topic, dissemination times, wherein the deepfake detection method further comprises: obtaining the metadata of the target fake data; retrieving a preset number of data entries that best match the target fake data from the vector index library; and constructing forgery relationship information of the target fake data according to the metadata of the target fake data and the metadata of the preset number of data entries, wherein the forgery relationship information includes at least one of the following: dissemination link, correlation graph.
[0008] Optionally, the deepfake detection method further comprises: processing the target fake data using a classification network to obtain category information of the target fake data, wherein the category information includes at least one of the following: forgery category, severity category, and the forgery category includes at least one of the following: face swapping, voice forgery, false subtitles.
[0009] Optionally, the deepfake detection method further comprises: receiving fake data collected by a user feedback platform, creating a new data entry based on the received fake data, and adding it to the vector index library.
[0010] Optionally, the deepfake detection method further includes: sending the detection result of the target fake data to a user feedback platform for the user feedback platform to output the detection result of the target fake data, and outputting user appeal information in response to receiving a user appeal request for the target fake data returned by the user feedback platform.
[0011] Optionally, the deepfake detection method further includes: in a case where no data entry matching the data to be detected is retrieved from the vector index library, using a preset deep reinforcement learning network to process the data to be detected to obtain action information, where the action information includes at least one of the following actions: detection passed, intercepted, reported, manual review.
[0012] Optionally, the deepfake detection method further includes: screening out data entries in the vector index library that meet a second preset forgery condition to obtain a vector index sub-library of the vector index library, where the second preset forgery condition is a variable condition.
[0013] According to a second aspect of the embodiments of the present disclosure, there is provided a deepfake detection device, where the deepfake detection device includes: a first acquisition unit configured to acquire data to be detected from a public network; an extraction unit configured to determine a multi-modal feature vector extracted from the data to be detected as a vector to be detected; a retrieval unit configured to perform a retrieval process on the data to be detected based on a vector index library and the vector to be detected, where the vector index library is constructed based on fake data that meets a first preset forgery condition collected from the public network, and each data entry stored in the vector index library includes corresponding fake data and a multi-modal feature vector; a determination unit configured to determine the data to be detected as target fake data in a case where a data entry matching the data to be detected is retrieved from the vector index library.
[0014] Optionally, the deepfake detection device further includes an addition unit configured to add the target fake data and its multi-modal feature vector as a new data entry to the vector index library.
[0015] Optionally, each data entry in the vector index library is associated with metadata, which includes at least one of the following metadata of the corresponding forged data: release platform, release time, release account, attribution topic, and dissemination times. The deepfake detection device further includes a second acquisition unit and a construction unit. The second acquisition unit is configured to acquire the metadata of the target forged data; the retrieval unit is further configured to retrieve a preset number of data entries that best match the target forged data from the vector index library; the construction unit is configured to construct forged relationship information of the target forged data according to the metadata of the target forged data and the metadata of the preset number of data entries, where the forged relationship information includes at least one of the following: dissemination link, correlation graph.
[0016] Optionally, the deepfake detection device further includes a classification unit, configured to process the target forged data using a classification network to obtain category information of the target forged data, where the category information includes at least one of the following: forgery category, severity category, and the forgery category includes at least one of the following: face swapping, voice forgery, and false subtitles.
[0017] Optionally, the deepfake detection device further includes a receiving unit, configured to receive forged data collected by a user feedback platform, create a new data entry based on the received forged data, and add it to the vector index library.
[0018] Optionally, the deepfake detection device further includes a feedback unit and an appeal unit. The feedback unit is configured to send the detection result of the target forged data to the user feedback platform for the user feedback platform to output the detection result of the target forged data. The appeal unit is configured to output user appeal information in response to receiving a user appeal request for the target forged data returned by the user feedback platform.
[0019] Optionally, the deepfake detection device further includes a processing unit, configured to, in the case where no data entry matching the data to be detected is retrieved from the vector index library, process the data to be detected using a preset deep reinforcement learning network to obtain action information, where the action information includes at least one of the following actions: detection passed, intercepted, reported, and manual review.
[0020] Optionally, the deepfake detection device further includes a screening unit, configured to screen out data entries that meet a second preset forgery condition from the vector index library to obtain a vector index sub-library of the vector index library, where the second preset forgery condition is a variable condition.
[0021] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: at least one processor; and at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute a deepfake detection method based on multimodal features according to an exemplary embodiment of the present disclosure.
[0022] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, and when instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute a deepfake detection method based on multimodal features according to an exemplary embodiment of the present disclosure.
[0023] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product including computer instructions, and when the computer instructions are run by at least one processor, the at least one processor is caused to execute a deepfake detection method based on multimodal features according to an exemplary embodiment of the present disclosure.
[0024] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: According to the deepfake detection method and the electronic device based on multimodal features of the present disclosure, a multimodal forgery detection scheme is proposed. Compared with forgery detection schemes that rely on a single data modality, such as technical solutions that only use video data or audio data for forgery detection, forgery recognition can be performed through cross-modal features of multiple modalities, effectively integrating multiple data sources, enhancing the comprehensiveness and accuracy of detection, and significantly improving the recognition ability for complex forged content.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0027] Figure 1 is a flowchart of a deepfake detection method based on multimodal features according to an exemplary embodiment of the present disclosure.
[0028] Figure 2 is a logical schematic diagram of a deepfake detection method based on multimodal features according to an exemplary embodiment of the present disclosure.
[0029] Figure 3 is a block diagram of a deepfake detection device based on multimodal features according to an exemplary embodiment of the present disclosure.
[0030] Figure 4It is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0031] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The implementation manners described in the following embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0033] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of the following three parallel situations: "any one of the several items", "a combination of any multiple of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.
[0034] Next, a method and an electronic device for deepfake detection based on multi-modal features according to an exemplary embodiment of the present disclosure will be described in detail with reference to the accompanying drawings.
[0035] Figure 1 It is a flowchart of a method for deepfake detection based on multi-modal features according to an exemplary embodiment of the present disclosure. This method can be executed on an electronic device with sufficient computing power. Figure 2 It is a logical schematic diagram of a method for deepfake detection based on multi-modal features according to an exemplary embodiment of the present disclosure. It should be noted that Figure 2 it is directed to another exemplary embodiment of the present disclosure with more functions, and thus does not correspond exactly to Figure 1 completely.
[0036] Refer to Figure 1 , in step S101, the data to be detected is obtained from the public network.
[0037] The public network includes, for example, but is not limited to, the Internet. Specifically, the data to be detected can be collected from multiple content publishing platforms on the Internet. The content publishing platforms include, for example, but are not limited to, social media platforms, news websites, and personal web pages. Each piece of data to be detected contains all the data that a user publishes at one time on the content publishing platform. In other words, it can be multimodal data including video data, audio data, image data, text data, etc. For example, when a user posts a picture and text update on a social media platform, the image data and text data therein together constitute a piece of data to be detected.
[0038] As an example, the method of the present disclosure can be implemented as an algorithm of a third-party supervision platform to centrally collect multiple pieces of data from the public network, or can be implemented as an electronic device plugin for detecting the data uploaded by the electronic device to the public network. The present disclosure does not limit this.
[0039] In step S102, the multimodal feature vectors extracted from the data to be detected are determined as the vectors to be detected.
[0040] As mentioned above, each piece of data to be detected may contain multiple data of different modalities. By extracting the feature vectors of the corresponding modalities from these data and fusing them, the multimodal feature vectors of the data to be detected can be obtained. As an example, a deep learning model can be first used to extract feature vectors from the data of each modality, and then the feature vectors of each modality are fused. For the extraction of the feature vectors of each modality, for example Figure 2 as shown, for image data and video data, a convolutional neural network (CNN) can be used to extract visual feature vectors, for audio data, a recurrent neural network (RNN) or a dedicated deep network for sound processing can be used to extract acoustic feature vectors, and for text, natural language processing techniques such as BERT (Bidirectional Encoder Representations from Transformers) or other transformer models can be used to extract semantic feature vectors. For the fusion of the feature vectors of each modality, a machine learning algorithm can be used to train a feature fusion model. For example, multimodal fusion techniques such as a model based on an attention mechanism can be applied, such as Figure 2 the Transformer Blocks and the Multilayer Perceptron Head (MLPHead) shown, to identify and compare the associations and conflicts between the feature vectors of different modalities.
[0041] As an example, before extracting the multi-modal feature vectors, the data to be detected can be preprocessed first, such as including but not limited to format standardization, noise removal, and necessary data augmentation, which helps to ensure the stability and effectiveness of subsequent data processing and reduce the degradation or errors in data processing performance caused by differences in data sources. For example, video data and image data can be decoded first and then subjected to format standardization processing such as size adjustment and normalization; audio data can be subjected to format standardization such as unified sampling rate, and silent parts can also be clipped; text data can be subjected to language detection and transcoding. Another example is that for different modal data in the same data to be detected, consistency processing can also be performed in terms of time axis, semantic content, and data format to support cross-modal analysis and joint detection.
[0042] In step S103, based on the vector index library and the vector to be detected, retrieval processing is performed on the data to be detected.
[0043] The vector index library is constructed based on forged data that meets the first preset forgery condition collected from the public network. Each data entry stored in the vector index library includes the corresponding forged data and multi-modal feature vectors. The first preset forgery condition is used to represent the conditions possessed by the forged content that the current deep forgery detection method focuses on. For example, the text or speech contains specified personal names, specified vocabulary, etc., and another example is that the image contains specified faces, specified objects (such as including but not limited to specified buildings), etc. For example Figure 2 As shown, the vector index library can be a distributed vector index library. As an example, expired data entries in the vector index library can be periodically cleared to ensure the timeliness of the data.
[0044] When performing the retrieval processing, specifically, it can be to calculate the similarity between the vector to be detected and the multi-modal feature vectors stored in the vector index library. If the similarity indicates that the two are sufficiently similar, it is considered that a matching data entry is retrieved; otherwise, it is considered that no matching data entry is retrieved.
[0045] Further, multi-modal feature vectors can be randomly selected, and their similarity with the vector to be detected can be calculated. Alternatively, the similarities between multiple multi-modal feature vectors and the vector to be detected can be calculated in batches. The present disclosure does not limit this. If a multi-modal feature vector with a sufficiently high similarity appears during the calculation, the calculation can be stopped, or the calculation can be stopped after calculating the similarities between all multi-modal feature vectors and the vector to be detected. In practice, it can be set as needed, and the present disclosure also does not limit this. The similarity can be represented by calculation methods such as cosine similarity, Euclidean distance, and Manhattan distance. And according to the positive or negative correlation between these similarities and the degree of similarity between the two vectors, thresholds representing "sufficiently similar" can be set respectively. For example, when the cosine similarity is greater than the threshold, the Euclidean distance is less than the threshold, or the Manhattan distance is less than the threshold, it indicates that the two vectors are sufficiently similar, that is, a data entry matching the data to be detected is retrieved. Otherwise, it indicates that the two vectors are not similar. If each multi-modal feature vector in the vector index library is not similar to the vector to be detected, no data entry matching the data to be detected is retrieved.
[0046] In step S104, in the case where a data entry matching the data to be detected is retrieved from the vector index library, the data to be detected is determined to be target forged data.
[0047] It should be understood that the target forged data is the data that is considered to also meet the first preset forgery condition, thereby achieving targeted detection of forged data. In other words, the present disclosure is not intended to detect all AIGC data in the public network (the workload is huge and the actual benefit is extremely low), but only to detect forged data with certain harms, such as forged data that has a negative impact on individuals and society. Such data often has clear text content and / or image content, making the present disclosure have sufficient feasibility.
[0048] According to the multi-modal feature-based deep forgery detection method of the exemplary embodiments of the present disclosure, a multi-modal forgery detection scheme is proposed. Compared with forgery detection schemes that rely on a single data modality, such as technical schemes that only use video data or audio data for forgery detection, forgery recognition can be performed through cross-modal features of multiple modalities, effectively integrating multiple data sources, enhancing the comprehensiveness and accuracy of detection, and significantly improving the recognition ability for complex forged content.
[0049] Next, the multi-modal feature-based deep forgery detection method according to the exemplary embodiments of the present disclosure will be further introduced.
[0050] Optionally, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: adding the target fake data and its multi-modal feature vector as a new data entry to the vector index library. By adding the newly detected target fake data and its multi-modal feature vector to the vector index library, the data in the library can be expanded, and the coverage of the vector index library can be improved. In addition, combining this embodiment with the embodiment of constructing fake relationship information in the following text helps to construct more complete fake relationship information.
[0051] Optionally, each data entry in the vector index library is associated with metadata, and the metadata includes at least one of the following metadata of the corresponding fake data: release platform, release time, release account, affiliated topic, number of transmissions. Among them, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: obtaining the metadata of the target fake data; retrieving a preset number of data entries that best match the target fake data from the vector index library; constructing the fake relationship information of the target fake data according to the metadata of the target fake data and the metadata of the preset number of data entries, where the fake relationship information includes at least one of the following: propagation link, correlation graph. By matching the most similar preset number of data entries from the vector index library and analyzing the relationships between these data entries in combination with the metadata, such as including but not limited to propagation links, correlation graphs, the propagation situation of the currently detected target fake data in the public network can be understood, and the potential sources of the fake data can be traced, realizing effective traceability and providing reliable evidence support for legal litigation. As an example, if the preset number of retrieved data entries are discrete and have no association with each other in terms of metadata, then the data entry with the earliest release time can be determined as the fake source. If the data entry with the earliest release time is also associated with at least one of the later released data entries in terms of metadata, such as including but not limited to being released by the same release account, the release time interval being less than a preset duration, belonging to the same topic, etc., then these associated data entries are determined together as a fake source group.
[0052] Regarding the propagation link, for example, it can be as Figure 2 shown, represented as the original screenshots of multiple data entries arranged in the propagation order, or can be represented as multiple nodes and edges connecting different nodes. Each node represents a data entry, and each edge represents that there is a data propagation relationship between the connected nodes. And the edge can specifically be a directed edge, pointing from the node with an earlier release time to the node with a later release time, to indicate that the fake data is propagated from the node with an earlier release time to the node with a later release time, thus forming a propagation link with a propagation direction. Of course, other reasonable ways can also be used to represent the propagation link, and the present disclosure does not limit this.
[0053] Regarding a correlation graph, for example, it can be represented as multiple nodes and edges connecting different nodes. Each node represents a data entry, and each edge represents the existence of a correlation between the connected nodes. Specifically, different types of lines, thicknesses, and / or colors of the edges can be used to represent different correlation relationships. For example, the correlation relationships published by the same account, the correlation relationships with a time interval between publications less than a preset duration, and the correlation relationships belonging to the same topic can be represented distinguishably. Of course, other reasonable ways can also be adopted to represent the correlation graph, and the present disclosure places no restrictions thereon.
[0054] Optionally, the deepfake detection method based on multi-modal features according to an exemplary embodiment of the present disclosure further includes: processing target forged data using a classification network to obtain category information of the target forged data, where the category information includes at least one of the following: forgery category, severity category. The forgery category includes at least one of the following: face swapping, voice forgery, false subtitles. By classifying the target forged data using the classification network, richer information can be provided for the detected target forged data, facilitating further processing of the target forged data subsequently. As an example, the severity category can comprehensively judge the severity of the target forged data according to dimensions such as misleadingness, dissemination potential, and sensitivity level, and be divided into multiple severity levels (for example, four severity levels from L1 to L4), and this can be used to drive subsequent adjustment of the warning level and traceability intensity. As an example, each type of category information can be obtained by processing with a separate classification network. The classification network can be trained using a supervised training method and can achieve incremental training as the sample data increases. As an example, the classification network can adopt a multi-modal classification network structure, including a visual feature recognition model Swin Transformer, an audio forgery detection network wav2vec2, and a text recognition model BERT.
[0055] Optionally, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: receiving forged data collected by a user feedback platform, creating a new data entry based on the received forged data, and adding it to the vector index library. By providing a user feedback platform, forged data reported by users can be received, providing more data sources for the construction and expansion of the vector index library. As an example, for the forged data reported by the user feedback platform, screening can be performed first, and then the screened forged data can be used for the expansion of the vector index library. For example, a large language model (LLM) can be used to analyze the content of the reported forged data, extract forged clues and context semantics, and then the preliminary screening of the forged data can be directly performed according to the extracted content. The remaining content after the preliminary screening can be handed over to humans for further screening, or the extracted content can be used as auxiliary information to provide reference for human screening. In addition, for the embodiment of classifying the target forged data introduced above, as an example, the category information of the screened forged data can also be obtained by manual analysis, or the category information of the screened forged data can be obtained semi-automatically by humans in combination with the model. Then, these forged data and their category information can be used as samples for the training of the classification network.
[0056] Optionally, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: sending the detection result of the target forged data to the user feedback platform for the user feedback platform to output the detection result of the target forged data, and in response to receiving a user appeal request for the target forged data returned by the user feedback platform, outputting user appeal information. By outputting the detection result of the target forged data on the user feedback platform, that is, the data to be detected is detected as the target forged data, and providing a user appeal function, user feedback can be received, reducing the adverse effects caused by misjudging forged data and serving as a basis for optimizing the detection method. For example, the threshold of the similarity used in step S103 can be adjusted accordingly, which helps to optimize the detection performance.
[0057] Optionally, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: in the case where no data entry matching the data to be detected is retrieved from the vector index library, using a preset deep reinforcement learning network to process the data to be detected to obtain action information, where the action information includes at least one of the following actions: detection passed, intercepted, reported, and manual review. By additionally configuring a preset deep reinforcement learning network to perform a fallback check on the data to be detected for which no matching data entry is retrieved, that is, to detect harmful forged content that has not yet appeared in the vector index library, the recognition accuracy of new harmful forged content can be improved, the risk of missed detection can be reduced, and the detection performance can be enhanced. In addition, since the deep reinforcement learning network (Deep Q-Network, abbreviated as DQN) can use an agent to make action decisions and store the agent's experience in the replay buffer through the experience replay mechanism for multiple learning, with specific dynamic adaptive learning capabilities, it can also dynamically optimize the preset deep reinforcement learning network using new harmful forged content, improve data utilization, achieve efficient learning, effectively improve the detection accuracy of new harmful forged data, and comprehensively enhance the detection performance. It should be understood that among the above several actions, "detection passed" means that the data to be detected is not harmful forged data, while "intercepted" and "reported" mean that the data to be detected is harmful forged data, but the processing methods are different. "Intercepted" means directly determining it as harmful forged data, "reported" means further sending a notice for the convenience of staff to make other processing, and "manual review" means that it is not yet clear whether the data to be detected is harmful forged data and needs to be handed over to manual review. As an example, for the data that is intercepted, reported, and determined to be harmful forged data after manual review, it can also be used for the expansion of the vector index library. Further, as an example, for the data to be detected for which no matching data entry is retrieved, a preliminary screening model (such as a large language model) can be used to perform a preliminary screening on it to screen out the data that may be harmful forged content, and then the screened data can be processed using a preset deep reinforcement learning network to obtain action information, and the present disclosure does not limit this.
[0058] Regarding the preset deep reinforcement learning network, specifically, first, define several basic concepts.
[0059] : The system state at time t, containing information of all modalities.
[0060] : The action executed by the agent at time t, marking whether a multi-modal input is forged.
[0061] : After executing the action The reward received.
[0062] R: The cumulative reward obtained by the agent over a period of time.
[0063] Reward function for multi-modal forgery detection The goal is to maximize the long-term reward, which is typically achieved by optimizing the following objective function:
[0064] In the above formula, γ is the discount factor used to balance the importance of immediate and future rewards, and T is the total number of time steps or decision points. The definitions of rewards and penalties are as follows.
[0065] Correctly identifying a forgery (TP: True Positive):
[0066] Correctly identifying a genuine item (TN: True Negative):
[0067] Misclassifying a genuine item as a forgery (FP: False Positive):
[0068] Failing to identify a forgery (FN: False Negative):
[0069] Then, the probability of selecting action a given state s is defined using the policy π:
[0070] The goal of adaptive learning is to find a policy to maximize the expected long-term reward:
[0071] A policy gradient algorithm or value function approximation method is used to update the policy or value function. For multi-modal input, a deep neural network is used to fuse the features of various modalities and output the forgery probability or action value. Here, the deep DQN network is used, and the update rule is as follows:
[0072] where α is the learning rate. Through such a design and formula, the multi-modal forgery detection system can utilize information from various sources and make accurate forgery detection decisions in complex forgery data.
[0073] Optionally, the multi-modal feature-based deepfake detection method according to an exemplary embodiment of the present disclosure further includes: screening out data entries that meet the second preset forgery condition from the vector index library to obtain a vector index sub-library of the vector index library, where the second preset forgery condition is a variable condition. Since the themes of harmful forged content that receive high attention in different periods may vary, based on this, by configuring the variable second preset forgery condition, a vector index sub-library can be further screened out from the vector index library, so as to centrally store data entries related to the current hot topics of concern. When using this vector index sub-library to retrieve the data to be detected, the data range for retrieval processing can be narrowed, which helps to improve the retrieval accuracy and retrieval accuracy. As an example, when retrieving, the vector index sub-library can be targeted first, or the vector index library can be targeted first, and the present disclosure does not limit this. It should be understood that when the second preset forgery condition changes, it is necessary to re-screen the vector index sub-library accordingly. In addition, as an example, during the period when the second preset forgery condition does not change, other update conditions can also be configured, such as configuring an update period, to update the vector index sub-library, and the present disclosure does not limit this either.
[0074] Figure 3 is a block diagram of a multi-modal feature-based deepfake detection device according to an exemplary embodiment of the present disclosure. Referring to Figure 3 FIG. 3, the multi-modal feature-based deepfake detection device 300 includes a first acquisition unit 301, an extraction unit 302, a retrieval unit 303, and a determination unit 304.
[0075] The first acquisition unit 301 can acquire the data to be detected from the public network.
[0076] The extraction unit 302 can determine the multi-modal feature vector extracted from the data to be detected as the vector to be detected.
[0077] The retrieval unit 303 can perform retrieval processing on the data to be detected based on the vector index library and the vector to be detected, where the vector index library is constructed based on forged data that meets the first preset forgery condition collected from the public network, and each data entry stored in the vector index library includes the corresponding forged data and multi-modal feature vector.
[0078] The determination unit 304 can determine that the data to be detected is target forged data when a data entry matching the data to be detected is retrieved from the vector index library.
[0079] Optionally, the multi-modal feature-based deepfake detection device 300 according to an exemplary embodiment of the present disclosure further includes an addition unit (not shown in the figure), which can add the target forged data and its multi-modal feature vector as a new data entry to the vector index library.
[0080] Optionally, each data entry in the vector index library is associated with metadata, which includes at least one of the following metadata of the corresponding forged data: release platform, release time, release account, attribution topic, and dissemination times. The deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a second acquisition unit and a construction unit (not shown in the figure). The second acquisition unit can acquire the metadata of the target forged data; the retrieval unit 303 can also retrieve a preset number of data entries that best match the target forged data from the vector index library; the construction unit can construct forged relationship information of the target forged data according to the metadata of the target forged data and the metadata of the preset number of data entries, where the forged relationship information includes at least one of the following: dissemination link, correlation graph.
[0081] Optionally, the deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a classification unit (not shown in the figure), which can process the target forged data using a classification network to obtain category information of the target forged data, where the category information includes at least one of the following: forgery category, severity category, and the forgery category includes at least one of the following: face swap, voice forgery, false caption.
[0082] Optionally, the deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a receiving unit (not shown in the figure), which can receive the forged data collected by the user feedback platform, create a new data entry based on the received forged data, and add it to the vector index library.
[0083] Optionally, the deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a feedback unit and an appeal unit (not shown in the figure). The feedback unit can send the detection result of the target forged data to the user feedback platform for the user feedback platform to output the detection result of the target forged data. The appeal unit can output user appeal information in response to receiving a user appeal request for the target forged data returned by the user feedback platform.
[0084] Optionally, the deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a processing unit (not shown in the figure), which can use a preset deep reinforcement learning network to process the data to be detected to obtain action information in the case where no data entry matching the data to be detected is retrieved from the vector index library, where the action information includes at least one of the following actions: detection passed, intercepted, reported, manual review.
[0085] Optionally, the deepfake detection device 300 based on multimodal features according to an exemplary embodiment of the present disclosure further includes a screening unit (not shown in the figure), which can screen out data entries that meet the second preset forgery condition from the vector index library to obtain a vector index sub-library of the vector index library, where the second preset forgery condition is a variable condition.
[0086] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0087] Figure 4 The structural block diagram of an electronic device 400 according to an exemplary embodiment of the present disclosure is shown.
[0088] Referring to Figure 4 , the electronic device 400 includes: at least one memory 401 and at least one processor 402. Computer-executable instructions are stored in the at least one memory 401. When the computer-executable instructions are run by the at least one processor 402, the at least one processor is caused to execute the deepfake detection method based on multimodal features as described in the above exemplary embodiments.
[0089] As an example, the electronic device 400 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 400 does not have to be a single electronic device 400, but can also be any assembly of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 400 can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device 400 that is interconnected locally or remotely (e.g., via wireless transmission).
[0090] In the electronic device 400, the processor 402 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example and not a limitation, the processor 402 can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0091] The processor 402 can run the instructions or code stored in the memory 401, where the memory 401 can also store data. The instructions and data can also be sent and received via a network interface device through the network, where the network interface device can adopt any known transmission protocol.
[0092] The memory 401 can be integrated with the processor 402. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 401 can include independent devices, such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 401 and the processor 402 can be operatively coupled or can communicate with each other, for example, through I / O ports, network connections, etc., so that the processor 402 can read files stored in the memory.
[0093] In addition, the electronic device 400 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 400 can be connected to each other via a bus and / or a network.
[0094] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions can also be provided. When the instructions are run by at least one processor, the at least one processor is caused to execute the multi-modal feature-based deepfake detection method as described in the above exemplary embodiment. Examples of such computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as a multimedia card, a Secure Digital (SD) card, or an eXtreme Digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device that is configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or a computer so that the processor or the computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system so that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0095] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including computer instructions that, when executed by at least one processor, perform the method for deepfake detection based on multimodal features as described in the above exemplary embodiment.
[0096] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
[0097] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A deepfake detection method based on multimodal features, characterized in that, The deepfake detection method includes: Obtaining data to be detected from a public network; Determining the multi-modal feature vectors extracted from the data to be detected as the vectors to be detected; Based on a vector index library and the vectors to be detected, performing a retrieval process on the data to be detected, where the vector index library is constructed based on forged data that meets a first preset forgery condition collected from a public network, and each data entry stored in the vector index library includes corresponding forged data and multi-modal feature vectors; In the case where a data entry matching the data to be detected is retrieved from the vector index library, determining the data to be detected as target forged data.
2. The deepfake detection method according to claim 1, wherein, The deepfake detection method further includes: Adding the target forged data and its multi-modal feature vectors as a new data entry to the vector index library.
3. The deepfake detection method according to claim 1, wherein Each data entry in the vector index library is associated with metadata, and the metadata includes at least one of the following metadata of the corresponding forged data: release platform, release time, release account, attribution topic, number of dissemination times, where the deepfake detection method further includes: Obtaining the metadata of the target forged data; Retrieving a preset number of data entries from the vector index library that best match the target forged data; Constructing forgery relationship information of the target forged data according to the metadata of the target forged data and the metadata of the preset number of data entries, where the forgery relationship information includes at least one of the following: dissemination link, correlation graph.
4. The deepfake detection method according to claim 1, wherein, The deepfake detection method further includes: Processing the target forged data using a classification network to obtain category information of the target forged data, where the category information includes at least one of the following: forgery category, severity category, and the forgery category includes at least one of the following: face swapping, voice forgery, false subtitles.
5. The deepfake detection method according to claim 1, wherein The deepfake detection method further includes: Receiving forged data collected by a user feedback platform, creating a new data entry based on the received forged data, and adding it to the vector index library; and / or Sending the detection result of the target forged data to the user feedback platform for the user feedback platform to output the detection result of the target forged data, and outputting user appeal information in response to receiving a user appeal request for the target forged data returned by the user feedback platform.
6. The deepfake detection method according to any one of claims 1 to 5, characterized in that The deepfake detection method further includes: In the case where a data entry matching the data to be detected is not retrieved from the vector index library, processing the data to be detected using a preset deep reinforcement learning network to obtain action information, where the action information includes at least one of the following actions: detection passed, intercepted, reported, manual review.
7. The deepfake detection method according to any one of claims 1 to 5, characterized in that, The deepfake detection method further includes: Filtering out data entries that meet a second preset forgery condition from the vector index library to obtain a vector index sub-library of the vector index library, where the second preset forgery condition is a variable condition.
8. An electronic device, characterized in that, Including: At least one processor; At least one memory storing computer-executable instructions Wherein, when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the multi-modal feature-based deepfake detection method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are run by at least one processor, the at least one processor is caused to execute the multi-modal feature-based deepfake detection method according to any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that, When the computer instructions are run by at least one processor, the at least one processor is caused to execute the multi-modal feature-based deepfake detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Dual-engine-driven multi-modal data retrieval method, device and system
CN115455249A
Forgery video detection method and device, electronic equipment and storage medium
CN118038155A
Deep counterfeit video detection method and device fusing multi-modal information
CN119251738A
Information forgery detection method, computer equipment, storage medium and program product
CN119397255A
Anti-fraud suspicious list association analysis method and analysis device thereof comprising a storage module and a processing module for obtaining suspicious transaction information
TW202505453A
Cited By
Depth counterfeit image identification method based on differential feature search
CN121583008A