Face Deepfake Forensics Method, System, Device and Storage Medium

By embedding the robustness and semi-fragile watermarks of multi-level regional semantic features in the face image, the lack of performance of the existing technology in deep forgery and evidence collection in an open environment is solved, and the blind detection and identity traceability of high generalization is achieved, which significantly improves the detection performance.

CN119723689BActive Publication Date: 2025-05-27UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510234677.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-27
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing deep forgery defense methods have poor performance when dealing with unknown forgery technologies, and cannot be applied to deep forgery forensic tasks in open environments, and cannot achieve evidence collection in blind detection scenarios.

Method used

By embedding the robust watermark and semi-fragile watermark of multi-level regional semantic features in the face image, the two watermarks are extracted and cross-compared by using the characteristics that are difficult to maintain intact watermarks, to determine whether the image has been deeply forged and to realize identity traceability.

Benefits of technology

It realizes blind detection with high generalization, significantly improves the traceability and detection performance of deep forgery, and the experimental results on multiple data sets have reached the leading level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723689B_ABST
    Figure CN119723689B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and storage medium for face deepfake forensics, which are corresponding solutions. In the solutions: multi-level regional semantic features of the face are extracted, and a robust watermark and a semi-fragile watermark are adaptively embedded into different semantic regions, and finally a watermarked image is reconstructed and then can be released externally; moreover, by using the characteristic that it is difficult to maintain the integrity of the watermark during the deepfake process, two kinds of watermark information are extracted from the obtained watermarked image. On the one hand, identity tracing can be realized, and on the other hand, deepfake detection can be achieved through cross-comparison; the above solution provided by the present invention is a detection solution that actively traces the source and is independent of the specific deepfake method, which can significantly improve the tracing and detection performance of deepfakes, and the experimental results on multiple data sets all show that it has reached the leading level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face deepfake forensics, and in particular to a method, system, device and storage medium for face deepfake forensics. Background Art

[0002] The face deepfake technology has been widely used in the entertainment field with its powerful generation ability, such as film and television production, game character modeling, and digital special effects. However, while bringing innovation and convenience, this technology also has the risk of being misused, resulting in some non-negligible negative impacts. Especially without authorization, criminals may use deepfake technology to forge the facial images and videos of public figures or ordinary citizens. In addition, the abuse of deepfake technology provides technical support for some illegal acts, further weakening people's trust in the authenticity of images and videos. Therefore, how to effectively identify and defend against the malicious applications of face deepfake technology has become an urgent problem to be solved.

[0003] Existing deepfake defenses mainly adopt passive detection methods, regarding it as a binary classification task, and learning forgery features through convolutional neural networks and outputting classification results. However, these methods perform poorly when dealing with unknown forgery technologies and are not applicable to deepfake forensics tasks in open environments. One effective solution is to embed traceable or auxiliary information in images through digital watermarking technology. Existing digital watermarking technologies usually focus on the robustness of watermarks, requiring the embedded watermarks to remain stable under various image transformations and distortions. When implementing deepfake detection, robust watermarks are further extended to semi-fragile watermarks, making them resistant to common image distortions while being sensitive to various deepfake methods, so as to achieve detection by comparing the decoded watermark with the original embedded watermark. However, this detection method requires prior knowledge of the original embedded watermark information and cannot achieve the forensics task in blind detection scenarios.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The object of the present invention is to provide a method, system, device and storage medium for face deepfake forensics, which can obtain different types of watermarks through notification decoding, and determine whether the image has been processed by deepfake through cross-comparison, so as to realize the verification of image authenticity. At the same time, the identity traceability can be realized by using the decoded watermark. Compared with the existing solutions, high-generalization blind detection can be achieved.

[0006] The object of the present invention is achieved by the following technical solutions:

[0007] A method for face deepfake forensics includes:

[0008] Watermark addition stage: Receive the face image to be protected and the binary sequence, perform multi-level semantic extraction and regional semantic segmentation on the face image respectively, and synthesize them into multi-level regional semantic features; for each level of regional semantic features, adaptively generate one of the two types of set watermarks based on the binary sequence, and embed it into the corresponding level of regional semantic features, and reconstruct the watermark image by synthesizing all the regional semantic features with embedded watermarks;

[0009] Watermark extraction and verification stage: Extract the two types of watermarks from the input watermark image respectively, trace the source through the proposed watermark, and judge whether the input watermark image has been deepfaked by cross-comparison.

[0010] A face deepfake forensics system, including: a face deepfake forensics model; implementing the foregoing method based on the face deepfake forensics model, and the face deepfake forensics model includes:

[0011] A watermark addition module, configured to receive the face image to be protected and the binary sequence, perform multi-level semantic extraction and regional semantic segmentation on the face image respectively, and synthesize them into multi-level regional semantic features; for each level of regional semantic features, adaptively generate one of the two types of set watermarks based on the binary sequence, and embed it into the corresponding level of regional semantic features, and reconstruct the watermark image by synthesizing all the regional semantic features with embedded watermarks;

[0012] A watermark extraction and verification module, configured to extract the two types of watermarks from the input watermark image respectively, trace the source through the proposed watermark, and judge whether the input watermark image has been deepfaked by cross-comparison.

[0013] A processing device, including: one or more processors; a memory for storing one or more programs;

[0014] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the foregoing method.

[0015] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.

[0016] As can be seen from the technical solution provided by the present invention above, multi-level regional semantic features of the face are extracted, and different types of watermarks (i.e., robust watermarks and semi-fragile watermarks) are adaptively embedded into different semantic regions, and finally the watermarked image is reconstructed and then can be published outward; moreover, by using the characteristic that it is difficult to maintain the integrity of the watermark during the deepfake process, two types of watermark information are extracted from the obtained watermarked image. On the one hand, identity tracing can be achieved, and on the other hand, deepfake detection can be achieved through cross-comparison; the above solution provided by the present invention is an active tracing and detection solution that is independent of specific deepfake methods, which can significantly improve the tracing and detection performance of deepfakes, and the experimental results on multiple data sets all show that it has reached the leading level. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0018] Figure 1 It is a flowchart of a face deepfake forensics method provided by an embodiment of the present invention;

[0019] Figure 2 It is a schematic diagram of the visualization result of the generated watermarked image provided by an embodiment of the present invention;

[0020] Figure 3 It is a schematic diagram of the training framework of a face deepfake forensics model provided by an embodiment of the present invention;

[0021] Figure 4 It is a schematic diagram of a processing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] First, the terms that may be used in this article are described as follows:

[0024] Descriptions with terms such as "including", "comprising", "containing", "having" or other similar semantics shall be construed as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction condition, processing condition, parameter, algorithm, signal, data, product or article, etc.) shall be construed as not only including the explicitly listed technical feature element, but also including other technical feature elements well-known in the art that are not explicitly listed.

[0025] The term "consisting of" means excluding any technical feature element that is not explicitly listed. If this term is used in a claim, it will make the claim a closed type, so that it does not include technical feature elements other than the explicitly listed ones, except for conventional impurities related thereto. If this term only appears in a sub-clause of a claim, then it only limits the elements explicitly listed in that sub-clause, and the elements recorded in other sub-clauses are not excluded from the overall claim.

[0026] The following provides a detailed description of a face deepfake forensics method, system, device and storage medium provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well-known to those skilled in the art. Conditions not specified in the embodiments of the present invention are carried out according to the conventional conditions in the art or the conditions recommended by the manufacturer. Reagents or instruments not specified in the embodiments of the present invention as to the manufacturer are all conventional products that can be obtained by commercial purchase.

[0027] Embodiment 1

[0028] The embodiment of the present invention provides a face deepfake forensics method, as Figure 1 shown, which mainly includes:

[0029] (1) Watermark addition stage.

[0030] The main process of this stage includes: receiving a face image to be protected and a binary sequence, performing multi-level semantic extraction and regional semantic segmentation on the face image respectively, and synthesizing them into multi-level regional semantic features; for each level of regional semantic features, adaptively generating one of two types of set watermarks based on the binary sequence, and embedding it into the corresponding level of regional semantic features, and reconstructing a watermarked image by synthesizing all the regional semantic features with embedded watermarks.

[0031] In the embodiments of the present invention, the two types of watermarks set include: robust watermarks and semi-fragile watermarks; corresponding types of watermarks are generated through a diffusion process based on the binary sequence; the diffusion process includes: a fully connected neural network processing process and an upsampling process; the adaptively generated watermark is spliced with the corresponding level regional semantic features and fused through an attention mechanism to obtain the regional semantic features embedded with the watermark; all the regional semantic features embedded with the watermark are synthesized, and a watermark image is reconstructed through a face reconstruction network.

[0032] (2) Watermark extraction and verification stage.

[0033] The main processes in this stage include: extracting the two types of watermarks from the input watermark image respectively, tracing the source through the proposed watermark, and judging whether the input watermark image has been deepfaked through cross-comparison.

[0034] In the embodiments of the present invention, the source can be traced through the robust watermark; the bit error rate difference between the robust watermark and the semi-fragile watermark can be obtained through cross-comparison, and thus it can be judged whether the input watermark image has been deepfaked: when the bit error rate difference is less than the set decision threshold, it is determined that the input watermark image is a real image, otherwise, it is determined that the input watermark image has been deepfaked.

[0035] Preferably, the watermark addition stage is implemented through a watermark addition module, and the watermark extraction and verification stage is implemented through a watermark extraction and verification module. The two together constitute a face deepfake forensics model; the face deepfake forensics model is pre-trained, and a progressive noise layer and a perturbation adversarial generation network are introduced during the training process; the training process is divided into two parts, the first part is called the early training period, and the second part is called the late training period; in the early training period, a predefined number of image transformations are randomly selected as noise through the progressive noise layer to apply image perturbations to the watermark image; in the late training period, an image transformation operation is generated through the perturbation adversarial generation network to apply image perturbations to the watermark image, and it is alternately trained with the subsequent watermark extraction and verification module in an adversarial training manner.

[0036] From the perspective of specific implementation: the watermark addition stage mainly processes the face image provided by the user to be protected to obtain a watermark image, which has embedded watermark information and can achieve accurate source tracing. Therefore, the watermark image can be published to the network side; when detection is required, the watermark image can be obtained from the network side, and then the two types of watermarks are extracted, and thus identity tracing and deepfake detection can be realized.

[0037] In the specific implementation process, the above-mentioned face deepfake forensics model can be deployed on a computer or server, which can automatically embed watermarks into face images and determine whether the images have been forged. It can be widely applied to various social platforms, such as short video platforms, photo sharing websites, etc., to provide effective content protection and authenticity verification functions.

[0038] To more clearly demonstrate the technical solutions provided by the present invention and the resulting technical effects, the methods provided by the embodiments of the present invention will be described in detail below with specific examples.

[0039] I. Overall overview and effect description of the solution.

[0040] The embodiments of the present invention provide a face deepfake forensics method, which is based on multi-level facial semantic watermarking technology. Different from the traditional method of passively detecting the generated forged content, the present invention adopts an active strategy to pre-embed watermark information in the original image to achieve accurate source tracing. By utilizing the characteristic that it is difficult to maintain the integrity of the watermark during the deepfake process, the present invention demonstrates a highly generalized deepfake detection ability.

[0041] In the embodiments of the present invention, a facial feature decoupling and reconstruction architecture is introduced into the watermark addition module, aiming to provide an active traceability and a detection strategy independent of specific deepfake methods. In this architecture, the present invention successfully embeds watermarks with different characteristics, namely robustness and semi-fragility, into different levels of semantic regions. The embedding of the robust watermark enables it to resist common distortion operations and deepfake tampering, ensuring reliable traceability verification and tracking. Once the facial semantics are tampered with, the semi-fragile watermark therein will be damaged, thus quickly identifying the forgery traces of the image. After training and testing on the high-definition face datasets CelebA and CelebA-HQ, the present invention demonstrates excellent performance. In the traceability and robustness tests, the bit error rate of the robust watermark is only 0.0476%, and the bit error rate of the semi-fragile watermark is also only 0.0982%. In the deepfake perturbation test, the bit error rate of the robust watermark performs excellently, and the detection success rate using the semi-fragile watermark is as high as 98.31%. In addition, Figure 2 Intuitively shows the watermark image effect generated by the present invention, where the watermark information is almost invisible, fully demonstrating the high-quality performance of the present invention in watermark generation. Figure 2Among them, the first five columns are the results of the present invention. The first column is JPEG compression with a compression quality of 50; the second column is Gaussian blur with a Gaussian kernel standard deviation of 2 and a Gaussian kernel size of 3; the third column is median filtering with a filter kernel size of 3; the fourth column is salt-and-pepper noise with a noise probability of 0.1; the fifth column is Gaussian noise with a mean of 0 and a standard deviation of 0.1. After that, the 6th to 12th columns are the results of existing solutions. SimSwap is a high-fidelity image face-swapping model, StarGAN is a face attribute editing based on a multi-domain transfer generative adversarial network, DiffFace is an image face-swapping model based on a diffusion model, HFGI is a face attribute editing based on the inverse mapping of a high-fidelity generative adversarial network, StyleMask is a facial manipulation based on the inversion of a stylized generative adversarial network, E4S is a fine-grained face-swapping model based on the inversion of a regional generative adversarial network, and InfoSwap is an image face-swapping model based on information bottleneck decoupling.

[0042] II. Detailed Introduction of the Solution.

[0043] 1. Watermark Addition Stage.

[0044] In this stage, regional semantic watermark injection is mainly carried out. Specifically: through a multi-level semantic encoding network for decoupling and semantic region segmentation to obtain levels of regional semantic attributes ; further, use a watermark encoder to adaptively embed the binary sequence into each level of semantic attributes; finally, use a face reconstruction network to output the watermark image. As mentioned before, this stage can be implemented through a watermark addition module, and the specific design of the internal modules is as follows:

[0045] (1) Multi-level Semantic Encoding Network.

[0046] Through the multi-level semantic encoding network extract multi-level semantics (such as facial contour, hair color, expression attributes, identity information, etc.) of the face image to be protected, expressed as:

[0047] ;

[0048] where, is the face image to be protected, is the semantic feature of the th level, is the number of levels.

[0049] Exemplarily, a feature extraction network with a U-Net architecture can be used as the multi-level semantic encoding network, and multi-level semantic features can be learned in an end-to-end manner to obtain groups of semantic feature maps with different resolution levels. The resolution corresponding to the th level is , and Let \(H\) and \(W\) be the height and width, and \(C\) be the number of channels. For example, when the face image is an RGB (Red, Green, Blue channels) image, \(C = 3\).

[0050] (2) Semantic Region Segmentation Module.

[0051] Perform region semantic segmentation on the face image to be protected through a facial region semantic segmentation network as follows:

[0052] ;

[0053] where is the \(i\)-th semantic region mask, \(N\) is the number of semantic region masks obtained by region semantic segmentation (for example, \(N = 12\)), , is the symbol of the set of real numbers, , correspond to the height and width of the input image respectively. The segmentation results at each level represent semantic regions such as background, hair, nose, mouth, eyes, and facial regions.

[0054] Exemplarily, the facial region semantic segmentation network can be composed of a bilateral semantic segmentation network (BiSeNet) pre-trained on the CelebAMask-HQ dataset.

[0055] By combining the outputs of the above two modules, regional semantic features are obtained. Specifically: for each level, they are combined into regional semantic features through the following formula:

[0056] ;

[0057] where is the regional semantic feature at the \(i\)-th level, which contains \(N\) semantic region masks, is the Hadamard product, means downsampling to the corresponding size.

[0058] (3) Watermark Encoder.

[0059] In the implementation of the present invention, to support the embedding of robust watermarks and semi-fragile watermarks, two independent watermark encoders are used. First, the binary sequence to be embedded is diffused to the same dimension as \(k\) semantic attribute feature maps through a fully connected neural network and an upsampling process ; Then, the diffused binary sequences are respectively concatenated with the regional semantic features at each level and fused through an attention mechanism, and finally k semantic features embedded with watermarks are output.

[0060] Among them, the watermark embedding process of the regional semantic features at each level is expressed as:

[0061] ;

[0062] Among them, is the regional semantic feature at the th level, represents the watermark generated based on the binary sequence through the diffusion process , which is a robust watermark or a semi-fragile watermark. The symbol represents concatenation, represents the attention convolution module with an attention mechanism, represents the semantic feature embedded with a watermark at the th level.

[0063] In the embodiment of the present invention, the watermark is adaptively embedded into the regional semantic features at each level. That is to say, it does not explicitly distinguish which type of watermark should be embedded in the regional semantic features at each level, but cooperates with the subsequent decoding process to separate the two types of watermarks and adopts end-to-end overall training, so that the model can automatically learn the specific type of watermark that should be embedded in the regional semantic features at each level.

[0064] (4) Facial reconstruction network.

[0065] In the facial reconstruction network, the k semantic features embedded with watermarks are successively input into the generation network, and semantic fusion is realized through an Adaptive Instance Normalization (AdaIN) module at each level. Finally, the watermark image is gradually output.

[0066] 2. Watermark extraction and verification phase.

[0067] Similarly, this phase can be realized through a watermark verification and extraction module. Since the watermark is embedded in a learnable manner, a depth-separable watermark decoder and are configured in the watermark verification and extraction module to respectively extract the robust watermark and the semi-fragile watermark, and the process is as follows:

[0068] ;

[0069] ;

[0070] Among them, the robust watermark is robust to both conventional image perturbations and malicious deep fakes, while the semi-fragile watermark is robust to conventional image perturbations but vulnerable to malicious deep fakes; is the watermarked image input to the watermark verification and extraction module, indicating that the watermarked image has been perturbed.

[0071] After obtaining the two types of watermarks, the robust watermark among them can be used for traceability; by comparing the difference in the bit error rates of the robust watermark and the semi-fragile watermark , it can be verified whether the image has been deep faked. The principle is as follows: for a real image, its error expectation ; for a forged image, the error expectation . That is, the average semi-fragile bit error rates of real images and forged images respectively follow the following distributions:

[0072] ;

[0073] ;

[0074] Among them, p(.) represents the probability value of the expression in the parentheses, c represents the category of the image to be detected, represents that the input image is a real image, represents that the input image is a forged image, represents the variance corresponding to the normal distribution N(,).

[0075] Therefore, the selection of the decision threshold needs to satisfy the maximum likelihood:

[0076] ;

[0077] Among them, is the watermarked image input to the watermark verification and extraction module, that is, the mentioned above. It has undergone unknown perturbations during the Internet transmission process (that is, it may have undergone conventional image perturbations or malicious forgeries during the transmission process).

[0078] According to the distributions of the average semi-fragile bit error rates of real images and forged images, it can be seen that their variances are the same, while the means are 0 and 0.5 respectively, that is, when it can satisfy the above maximum likelihood.

[0079] That is, when the difference between the robust and semi-fragile bit error rates is less than the decision threshold When it is judged as a real image (for example, 0.25), that is, the generated watermark image only undergoes conventional image perturbations during the propagation process and does not experience deep forgery, otherwise it is a deep forgery image. Thus, it can be verified whether the image has been maliciously tampered with by deep forgery technology.

[0080] III. Model training scheme.

[0081] The above watermark addition module and watermark extraction and verification module together constitute a face deep forgery forensics model. The face deep forgery forensics model is pre-trained. In order to achieve the robustness of the robust watermark and semi-fragile watermark against conventional image perturbations, a progressive noise layer and a perturbation adversarial generation network are introduced during the training process. Moreover, an adversarial discriminator is introduced to respectively discriminate the input face image to be protected and the watermark image, as Figure 3 shown.

[0082] The training process is divided into two parts, the first part is called the early training stage, and the second part is called the late training stage; in the early training stage, to keep the training process stable, a predefined number of image transformations are randomly selected as noise through the progressive noise layer, and image perturbations are applied to the watermark image; exemplarily, 4 types of image transformations can be selected as noise, including Gaussian blur, Gaussian noise, median filtering, and JPEG (Joint Photographic Experts Group) compression. In the late training stage, more challenging image transformation operations are generated through the perturbation adversarial generation network (Generative Adversarial Network, GAN), image perturbations are applied to the watermark image, and it is alternately trained with the watermark extraction and verification module in an adversarial training manner, further enhancing the robustness of the watermark against unknown image distortions.

[0083] The training loss is expressed as:

[0084] ;

[0085] where is the training loss, , , correspondingly represent the watermark decoding loss term, the perceptual loss term of the watermark image, and the adversarial loss term, , and are the weights of the corresponding loss terms.

[0086] The watermark decoding loss term is obtained by calculating the differences between the two types of watermarks extracted and the binary sequence respectively, and is expressed as:

[0087] ;

[0088] Among them, represents the mean square error function.

[0089] The perceptual loss term of the watermark image is obtained by calculating the difference between the face image to be protected and the watermark image, and is expressed as:

[0090] ;

[0091] Among them, is the symbol of the 2-norm, LPIPS is the perceptual loss function, is the watermark image.

[0092] The adversarial loss term is calculated through the discrimination result output by the introduced adversarial discriminator, and is expressed as:

[0093] ;

[0094] Among them, is the adversarial discriminator.

[0095] Figure 3 In , -1 is the watermark decoder and the internal decoding process. For there are three types of situations in the decoding process. In training, these three situations are simulated to achieve complete end-to-end training. These three situations are as follows:

[0096] (1) Passing through regular perturbations: The semi-fragile watermark is robust to regular image perturbations. Therefore, the watermark exists after semi-fragile decoding, that is, the BER should tend to 0. BER is the bit error rate, which refers to the bit error rate of the decoded binary sequence and the given binary sequence for the watermark (that is, the ratio of the number of wrongly predicted bits to the total length).

[0097] (2) Passing through malicious tampering: The semi-fragile watermark is vulnerable to malicious tampering perturbations. Therefore, the watermark does not exist after semi-fragile decoding, that is, the BER should tend to 0.5.

[0098] (3) Passing through various perturbations (including two small situations of regular perturbations and malicious tampering): The robust watermark is robust to various perturbations. Therefore, the watermark exists after passing through the robust decoder, that is, the BER should tend to 0.

[0099] In terms of training details: A multi-level GAN structure is selected to enhance the discriminability of the detail area. The model in the present invention is trained on a single GPU card, and 32 face images are input at a time. The Adam (Adaptive Moment Estimation) optimizer is used for optimization. The initial learning rate of the model is set to 0.001, and the initial learning rate of the adversarial discriminator is set to 0.0001. For more sufficient training, the learning rate is adjusted to 1×10 -8 after 100 epochs by means of cosine annealing decay of the learning rate. The related training process involved can refer to the conventional technology, and the present invention will not elaborate.

[0100] The above solution provided by the embodiment of the present invention extracts multi-level facial semantics in an end-to-end manner, and further adaptively embeds a robust watermark and a semi-fragile watermark into different semantic regions through a watermark encoder, and finally generates a watermarked image through a face reconstruction network; in the watermark extraction and detection stage, the robust watermark and the semi-fragile watermark can be decoded simultaneously to achieve highly generalized blind detection. Thanks to the above improvements, the present invention significantly improves the traceability and detection performance of deepfakes, and the experimental results on multiple data sets all show that it has reached the leading level.

[0101] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0102] Embodiment 2

[0103] The present invention also provides a face deepfake forensics system, which mainly includes: a face deepfake forensics model, which can implement the method provided by the foregoing embodiment based on the foregoing face deepfake forensics model. The system mainly includes:

[0104] A watermark adding module, which is used to receive a face image to be protected and a binary sequence, perform multi-level semantic extraction and regional semantic segmentation on the face image respectively, and synthesize them into multi-level regional semantic features; for each level of regional semantic features, based on the binary sequence, adaptively generate one of two types of set watermarks and embed it into the corresponding level of regional semantic features, and reconstruct a watermarked image by synthesizing all the regional semantic features embedded with watermarks;

[0105] The watermark extraction and verification module is used to extract two types of watermarks from the input watermark image respectively, trace the source through the proposed watermark, and determine whether the input watermark image has been deepfaked by cross-comparison.

[0106] Considering that the relevant technical details involved in this model have been introduced in detail in the foregoing embodiments, they will not be elaborated here.

[0107] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0108] Embodiment III

[0109] The present invention also provides a processing device, as Figure 4 shown, which mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0110] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.

[0111] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:

[0112] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;

[0113] The output device can be a display terminal;

[0114] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.

[0115] Embodiment IV

[0116] The present invention also provides a readable storage medium storing a computer program, which implements the method provided in the foregoing embodiments when the computer program is executed by a processor.

[0117] In the embodiments of the present invention, the readable storage medium, as a computer-readable storage medium, may be disposed in the aforementioned processing device. For example, it may be a memory in the processing device. In addition, the readable storage medium may also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc.

[0118] As described above, the foregoing are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.

Claims

1. A method for collecting evidence of deep fake face, characterized in that: include: Watermarking stage: receiving the face image to be protected and the binary sequence, performing multi-level semantic extraction and regional semantic segmentation on the face image, and synthesizing them into multi-level regional semantic features; for each level of regional semantic features, adaptively generating one of the two types of watermarks set based on the binary sequence, and embedding it into the regional semantic features of the corresponding level, and reconstructing the watermark image by synthesizing all the regional semantic features embedded in the watermark; Watermark extraction and verification stage: extract two types of watermarks from the input watermark image, trace the source through the proposed watermarks, and determine whether the input watermark image has been deeply forged through cross-comparison; The multi-level semantic extraction and regional semantic segmentation of the face image are performed respectively, and the multi-level regional semantic features are synthesized, including: Through a multi-level semantic encoding network Multi-level semantic extraction is performed on the face image to be protected, which is expressed as: ; in, is the face image to be protected, For the The semantic features of the is the series; Through facial region semantic segmentation network The face image to be protected is subjected to regional semantic segmentation, which is expressed as: ; in, For the semantic region masks, N is the number of semantic region masks obtained by regional semantic segmentation; For each level, the following formula is used to synthesize regional semantic features: ; in, For the The regional semantic features of For Hadamard, Indicates that Downsample to The corresponding size.

2. A method for collecting evidence of deep fake face according to claim 1, characterized in that: The two types of watermarks set include: robust watermark and semi-fragile watermark; based on the binary sequence, the corresponding type of watermark is generated through a diffusion process; the diffusion process includes: a fully connected neural network processing process and an upsampling process.

3. A method for collecting evidence of deep fake face according to claim 1 or 2, characterized in that: The adaptive generation of one of the two types of watermarks set based on the binary sequence and embedding into the corresponding level regional semantic features includes: The adaptively generated watermark is concatenated with the corresponding level regional semantic features, and fused through the attention mechanism to obtain the regional semantic features embedded in the watermark; Among them, the watermark embedding process of each level of regional semantic features is expressed as: ; in, For the The regional semantic features of Represents a binary sequence Through the diffusion process Generated watermark, symbol Indicates splicing, represents the attention convolution module with attention mechanism, Indicates The semantic features of the watermark are embedded in the data level.

4. A method for collecting evidence of deep fake face according to claim 1, characterized in that: The extracting of two types of watermarks from the input watermark image, tracing the source through the proposed watermarks, and judging whether the input watermark image has been deeply forged through cross comparison include: The two types of watermarks include: robust watermarks and semi-fragile watermarks; Traceability through robust watermarking; Through cross-comparison, the bit error rate difference between the robust watermark and the semi-fragile watermark is obtained, thereby judging whether the input watermark image has been deeply forged: when the bit error rate difference is less than the set judgment threshold, the input watermark image is judged to be a real image, otherwise, the input watermark image is judged to be deeply forged.

5. The method for collecting evidence of deep fake face according to claim 1, characterized in that: Also includes: The watermark adding stage is implemented by a watermark adding module, and the watermark extraction and verification stage is implemented by a watermark extraction and verification module, which together constitute a face deep forgery forensics model; The face deep fake forensics model is pre-trained, and a progressive noise layer and a perturbation adversarial generative network are introduced during the training process; and an adversarial discriminator is introduced to distinguish the input face image to be protected and the watermark image respectively; The training process is divided into two parts, the former is called the early training period and the latter is called the late training period. In the early stage of training, a plurality of predefined image transformations are randomly selected as noises through the progressive noise layer to apply image disturbance to the watermark image; In the later stage of training, the perturbation adversarial generative network generates an image transformation operation, applies image perturbation to the watermark image, and is trained alternately with the watermark extraction and verification module in an adversarial training manner.

6. A method for collecting evidence of deep fake face according to claim 5, characterized in that: The training loss is expressed as: ; in, is the training loss, , , Correspondingly, they represent the watermark decoding loss term, the watermark image perceptual loss term, and the adversarial loss term. , and is the weight of the corresponding loss item; the watermark decoding loss item is obtained by calculating the difference between the two types of watermarks extracted and the binary sequence respectively; the perceptual loss item of the watermark image is obtained by calculating the difference between the face image to be protected and the watermark image; the adversarial loss item is calculated by the discrimination result output by the adversarial discriminator.

7. A facial deep fake evidence collection system, characterized in that: include: Face deep fake forensics model; The method according to any one of claims 1 to 6 is implemented based on the face deep fake forensics model, wherein the face deep fake forensics model comprises: The watermark adding module is used to receive the face image to be protected and the binary sequence, perform multi-level semantic extraction and regional semantic segmentation on the face image, and synthesize them into multi-level regional semantic features; for each level of regional semantic features, adaptively generate one of the two types of watermarks set based on the binary sequence, and embed it into the regional semantic features of the corresponding level, and reconstruct the watermark image by synthesizing all the regional semantic features embedded in the watermark; The watermark extraction and verification module is used to extract two types of watermarks from the input watermark image, trace the source through the proposed watermarks, and determine whether the input watermark image has been deeply forged through cross-comparison.

8. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Digital watermark processing method and device, electronic equipment and storage medium

    CN111754379A

  • Face deep counterfeiting evidence obtaining method, system and device and storage medium

    CN118279995A