Self-supervised deep pseudo image detection method and system

Through the self-supervised framework and the deep manifold registration polymeric expression, the self-supervised deep pseudo-image detection method solves the problem of dependence on labeled samples in the prior art, and improves the accuracy and generalization ability of detection.

CN120088626APending Publication Date: 2025-06-03XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411957505.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing deep pseudo-image detection methods rely on supervised learning and require a large number of samples to be marked, which leads to high cost and limited generalization capabilities, making it difficult to effectively detect new deep pseudo-images.

Method used

A self-supervised deep pseudo-image detection method is proposed. By constructing a self-supervised framework, using deep pseudo-high-frequency information self-generating technology and deep manifold registration polymerization expressions, reducing dependence on labeled samples and improving detection generalization capabilities.

Benefits of technology

This method can effectively reduce the cost of sample acquisition, improve the accuracy and generalization ability of deep pseudo-image detection, clarify decision-making boundaries, and is suitable for detecting various deep pseudo-images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088626A_ABST
    Figure CN120088626A_ABST
Patent Text Reader

Abstract

The invention provides a self-supervised deep pseudo image detection method and system, and the method comprises the steps: constructing a data set, building a self-supervised frame named as a model F, and mainly comprising a deep pseudo high-frequency information generation module, a comparative learning and sample processing frame and a deep manifold registration aggregation unit. And introducing a feature embedding mask strategy, and performing sample distribution processing and loss calculation by using a distribution shaper to finish fine adjustment and detection of the model. The invention provides a self-supervision framework, and aims to simulate extra high-frequency signals formed by most of the generative models in a decoding stage by utilizing a low-cost simulation auto-encoder model by analyzing the generality of the generative models, so that a certain online deep pseudo high-frequency information self-generation mode is realized. In addition, according to the scheme of the invention, a deep manifold registration aggregation generator is introduced to strengthen the effect of a deep pseudo detection algorithm, so as to extract features with higher discrimination capability from a deep pseudo image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a self-supervised deepfake image detection method and system. Background Art

[0002] With the rapid development of digital technology, deepfake technology makes it increasingly difficult to distinguish the generated deepfake images.

[0003] Hebei University of Technology proposed a deepfake image detection method that combines depth and width learning. By using the centralized frequency domain amplitude spectrum combined with the attention mechanism, and using techniques such as channel splicing, feature fusion, and width learning at the macro level, it focuses on the tampering traces from the global information center. Although it makes the neural network focus on the features related to the tampering traces to a certain extent, its framework is essentially a supervised learning strategy, which has high requirements for sample collection and manufacturing, and cannot reduce the relevant implementation costs;

[0004] Li Da, Yang Ke, etc. proposed a face forgery detection method, device, equipment and storage medium, which mainly constructs a series of basic neural operation units and then stacks them into a new neural network backbone, and combines prior attention to strengthen the feature selection mechanism in the information processing process of the network. Although it alleviates the phenomena of overly broad feature utilization and insufficient causal support of evidence in deepfake detection to a certain extent, the clarity of the prior attention proposed by it is insufficient, and the influence on the inference result remains to be verified, and it does not solve the problem of too high annotation cost of such algorithms.

[0005] Most of the existing deepfake image detection methods rely on supervised learning and require a large number of labeled samples. However, obtaining these labeled samples is not only costly, but also often difficult to cover all types of deepfake images, resulting in limited generalization ability of the detection model.

[0006] In addition, the features of deepfake images are complex and elusive. Traditional detection methods have low detection accuracy and reliability when facing new types of deepfake images, and it is difficult to meet the needs of practical applications. Therefore, it is of great practical significance to develop an efficient, accurate and well-generalized deepfake image detection system. Summary of the Invention

[0007] The present invention provides a self-supervised deepfake image detection method and system, which improves the problems of overly strong inductive bias and high training cost of the supervised algorithm based on the fine-grained classification model construction strategy and feature selection mechanism, and proposes a new self-supervised framework composed of the deepfake high-frequency information self-generation technology and the deep manifold registration aggregation expression.

[0008] In the first aspect, the present invention proposes a self-supervised deepfake image detection method, which includes the following steps:

[0009] S1. Build a dataset: Create a portrait image dataset. The image dataset contains a number of portrait images, including different age states of both male and female genders. The image resolution meets the preset minimum standard, and after processing, a real dataset R is obtained. d Create a deepfake labeled dataset, which contains a preset number of deepfake images.

[0010] S2. Build a model architecture: Build a self-supervised framework named Model F. Model F mainly consists of three parts: a deepfake high-frequency information generation module, a contrastive learning and sample processing framework, and a deep manifold registration aggregation unit. The contrastive learning and sample processing framework is used to distinguish an image from its deepfake high-frequency information pattern version, and the deep manifold registration aggregation unit is used to macroscopically implement a deep adversarial mechanism-based deep manifold registration aggregator for spatial metric learning.

[0011] S3. Feature optimization operation: By introducing a feature embedding mask strategy and setting an image encoder and a corresponding decoder, the introduced mask strategy is applied to the encoded output features, which are further used as the input of the encoder. Finally, a reconstructed image is obtained. Samples with different deepfake high-frequency information patterns in the output are grouped into an autoencoding convolutional neural network group output dataset I. r ;

[0012] S4. Sample distribution processing and loss calculation: Use a distribution shaper to map the samples in the real dataset R d and the autoencoding convolutional neural network group output dataset I r to latent space feature encodings through latent space metric optimization of the manifold registration criterion. Then, randomly select anchor samples, real samples, and deepfake samples from the output sample set, execute a triplet pairing sampling strategy, and calculate using a metric learning loss function.

[0013] S5. Model fine-tuning and detection: Use the deepfake labeled dataset to perform supervised fine-tuning on Model F, calculate using a cross-entropy loss function, and obtain a final deepfake detection model after training. The deepfake detection model can output a digital label representing whether the input image is forged after receiving the input image, for further related inference analysis tasks.

[0014] Preferably, in the model architecture building step of S2, the deepfake high-frequency information generation module, the contrastive learning and sample processing framework, and the deep manifold registration aggregation unit specifically include:

[0015] The deepfake high-frequency information generation module: Based on the autoencoder convolutional neural network, construct no less than twenty kinds of image encoding and decoding networks. The number of encoder layers is randomly determined within the range of 3 to 6 layers, and the number of decoder layers is randomly selected within the range of 3 to 9 layers. The upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of the UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused to generate different deepfake high-frequency information patterns of the same picture, constituting a set of deepfake high-frequency information generation methods;

[0016] The contrastive learning and sample processing framework: Construct a self-supervised architecture based on contrastive learning and a queue-based sample processing pool to effectively distinguish pictures from their deepfake high-frequency information pattern versions and enhance the model's ability to identify the authenticity differences in images;

[0017] The deep manifold registration and aggregation unit: Set up a deep manifold registration and aggregation expression based on spatial metric learning to implement a deep adversarial mechanism at the macroscopic level and assist the model in accurately judging the authenticity of samples.

[0018] Preferably, in the feature optimization operation step of S3, it specifically includes:

[0019] Introduce a feature embedding mask strategy, set the image encoder as f e , and the corresponding decoder as f d . For the input image I n , after being encoded by the encoder f e , the feature z e = f e (I n ) is obtained;

[0020] By randomly generating a mask M e , make the mask strategy act on the encoded feature z e . Take z e ⊙ M e as the input of the decoder f d , and finally obtain the reconstructed image I r = f d (z e ⊙ M e ), so as to prompt the decoder to achieve fine-grained restoration and simulate the information completion pressure in the deepfake model generation stage.

[0021] Preferably, in the sample distribution processing and loss calculation step of S4, randomly select anchor samples, real samples, and deepfake samples from the output sample set, execute a triplet pairing sampling strategy, and calculate using a metric learning loss function, specifically including:

[0022] Anchor samples, real samples and deep fake samples are randomly selected from the sample set, and the loss function is used to shorten the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample, which includes a guidance parameter γ that controls the amplitude of the triplet loss regularization direction and a regularization function Cos that performs constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

[0023] Preferably, in the model fine-tuning and detection step of S5, the cross entropy loss function is:

[0024]

[0025] Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

[0026] Preferably, in the step of constructing a data set in S1, the following steps are specifically included:

[0027] The portrait image dataset N is ≥ 200,000, covering both male and female genders, and including the age status of infants, children, teenagers, young people, middle-aged and elderly people, and the image resolution is above 224*224;

[0028] The deep fake labeled dataset M created is ≥ 3000, and the forgery types cover deepfake, faceswap, stylegan model and conditional generation model.

[0029] In a second aspect, an embodiment of the present invention provides a self-supervised deep fake image detection system, comprising:

[0030] The data storage module 41 is configured to create a portrait image data set, wherein the image data set includes a number of portrait images, including male and female images of different ages, and the image resolution meets the preset minimum standard. After processing, the real data set R is obtained. d ; Creating a deep fake labeled dataset, wherein the deep fake labeled dataset contains a preset number of deep fake images;

[0031] A model building module is configured to build a self-supervised framework named Model F, which mainly consists of three parts: a deep fake high-frequency information generation submodule, a contrastive learning and sample processing submodule, and a deep manifold registration aggregation submodule;

[0032] A feature processing module, configured to set an image encoder and a corresponding decoder by introducing a feature embedding masking strategy, so that the introduced masking strategy acts on the encoded output features, and further serves as the input of the encoder, and finally obtain a reconstructed image, and form an output dataset I of an autoencoder convolutional neural network group from samples with different deepfake high-frequency information patterns r ;

[0033] A sample processing and loss calculation module, configured to use a distribution shaper to optimize the manifold registration criterion through latent space metric for the real dataset R d and the output dataset I of the autoencoder convolutional neural network group r Map the samples in to latent space feature encodings, then randomly select anchor samples, real samples and deepfake samples from the output sample set, execute a triplet pairing sampling strategy, and calculate using a metric learning loss function;

[0034] A model training and detection module, uses the deepfake labeled dataset to perform supervised fine-tuning on the model F, calculates using a cross-entropy loss function, and obtains a final deepfake detection model after training. The deepfake detection model can output a digital label representing whether the input image is forged after receiving the input image, for further related inference analysis tasks.

[0035] Preferably, in the model construction module, it further includes:

[0036] A deepfake high-frequency information generation sub-module: configured to build at least twenty picture encoding and decoding units based on an autoencoder convolutional neural network, randomly determine the number of encoder layers within the range of 3 to 6 layers, and randomly select the number of decoder layers within the range of 3 to 9 layers. The upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of the UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused, and then different deepfake high-frequency information patterns of the same picture can be generated to form a set of deepfake high-frequency information generation methods;

[0037] A contrast learning and sample processing sub-module: configured to build a self-supervised architecture based on contrast learning and a queued sample processing pool to effectively distinguish pictures from their deepfake high-frequency information pattern versions and enhance the model's ability to identify the authenticity differences of images;

[0038] A deep manifold registration aggregation sub-module, configured to set up a deep manifold registration aggregation expression based on spatial metric learning to achieve a deep adversarial mechanism at the macroscopic level and assist the model in accurately judging the authenticity of samples.

[0039] Preferably, in the feature processing module, it specifically includes:

[0040] Introduce the feature embedding mask strategy and set the image encoder to f e , the corresponding decoder is f d , for the input image I n , through the encoder f e Encoding gets feature z e =f e (I n );

[0041] By randomly generating a mask M e , so that the mask strategy acts on the encoded feature z e , z e ⊙M e As decoder f d The reconstructed image I is finally obtained. r =f d (z e ⊙M e ), thereby forcing the decoder to achieve fine-grained restoration and simulating the information completion pressure in the generation stage of deep fake models.

[0042] Preferably, in the sample processing and loss calculation module, the output sample set randomly selects anchor samples, real samples and deep fake samples, executes the triple pairing sampling strategy, and uses the metric learning loss function for calculation, specifically including:

[0043] Anchor samples, real samples and deep fake samples are randomly selected from the sample set, and the loss function is used to shorten the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample, which includes a guidance parameter γ that controls the amplitude of the triplet loss regularization direction and a regularization function Cos that performs constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

[0044] Preferably, in the model training and detection module, the cross entropy loss function is configured as:

[0045]

[0046] Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

[0047] Preferably, during the construction of the model building module, the output of the deepfake high-frequency information generation sub-module is used as one of the inputs of the contrast learning and sample processing sub-module, and the deep manifold registration aggregation sub-module receives the outputs of the contrast learning and sample processing sub-module and the distribution shaping unit to achieve the orderly transfer and collaborative work of data between various modules in the system, and improve the detection performance of the system for deepfake images.

[0048] In a third aspect, an embodiment of the present invention provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.

[0049] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0050] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0051] The present invention innovatively proposes a set of deepfake information self-generation strategies, which greatly alleviates the sample requirements in the research and development process of related models. By deeply understanding the principle mechanism of the deepfake algorithm for generating image content, analyzing various characteristics in the neural network information conduction process, and finding out the common information hidden in the content generated by the deepfake model, the inventor is guided to propose a resolution upsampling and downsampling strategy including multiple autoencoder-decoder structures. An innovative deep manifold registration aggregation generator is proposed to output a simulation result to judge the spatial distance between genuine and fake samples, clarify the decision boundary, and improve the generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the intended advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily drawn to scale. The same reference numerals refer to corresponding like parts.

[0053] Figure 1 is an exemplary device architecture diagram to which an embodiment of the present invention can be applied;

[0054] Figure 2 is a schematic flowchart of the self-supervised deepfake image detection method according to the embodiment of the present invention;

[0055] Figure 3Technical architecture diagram of the self-supervised deepfake image detection algorithm according to a specific embodiment of the present invention;

[0056] Figure 4 Schematic diagram of the architecture of the self-supervised deepfake image detection system according to an embodiment of the present invention;

[0057] Figure 5 Schematic diagram of the architecture of the model construction module according to a specific embodiment of the present invention;

[0058] Figure 6 Schematic diagram of the structure of the computer device of the electronic device suitable for implementing the embodiment of the present invention. Detailed description of specific embodiments

[0059] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that, for the sake of description, only the parts related to the relevant invention are shown in the drawings.

[0060] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.

[0061] Figure 1 Exemplary system architecture 100 to which the self-supervised deepfake image detection method or the self-supervised deepfake image detection system according to the embodiment of the present invention can be applied is shown.

[0062] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0063] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0064] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to smartphones, tablet computers, laptop portable computers, and desktop computers, etc. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0065] The server 105 can be a server that provides various services, such as a background information processing server that processes the verification request information sent by the terminal devices 101, 102, and 103. The background information processing server can analyze and process the received verification request information and obtain a processing result.

[0066] It should be noted that the self-supervised deepfake image detection method provided by the embodiments of the present invention is generally executed by the server 105. Correspondingly, the self-supervised deepfake image detection system is generally set in the server 105. In addition, the method for sending information provided by the embodiments of the present invention is generally executed by the terminal devices 101, 102, and 103. Correspondingly, the system for sending information is generally set in the terminal devices 101, 102, and 103.

[0067] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (such as for providing distributed services), or it can be implemented as a single software or multiple software modules. No specific limitation is made here.

[0068] It should be understood that Figure 1 the numbers of the terminal devices, network, and server in

[0069] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, network, and server. In the case where the data to be processed does not need to be obtained remotely, the above device architecture may not include a network, but only a server or a terminal device.

[0070] To solve the above problems, an embodiment of the present invention proposes a self-supervised framework. By analyzing the commonalities of generative models, it aims to use a low-cost simulation autoencoder model to simulate the additional high-frequency signals formed by most generative models during the decoding stage, thereby realizing a certain online deepfake high-frequency information self-generation mode. In addition, this solution introduces a deep manifold registration aggregation generator to enhance the effect of the deepfake detection algorithm, in order to extract more discriminative features from deepfake images.

[0071] In a first aspect, an embodiment of the present invention discloses a self-supervised deepfake image detection method, as Figure 2 and 3 shown, the method includes the following steps:

[0072] S1. Construct a dataset: Create a portrait picture dataset. The picture dataset contains a number of portrait pictures, including different age states of both male and female genders. The picture resolution meets the preset minimum standard, and after processing, a real dataset R is obtained d ; Create a deepfake labeled dataset, which contains a preset number of deepfake images;

[0073] Specifically, in this embodiment, constructing the dataset specifically includes:

[0074] S11. The created portrait picture dataset N≥200,000, covering both male and female genders, and including age states of infants, children, teenagers, young people, middle-aged people, and the elderly. The picture resolution reaches 224*224 or above;

[0075] S12. The created deepfake labeled dataset M≥3,000, where the forgery types cover deepfake, faceswap, stylegan models, and conditional generative models.

[0076] S2. Model architecture construction: Build a self-supervised framework named model F. Model F mainly consists of three parts: a deepfake high-frequency information generation module, a contrast learning and sample processing framework, and a deep manifold registration aggregation unit. The contrast learning and sample processing framework is used to distinguish between pictures and their deepfake high-frequency information pattern versions, and the deep manifold registration aggregation unit is used to realize a deep manifold registration aggregation expression based on spatial metric learning of a deep adversarial mechanism from a macroscopic perspective;

[0077] Specifically, in the model architecture construction of this step, the deepfake high-frequency information generation module, the contrast learning and sample processing framework, and the deep manifold registration aggregation unit specifically include:

[0078] Deepfake high-frequency information generation module: Based on the autoencoder convolutional neural network, construct no less than twenty kinds of image encoding and decoding networks. The number of encoder layers is randomly determined within the range of 3 to 6 layers, and the number of decoder layers is randomly selected within the range of 3 to 9 layers. The upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused, and then different deepfake high-frequency information patterns of the same image can be generated, constituting a set of deepfake high-frequency information generation methods;

[0079] Contrast learning and sample processing framework: Construct a self-supervised architecture based on contrast learning and a queue-style sample processing pool to effectively distinguish images from their deepfake high-frequency information pattern versions and enhance the model's ability to recognize the authenticity differences in images;

[0080] Deep manifold registration aggregation unit: Set up a deep manifold registration aggregation expression based on spatial metric learning to achieve a deep adversarial mechanism at the macroscopic level and assist the model in accurately judging the authenticity of samples.

[0081] S3. Feature optimization operation: By introducing a feature embedding mask strategy, set up an image encoder and the corresponding decoder, so that the introduced mask strategy acts on the encoded output features, and further serves as the input of the encoder, and finally obtain a reconstructed image. The samples with different deepfake high-frequency information patterns output are composed into the output dataset I of the autoencoder convolutional neural network group r ;

[0082] Specifically, in this feature optimization operation step, it specifically includes:

[0083] S31. Introduce a feature embedding mask strategy, and set the image encoder as f e , and the corresponding decoder as f d . For the input image I n , after being encoded by the encoder f e , the feature z e = f e (I n ) is obtained;

[0084] S32. By randomly generating a mask M e , make the mask strategy act on the encoded feature z e , and use z e ⊙ M e as the input of the decoder f d , and finally obtain the reconstructed image I r = f d (z e ⊙ M e ), so as to prompt the decoder to achieve fine-grained restoration and simulate the information completion pressure in the deepfake model generation stage.

[0085] S4. Sample distribution processing and loss calculation: Using the distribution shaper, the real dataset R is optimized through the latent space metric manifold registration criterion. d Output dataset of autoencoder convolutional neural network group I r The samples in the output sample set are mapped to the latent space feature encoding, and then the anchor samples, real samples and deep fake samples are randomly selected from the output sample set, and the triple pairing sampling strategy is implemented, and the metric learning loss function is used for calculation;

[0086] Specifically, in this sample distribution processing and loss calculation step, anchor samples, real samples and deep fake samples are randomly selected from the output sample set, the triple pairing sampling strategy is executed, and the metric learning loss function is used for calculation, which specifically includes:

[0087] Anchor samples, real samples and deep fake samples are randomly selected from the sample set. The loss function is used to narrow the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample. It contains the guidance parameter γ that controls the amplitude of the triplet loss regularization direction and the regularization function Cos that enforces constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

[0088] S5. Model fine-tuning and detection: The model F is fine-tuned in a supervised manner using a deep fake labeled dataset. The cross entropy loss function is used for calculation. After training, the final deep fake detection model is obtained. After receiving the input image, the deep fake detection model can output a digital label representing whether the image is forged, which is used for further related reasoning and analysis tasks.

[0089] Specifically, in the model fine-tuning and detection steps, the cross entropy loss function is:

[0090]

[0091] Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

[0092] The embodiment of the present invention improves the problems of strong inductive bias and high training cost of supervised algorithms based on fine-grained classification model construction strategy and feature optimization mechanism, and proposes a new self-supervisory framework composed of deep pseudo high-frequency information self-generation technology and deep popular registration aggregation expressor. The deep pseudo high-frequency information generator takes real images as input and reconstructs high-frequency information of various subspaces of the image using the resolution up-down sampling strategy of multiple self-encoder-decoder structures.

[0093] The embodiments of the present invention aim to solve the problems of high sample acquisition cost, unbalanced sample distribution, and insufficient generalization ability for unknown domain deepfake samples in supervised deepfake image detection algorithms. Finally, a deep manifold registration aggregation generator is trained to output a simulation result, that is, to judge the distance between real and fake samples in space in the way of manifold registration and metric learning. The whole process can ignore the forged image dataset, thus forming a self-supervised learning process without manual annotation. In short, this patent alleviates the sample demand in the process of related model research and development, clarifies the decision boundary, and improves the overall generalization ability of the algorithm.

[0094] Figure 3 The technical architecture diagram of the self-supervised deepfake image detection algorithm according to a specific embodiment of the present invention is shown.

[0095] In a specific embodiment, as Figure 3 shown, the specific steps of the self-supervised deepfake image detection algorithm of the present invention are as follows:

[0096] Step 1: Create a portrait picture dataset with N≥200,000 images, covering both genders and including age states such as infants, children, teenagers, young people, middle-aged people, and the elderly. The picture resolution reaches 224*224 or above.

[0097] Step 2: Create a deepfake labeled dataset with M≥3,000 images, where the forgery types cover models such as deepfake, faceswap, stylegan, and conditional generation models.

[0098] Step 3: Build a self-supervised framework named Model F, which is specifically composed of three parts:

[0099] 1. A set of deepfake high-frequency information generation methods in Step 4, used to output different deepfake high-frequency information patterns of the same picture.

[0100] 2. A self-supervised framework based on contrastive learning and a queue-style sample processing pool, used to distinguish between a picture and its deepfake high-frequency information pattern version.

[0101] 3. A deep manifold registration aggregation expression based on spatial metric learning, which can achieve a deep adversarial mechanism from a macroscopic perspective.

[0102] Step 4: Propose a large-scale deepfake high-frequency information generation method, that is, based on the basic structure of an autoencoder convolutional neural network, establish no less than 20 picture encoding and decoding networks to form a set of deepfake high-frequency information generation methods.

[0103] The technical details of the above module are as follows: the number of encoder layers is randomly selected between 3 and 6 layers, and the number of decoder layers is randomly selected between 3 and 9 layers. The upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. Among them, UNET can combine the deep high-level features from the decoder with the shallow low-level features from the encoder, and the decoded information benefits from the effective fusion support of the shallow and deep features.

[0104] Step 5: To further force the decoder to achieve fine-grained restoration and simulate the information completion pressure faced by the deepfake model during the generation stage, this scheme introduces a feature embedding mask strategy to strengthen the information flow path of the above encoder-decoder. Its formal description is as follows:

[0105] Let the image encoder be f e , and the corresponding decoder be f d . Its output result is:

[0106] z e = f e (I n )

[0107] z e is the encoded feature. The mask strategy acts on z e and serves as the input to the decoder. Its formula is:

[0108] I r = f d (z e ⊙M e )

[0109] where M e is a randomly generated mask. I r is the reconstructed image.

[0110] Step 6: Let R d and I r represent samples in the real dataset and the dataset output by the autoencoder convolutional neural network group. The distribution shaper uses a latent space metric to optimize the manifold registration criterion to map the samples to the latent space feature encoding, and then uses a metric learning loss function to randomly select anchor samples, real samples, and deepfake samples and follow a certain triplet pairing sampling strategy.

[0111] When the loss function takes effect, it realizes pulling the distance between the positive example sample and the anchor sample and tries to push the negative sample to a farther distance. γ is the magnitude that controls the regularization direction of the triplet loss and is a guiding parameter. x a , x p are the positive example sample and the anchor sample respectively, and x n1 and x n2For negative example samples, Cos is a regularization function that can enforce certain constraints in distance control to keep it within a reasonable spatial range.

[0112] Step 7: Use the labeled dataset M in Step 1 to perform supervised fine-tuning on the model F, with the cross-entropy loss as the loss function:

[0113]

[0114] where y i represents the label of sample i. We specify the deepfake example class as number 1 and the normal example class as number 0. p i represents the probability that sample i is predicted as the deepfake example class.

[0115] Step 8: Obtain the final version of the deepest fake detection model, which takes an input image and outputs a digital label representing whether the image is forged, for use in related inference and analysis tasks.

[0116] The present invention innovatively proposes a set of deepfake information self-generation strategies, which greatly alleviates the sample requirements in the process of developing related models. By deeply understanding the principle mechanism of the deepfake algorithm for generating image content, analyzing various characteristics in the information conduction process of its neural network, and finding out the common information hidden in the content generated by the deepfake model, the inventor is guided to propose a resolution upsampling and downsampling strategy including multiple autoencoder-decoder structures. An innovative depth manifold registration aggregation generator is proposed to output a simulation result to judge the spatial distance between real and fake samples, clarify the decision boundary, and improve the generalization ability.

[0117] Further referring to Figure 4 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a self-supervised deepfake image detection system. This system embodiment corresponds to Figure 2 the method embodiment shown, and this system can be specifically applied to various electronic devices.

[0118] In a second aspect, an embodiment of the present invention discloses a self-supervised deepfake image detection system, as shown in Figure 4 . This system includes: a data storage module 41, a model construction module 42, a feature processing module 43, a sample processing and loss calculation module 44, and a model training and detection module 45.

[0119] In this embodiment, the data storage module 41 is configured to create a portrait picture dataset. The picture dataset contains a number of portrait pictures, including different age states of both male and female genders. The picture resolution meets a preset minimum standard, and after processing, a real dataset R is obtained d ; create a deepfake labeled dataset, which contains a preset number of deepfake images;

[0120] The model construction module 42 is configured to build a self-supervised framework named Model F, which mainly consists of three parts: the deepfake high-frequency information generation sub-module 421, the contrastive learning and sample processing sub-module 422, and the deep manifold registration and aggregation sub-module 423; the feature processing module 43 is configured to introduce a feature embedding mask strategy, set up an image encoder and a corresponding decoder, make the introduced mask strategy act on the encoded output features, and further use them as the input of the encoder, and finally obtain a reconstructed image. Samples with different deepfake high-frequency information patterns are composed into an autoencoder convolutional neural network group to output the dataset I. r ;

[0121] The sample processing and loss calculation module 44 is configured to use a distribution shaper to optimize the manifold registration criterion through latent space metric for the real dataset R d and the autoencoder convolutional neural network group output dataset I r in the samples are mapped to latent space feature encodings, and then anchor samples, real samples, and deepfake samples are randomly selected from the output sample set to execute a triplet pairing sampling strategy, and a metric learning loss function is used for calculation; the model training and detection module 45 uses a deepfake labeled dataset to perform supervised fine-tuning on Model F, calculates using a cross-entropy loss function, and obtains a final deepfake detection model after training. The deepfake detection model can output a digital label representing whether the input image is forged after receiving the input image, which is used for further related inference analysis tasks.

[0122] In a specific embodiment, as Figure 5 shown, in the model construction module 42, it further includes: the deepfake high-frequency information generation sub-module 421: is configured to build no less than twenty kinds of picture encoding and decoding units based on an autoencoder convolutional neural network, the number of encoder layers is randomly determined within the range of 3 to 6 layers, the number of decoder layers is randomly selected within the range of 3 to 9 layers, and the upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused, and then different deepfake high-frequency information patterns of the same picture are generated, constituting a set of deepfake high-frequency information generation methods;

[0123] The contrastive learning and sample processing sub-module 422: is configured to build a self-supervised architecture based on contrastive learning and a queue-based sample processing pool, which is used to effectively distinguish pictures from their own deepfake high-frequency information pattern versions and enhance the model's ability to identify the authenticity differences of images;

[0124] The deep manifold registration and aggregation sub-module 423 is configured to set up a deep manifold registration and aggregation expression based on spatial metric learning, and implement a deep adversarial mechanism from a macroscopic level to assist the model in accurately judging the authenticity of samples.

[0125] In the feature processing module 43, the following steps are specifically performed: introducing a feature embedding mask strategy, setting the image encoder to f e , the corresponding decoder is f d , for the input image I n , through the encoder f e Encoding gets feature z e =f e (I n ); by randomly generating a mask M e , so that the mask strategy acts on the encoded feature z e , z e ⊙M e As decoder f d The reconstructed image I is finally obtained. r =f d (z e ⊙M e ), thereby forcing the decoder to achieve fine-grained restoration and simulating the information completion pressure in the generation stage of deep fake models.

[0126] In the sample processing and loss calculation module 44, anchor samples, real samples and deep fake samples are randomly selected from the output sample set, and the triplet pairing sampling strategy is executed. The metric learning loss function is used for calculation, which specifically includes: randomly selecting anchor samples, real samples and deep fake samples from the sample set, and the loss function is used to shorten the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample, which includes the guidance parameter γ for controlling the amplitude of the triplet loss regularization direction and the regularization function Cos for performing constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

[0127] In the model training and detection module 45, the cross entropy loss function is configured as:

[0128]

[0129] Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

[0130] During the construction process of the model building module 42, the output of the deep fake high-frequency information generation submodule 421 is used as one of the inputs of the contrast learning and sample processing submodule 422, and the deep manifold registration aggregation submodule 423 receives the output of the contrast learning and sample processing submodule 422 and the distribution modeling unit 441 to achieve orderly flow and collaborative work of data between modules within the system, thereby improving the system's detection performance for deep fake images.

[0131] The functions of the above modules correspond to the methods and will not be elaborated here.

[0132] Next, refer to Figure 6 , which shows a schematic structural diagram of a computer device 600 suitable for use in implementing the electronic device (such as Figure 1 the server or terminal device shown) of the embodiments of the present invention. Figure 6 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0133] As Figure 6 shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 603 or the programs loaded from the storage section 609 into the random access memory (RAM) 604. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other via a bus 605. The input / output (I / O) interface 606 is also connected to the bus 605.

[0134] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including, for example, a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 can also be connected to the I / O interface 606 as needed. A removable medium 612, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 611 as needed so that the computer program read from it can be installed into the storage section 609 as needed.

[0135] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication section 610 and / or installed from the removable medium 612. When the computer program is executed by the central processing unit (CPU) 601 and the graphics processing unit (GPU) 602, the above functions defined in the method of the present invention are executed.

[0136] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium, a computer-readable medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus, or component, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution device, apparatus, or component. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0137] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based device that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0139] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor.

[0140] As another aspect, the present invention also provides a computer-readable medium, which can be included in the electronic device described in the above embodiments; or can exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to: execute the methods and steps described in the first aspect of the embodiments of the present invention.

[0141] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present invention.

Claims

1. A self-supervised deep fake image detection method, characterized in that: The method comprises the following steps: S1. Constructing a dataset: Creating a portrait image dataset, which contains a number of portrait images, including male and female images of different ages, with image resolutions that meet the preset minimum standard. After processing, the real dataset R is obtained. d ; Creating a deep fake labeled dataset, wherein the deep fake labeled dataset contains a preset number of deep fake images; S2. Model architecture construction: Build a self-supervised framework named Model F, which is mainly composed of three parts: a deep fake high-frequency information generation module, a contrastive learning and sample processing framework, and a deep manifold registration aggregation unit. The contrastive learning and sample processing framework is used to distinguish between the image and its own deep fake high-frequency information mode version, and the deep manifold registration aggregation unit is used to realize the deep manifold registration aggregation expressor based on spatial metric learning of the deep adversarial mechanism from a macro perspective; S3, feature optimization operation: by introducing the feature embedding mask strategy, setting the image encoder and the corresponding decoder, the introduced mask strategy acts on the encoded output features, further as the input of the encoder, and finally obtains the reconstructed image, and the output samples with different deep pseudo high-frequency information patterns are composed of the self-encoder convolutional neural network group output data set I r ; S4. Sample distribution processing and loss calculation: Using the distribution shaper, the real data set R is optimized through the latent space metric manifold registration criterion. d Output dataset of autoencoder convolutional neural network group I r The samples in the output sample set are mapped to the latent space feature encoding, and then the anchor samples, real samples and deep fake samples are randomly selected from the output sample set, and the triple pairing sampling strategy is implemented, and the metric learning loss function is used for calculation; S5. Model fine-tuning and detection: The model F is fine-tuned in a supervised manner using the deep fake labeled dataset, and the cross entropy loss function is used for calculation. After training, the final deep fake detection model is obtained. After receiving the input image, the deep fake detection model can output a digital label representing whether the image is forged, which is used for further related reasoning and analysis tasks.

2. The self-supervised deep fake image detection method according to claim 1, characterized in that: In the model architecture building step of S2, the deep pseudo high-frequency information generation module, contrastive learning and sample processing framework and deep manifold registration aggregation unit specifically include: The deep fake high-frequency information generation module: constructs no less than twenty kinds of image encoders and decoders based on the autoencoder convolutional neural network, the number of encoder layers is randomly determined in the range of 3 to 6 layers, the number of decoder layers is randomly selected in the range of 3 to 9 layers, and the upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of the UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused to generate different deep fake high-frequency information modes of the same image, forming a set of deep fake high-frequency information generation methods; The contrastive learning and sample processing framework: constructs a self-supervised architecture based on contrastive learning and queue-based sample processing pools to effectively distinguish between images and their own deep fake high-frequency information pattern versions, and enhances the model's ability to identify differences between true and false images; The deep manifold registration aggregation unit: establishes a deep manifold registration aggregation expressor based on spatial metric learning, realizes a deep adversarial mechanism from a macro level, and assists the model to accurately judge the authenticity of samples.

3. The self-supervised deep fake image detection method according to claim 1, characterized in that: The feature optimization operation steps in S3 include: Introduce the feature embedding mask strategy and set the image encoder to f e , the corresponding decoder is f d , for the input image I n , through the encoder f e Encoding gets feature z e =f e (I n ); By randomly generating a mask M e , so that the mask strategy acts on the encoded feature z e , z e ⊙M e As decoder f d The reconstructed image I is finally obtained. r =f d (z e ⊙M e ), thereby forcing the decoder to achieve fine-grained restoration and simulating the information completion pressure in the deep fake model generation stage.

4. The self-supervised deep fake image detection method according to claim 1, characterized in that: In the sample distribution processing and loss calculation step of S4, the output sample set randomly selects anchor samples, real samples and deep fake samples, executes the triple pairing sampling strategy, and uses the metric learning loss function for calculation, specifically including: Anchor samples, real samples and deep fake samples are randomly selected from the sample set, and the loss function is used to shorten the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample, which includes a guidance parameter γ that controls the amplitude of the triplet loss regularization direction and a regularization function Cos that performs constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

5. The self-supervised deep fake image detection method according to claim 1, characterized in that: In the model fine-tuning and detection step of S5, the cross entropy loss function is: Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

6. The self-supervised deep fake image detection method according to claim 1, characterized in that: The step of building a dataset in S1 specifically includes: The portrait image dataset N is ≥ 200,000, covering both male and female genders, and including the age status of infants, children, teenagers, young people, middle-aged and elderly people, and the image resolution is above 224*224; The deep fake labeled dataset M created is ≥ 3000, and the forgery types cover deepfake, faceswap, stylegan model and conditional generation model.

7. A self-supervised deep fake image detection system, characterized in that include: The data storage module is configured to create a portrait image dataset, wherein the image dataset contains a number of portrait images, including male and female images of different ages, and the image resolution meets the preset minimum standard. After processing, the real data set R is obtained. d ; Creating a deep fake labeled dataset, wherein the deep fake labeled dataset contains a preset number of deep fake images; A model building module is configured to build a self-supervised framework named Model F, which mainly consists of three parts: a deep fake high-frequency information generation submodule, a contrastive learning and sample processing submodule, and a deep manifold registration aggregation submodule; The feature processing module is configured to introduce a feature embedding mask strategy, set an image encoder and a corresponding decoder, so that the introduced mask strategy acts on the encoded output features, further serves as the input of the encoder, and finally obtains the reconstructed image, and the output samples with different deep pseudo high-frequency information patterns are composed of the self-encoder convolutional neural network group output data set I r ; The sample processing and loss calculation module is configured to use the distribution shaper to optimize the manifold registration criterion through the latent space metric to transform the real data set R d Output dataset of autoencoder convolutional neural network group I r The samples in the output sample set are mapped to the latent space feature encoding, and then the anchor samples, real samples and deep fake samples are randomly selected from the output sample set, and the triple pairing sampling strategy is implemented, and the metric learning loss function is used for calculation; The model training and detection module uses the deep fake labeled dataset to perform supervised fine-tuning on the model F, and uses the cross entropy loss function for calculation. After training, the final deep fake detection model is obtained. After receiving the input image, the deep fake detection model can output a digital label representing whether the image is forged, which is used for further related reasoning and analysis tasks.

8. The self-supervised deep fake image detection system according to claim 7, characterized in that: The model building module also includes: Deep fake high-frequency information generation submodule: configured to construct no less than twenty image encoders and decoders based on the autoencoder convolutional neural network, the number of encoder layers is randomly determined in the range of 3 to 6 layers, the number of decoder layers is randomly selected in the range of 3 to 9 layers, and the upsampling method is randomly selected from three methods: transposed convolution, differentiable bilinear interpolation, and UNET connection. With the help of the UNET connection, the deep high-level features of the decoder and the shallow low-level features of the encoder can be fused to generate different deep fake high-frequency information modes for the same image, forming a set of deep fake high-frequency information generation methods; Contrastive learning and sample processing submodule: It is configured to build a self-supervised architecture based on contrastive learning and queue-based sample processing pool, which is used to effectively distinguish between images and their own deep fake high-frequency information pattern versions, and enhance the model's ability to identify the difference between true and false images; The deep manifold registration aggregation submodule is configured to establish a deep manifold registration aggregation expressor based on spatial metric learning, implement the deep adversarial mechanism from a macro level, and assist the model to accurately judge the authenticity of samples.

9. The self-supervised deep fake image detection system according to claim 7, characterized in that: The feature processing module specifically includes: Introduce the feature embedding mask strategy and set the image encoder to f e , the corresponding decoder is f d , for the input image I n , through the encoder f e Encoding gets feature z e =f e (I n ); By randomly generating a mask M e , so that the mask strategy acts on the encoded feature z e , z e ⊙M e As decoder f d The reconstructed image I is finally obtained. r =f d (z e ⊙M e ), thereby forcing the decoder to achieve fine-grained restoration and simulating the information completion pressure in the deep fake model generation stage.

10. The self-supervised deep fake image detection system according to claim 7, characterized in that: In the sample processing and loss calculation module, the output sample set randomly selects anchor samples, real samples and deep fake samples, executes the triple pairing sampling strategy, and uses the metric learning loss function for calculation, specifically including: Anchor samples, real samples and deep fake samples are randomly selected from the sample set, and the loss function is used to shorten the distance between the positive sample and the anchor sample and push the negative sample away from the anchor sample, which includes a guidance parameter γ that controls the amplitude of the triplet loss regularization direction and a regularization function Cos that performs constraints in the distance control process to ensure that the distance between samples is within a reasonable range.

11. The self-supervised deep fake image detection system according to claim 7, characterized in that: In the model training and detection module, the cross entropy loss function is configured as: Where N represents the total number of samples, y i represents the true label of sample i, specifying the deep fake sample category as 1 and the normal sample category as 0, p i represents the probability that sample i is predicted to be a deep fake sample category.

12. The self-supervised deep fake image detection system according to claim 8, characterized in that: During the construction of the model building module, the output of the deep fake high-frequency information generation submodule is used as one of the inputs of the contrastive learning and sample processing submodule, and the deep manifold registration aggregation submodule receives the output of the contrastive learning and sample processing submodule and the distributed modeling unit to achieve orderly data flow and collaborative work among modules within the system, thereby improving the system's detection performance for deep fake images.

13. An electronic device comprising: one or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.