AI-Generated Facial Image Authenticity Discrimination Method, Device, and Storage Medium

By constructing a hybrid data set, adopting adaptive enhancement strategies and frequency domain feature extraction, combining deep convolutional networks and knowledge distillation frameworks, an incremental learning mechanism is introduced, which solves the problem that existing technology is difficult to detect high-fidelity AI to generate face images, achieves high-precision and low-complexity detection effects, and has the ability to continuously adapt to new generation technologies.

CN119992630BActive Publication Date: 2025-06-24BEIJING ZIXUAN XINLAI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510466361.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-24
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect forged traces in face images generated by high-fidelity AI, especially when facing complex images generated by diffusion models, detection reliability is significantly reduced.

Method used

By constructing a mixed data set containing real face images and multiple AI-generated face images, adopting adaptive dynamic enhancement strategies and frequency domain feature extraction, combining pre-trained deep convolutional networks and composite scaling strategies, a teacher-student knowledge distillation framework is designed, and an incremental learning mechanism is introduced to achieve the continuous adaptation of the model to the new generation technology.

Benefits of technology

It significantly improves the detection accuracy of AI-generated face images and the generalization ability of models, reduces the computational complexity, realizes robust detection of diversified generation technologies, and has the ability to continuously adapt to new generation technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992630B_ABST
    Figure CN119992630B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, and storage medium for discriminating the authenticity of AI-generated face images, belonging to the technical fields of computer vision and artificial intelligence security. The present invention proposes a detection framework based on the coordination of composite scaling optimization, dynamic enhancement, and incremental learning. Through the composite scaling strategy, the coordinated expansion of the depth, width, and resolution of the network architecture is optimized. Combining with the adaptive data augmentation technology, the training intensity is dynamically adjusted, and the incremental learning mechanism is introduced to enable the model to continuously adapt to new generation technologies and accumulate knowledge. While reducing the computational complexity, the detection robustness against diverse generation technologies is significantly improved. Verification based on authoritative test sets shows that the present invention is superior to existing methods in terms of detection accuracy, generalization ability, and continuous adaptability, and has a better model efficiency balance characteristic, providing reliable technical support for digital content security verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, device, and storage medium for discriminating the authenticity of AI-generated face images, belonging to the technical fields of computer vision and artificial intelligence security. Background Art

[0002] Currently, with the rapid development of technologies such as generative adversarial networks (GANs) and diffusion models (such as the Stable Diffusion model), the visual fidelity of AI-generated face images has approached the real level, posing a severe challenge to traditional detection methods. Methods based on frequency-domain analysis or local texture features are difficult to effectively capture the subtle forgery traces in high-fidelity generated images. Especially when dealing with complex images generated by diffusion models, the detection reliability significantly decreases.

[0003] Although existing deep learning detection schemes have improved the accuracy through complex network structures, they still have significant limitations: large models have a huge number of parameters and are difficult to meet the requirements of real-time detection; fixed-mode data augmentation strategies lack dynamic adaptability, resulting in insufficient generalization ability of the model in small-sample scenarios; the transferability across different generative technologies is poor. For example, a detection system trained for a specific generative model will experience a sharp decline in performance when faced with new generative technologies. In addition, existing methods generally lack the ability of continuous learning and are difficult to adapt to rapidly evolving generative technologies without complete retraining, severely restricting the adaptability and sustainable development of detection systems in real-world environments. Although multi-modal fusion methods can theoretically improve the detection ability, their complex designs significantly increase the computational cost, restricting the practical deployment value. Summary of the Invention

[0004] To improve the accuracy of discriminating the authenticity of AI-generated face images, while balancing the model efficiency, detection accuracy, generalization ability across different generative technologies, and continuous adaptability to new generative technologies, the present invention provides a method, device, and storage medium for discriminating the authenticity of AI-generated face images. The technical solutions are as follows:

[0005] The present invention provides a method for discriminating the authenticity of AI-generated face images. The construction of the discrimination model includes the following steps:

[0006] Step 1: Construct a mixed dataset containing real face images and various AI-generated face images, divide the training set, validation set, and test set using a stratified random sampling strategy, and construct a core sample memory bank;

[0007] Step 2: Implement an adaptive dynamic augmentation strategy for the images in the training set. Apply high-intensity augmentation in the initial stage and gradually reduce the augmentation intensity as the training progresses. Use a differential augmentation strategy for key samples in the memory bank, and extract frequency-domain features as auxiliary input channels;

[0008] Step 3: Use a pre-trained deep convolutional network as the feature extraction backbone, coordinate the expansion ratios of network depth, width, and resolution through a compound scaling strategy, construct an adaptive classification head and embed lightweight attention units;

[0009] Step 4: Design a teacher-student knowledge distillation framework, configure the temperature parameter and soft label weight, set up a feature distillation layer to capture intermediate layer feature representations and transfer them during the incremental learning stage;

[0010] Step 5: Execute the training process, adopt the automatic mixed precision training technique, configure the dynamic learning rate scheduling strategy and combine regularization techniques to prevent overfitting;

[0011] Step 6: Set up a performance monitoring module. When the model performance drops by more than a preset threshold or a new AI generation technology is detected, automatically trigger the incremental learning process. During incremental learning, jointly train the new data and the samples in the memory bank. The total loss consists of the classification loss and the knowledge distillation loss. The total loss is expressed as:

[0012] , where is the total loss function, is the knowledge retention weight factor, is the cross-entropy loss for the new task, is the knowledge distillation loss, and the knowledge distillation loss is expressed as:

[0013] , where, is the predicted output of the teacher model, is the predicted output of the student model;

[0014] Step 7: Calculate the sample importance score based on the gradient information, identify the key samples and update the memory bank; the formula for the importance score is:

[0015] , where represents the gradient of the sample with respect to the model parameters, is the Frobenius norm, represents the input sample, represents the model's predicted output, represents the true label, represents the model parameters;

[0016] Step 8: Calculate the feature similarity between the old and new tasks, and dynamically adjust the depth of the shared layer and the knowledge distillation weight;

[0017] Step 9: Automatically iterate the model version after each incremental learning is completed, and regularly perform model pruning and quantization operations;

[0018] Step 10: Evaluate the model test results and continuously optimize the model structure and training strategy according to the evaluation results.

[0019] Optionally, the composite scaling strategy is expressed as:

[0020] , where represents the global scaling factor, , , respectively represent the expansion ratios of the number of channels, the number of layers, and the input resolution.

[0021] Optionally, the adaptive dynamic enhancement strategy in Step 2 includes:

[0022] When epoch < 10, randomly crop according to the scaling ratio of 0.2 - 0.8, perform ±30% color jitter and 15° perspective transformation, which is expressed as:

[0023] , where represents the original input image, represents the random cropping transformation, represents the color jitter transformation, represents the perspective transformation;

[0024] When 10 ≤ epoch < 30, perform ±15% color jitter and 10° rotation;

[0025] When epoch ≥ 30, perform random horizontal flipping and ±5% brightness adjustment.

[0026] Optionally, Step 5 uses the AdamW optimizer and the cosine annealing learning rate scheduling strategy for training optimization, and the expression is:

[0027] , where , , represents the restart period, represents the current iteration number.

[0028] Optionally, during the training process of Step 5, perform a fast Fourier transform on the input image to extract frequency domain features, and fuse them with the spatial domain features through a cross-modal attention mechanism:

[0029] , where represents the discrete wavelet transform, is the channel concatenation operation, represents the frequency domain feature map, represents the spatial domain feature.

[0030] Optionally, the depth of the shared layer in step 8 is expressed as:

[0031] , where is the total depth of the network, is the similarity between the new and old tasks.

[0032] Optionally, the memory bank sample screening strategy in step 7 is as follows:

[0033] When the importance score satisfies , it is a high-importance old sample, and the retention ratio is 40%;

[0034] When the importance score satisfies , it is a medium-importance sample, and the retention ratio is 30%;

[0035] The retention ratio of the new AI generation technology samples is 30%.

[0036] Optionally, the model classifier adopts a two-level fully connected layer structure, embedding a dynamic regularization mechanism. The first level includes a 512-dimensional fully connected layer and ReLU activation. After dimensionality reduction through a 128-dimensional fully connected layer in the second level, the probability is output through the Sigmoid function. The Dropout probability decays linearly with the number of training epochs , and the formula is:

[0037] , where the initial value of the Dropout probability is set to 0.5.

[0038] The present invention provides an AI-generated face image authenticity discrimination device, including a memory and a processor;

[0039] The memory is used to store computer programs;

[0040] The processor is used to implement the AI-generated face image authenticity discrimination method as described in any one of the above when executing the computer program.

[0041] The present invention provides a computer-readable storage medium, characterized in that a computer program is stored on the storage medium, and when the computer program is executed by a processor, the AI-generated face image authenticity discrimination method as described in any one of the above is implemented.

[0042] The beneficial effects of the present invention are:

[0043] The present invention proposes an efficient, lightweight, and highly generalizable AI-generated face image detection solution. By optimizing the coordinated expansion of the depth, width, and resolution of the network architecture through a composite scaling strategy, combining an adaptive data augmentation technique to dynamically adjust the training intensity, and introducing an incremental learning mechanism to enable the model to continuously adapt to new generation technologies and accumulate knowledge, the computational complexity is reduced while significantly enhancing the detection robustness against diverse generation technologies. Verification based on authoritative test sets (such as DeepFaceGen) shows that the present invention is superior to existing methods in terms of detection accuracy, generalization ability, and continuous adaptability, and has a better model efficiency balance characteristic, providing reliable technical support for digital content security verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 is the workflow diagram of the AI-generated face image authenticity discrimination model of the present invention.

[0046] Figure 2 is the structural diagram of the AI-generated face image authenticity discrimination model of the present invention.

[0047] Figure 3 is the construction flowchart of the AI-generated face image authenticity discrimination model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.

[0049] Embodiment 1:

[0050] This embodiment provides an AI-generated face image authenticity discrimination method, as Figure 3 shown, including the following steps:

[0051] Step 1: Dataset preparation and preprocessing.

[0052] Experimental verification is carried out based on synthetic face datasets from multiple sources, including real face images and face images generated by various AI generation technologies (GAN, diffusion models, etc.) to form a mixed dataset. A stratified random sampling strategy is used to divide the training set, validation set, and test set to ensure the balanced distribution of each type of sample.

[0053] An additional core sample memory bank (CoreMemory) is constructed to store key samples during the incremental learning process, and the initial capacity is set to 20% of the original training set.

[0054] Step 2: Data augmentation and normalization.

[0055] An adaptive dynamic augmentation strategy is implemented for the training set images. At the beginning, high-intensity augmentations (random cropping, color jittering, perspective transformation) are applied, and the augmentation intensity is gradually reduced as the training progresses. A differential augmentation strategy is adopted for the key samples in the memory bank to prevent excessive feature perturbation. All images are uniformly scaled to the standard resolution, and Fourier transform (FFT) is performed to extract frequency domain features, which are used as auxiliary input channels to enhance feature representation.

[0056] Specifically, under the adaptive dynamic augmentation strategy, in the initial stage of training (epoch < 10), high-intensity augmentation operations are applied, including random large-scale cropping (scaling ratio 0.2 - 0.8), ±30% color jittering, and 15° perspective transformation. The mathematical expression is:

[0057] , where represents the original input image, represents the random cropping transformation, represents the color jittering transformation, represents the perspective transformation.

[0058] In the middle stage of training (10 ≤ epoch < 30), it transitions to medium intensity, retaining ±15% color jittering and 10° rotation; in the late stage of training (epoch ≥ 30), only basic augmentations (random horizontal flipping, ±5% brightness adjustment) are performed. This progressive adjustment strategy simulates the evolution of the model's learning state and gradually shifts from emphasizing data diversity to feature stability.

[0059] Step 3: Network initialization and configuration.

[0060] A pre-trained deep convolutional network is used as the feature extraction backbone. The compound scaling strategy is adopted to coordinate the expansion ratios of network depth, width, and resolution. An adaptive classification head is constructed at the end of the backbone network, and lightweight attention units are embedded in key layers to enhance cross-layer feature correlation. The pre-trained weights are loaded, and the shallow convolutional parameters are frozen to retain the general feature extraction ability, while the classification head parameters are specifically initialized.

[0061] Specifically, the compound scaling strategy is expressed as:

[0062] , where represents the global scaling factor, which is determined through experimental optimization; 、 、 They represent the expansion ratios of the number of channels, the number of layers, and the input resolution respectively. The optimal ratios are . The input resolution is preferentially expanded to 256×256 to enhance the ability to capture detailed features. High-frequency information is retained through bilinear interpolation, and then the network width and depth are gradually increased to ensure the efficient extraction of multi-scale features.

[0063] Step 4: Construction of the knowledge distillation framework.

[0064] Design a teacher-student knowledge distillation framework and configure the temperature parameter and the soft label weight. In the initial training stage, only the student model participates in the training; during incremental learning, the previously trained model is set as the teacher model to guide the new model to learn. Set up a feature distillation layer to capture the intermediate layer feature representations and transfer them during the incremental learning stage.

[0065] Step 5: Model training process.

[0066] Execute the training process on a high-performance GPU platform and adopt the mixed-precision training technique to improve efficiency. Configure a dynamic learning rate scheduling strategy and combine appropriate regularization techniques to prevent overfitting. Conduct a hierarchical design of the training parameters to ensure the efficiency of model learning while maintaining stability.

[0067] The classifier design in this embodiment adopts a two-level fully connected layer structure, and embeds a dynamic regularization mechanism to balance feature learning and generalization performance. The first level includes a 512-dimensional fully connected layer and ReLU activation. After dimensionality reduction through a 128-dimensional fully connected layer in the second level, the probability is output through the Sigmoid function. To avoid overfitting, the Dropout probability decreases linearly with the number of training epochs , and the formula is:

[0068] . The initial value of the Dropout probability is set to 0.5, and it decreases by 0.1 every 10 epochs, with a minimum of 0.2, so as to suppress the overfitting risk in the initial stage of training and gradually release the model capacity in the later stage.

[0069] In terms of training optimization, this embodiment adopts the AdamW optimizer (initial learning rate ) and the cosine annealing learning rate scheduling strategy, and its mathematical expression is:

[0070] , where , , and the restart period is [number of epochs]. Combine the automatic mixed precision (AMP) technology. The forward calculation of the backbone network uses FP16 precision, and the classification head maintains FP32 precision to ensure numerical stability. The gradient scaling factor is dynamically adjusted according to the gradient magnitude to avoid calculation overflow.

[0071] Step 6: Incremental learning trigger mechanism.

[0072] Set up a performance monitoring module to start incremental training when the accuracy of the test set drops by more than a preset threshold (3%) or a new type of AI generation technology appears, and the model version is automatically iterated.

[0073] After incremental learning is started, joint training is performed on new data and memory bank samples. The total loss is composed of classification loss and knowledge distillation loss, and the total loss is expressed as:

[0074] , where is the total loss function, is the knowledge retention weight factor, is the cross-entropy loss of the new task, is the knowledge distillation loss, and the knowledge distillation loss is expressed as:

[0075] , where, is the predicted output of the teacher model, is the predicted output of the student model.

[0076] The incremental learning process adopts a progressive fine-tuning strategy. First, freeze the backbone network parameters and only train the newly added classification head. Then, unfreeze the deep feature extraction module for fine-tuning. Finally, jointly optimize the entire network. The learning rate adopts a hierarchical design. The learning rate of the newly added layer is the benchmark value, and the learning rate of the old layer is 0.1 times the benchmark value, effectively preventing forgetting old knowledge due to over-adapting to new data.

[0077] Step 7: Sample importance evaluation.

[0078] After training is completed, calculate the importance scores of all samples based on gradient information to identify key samples. Update the memory bank according to the importance scores, and retain high- and medium-importance old samples and new generation technology samples in proportion. The memory bank capacity is dynamically managed and appropriately expanded as the model complexity increases.

[0079] The calculation method of the importance score of a sample is:

[0080] , where represents the gradient of the sample with respect to the model parameters, is the Frobenius norm, represents the input sample, represents the model prediction output, represents the true label, represents the model parameters.

[0081] The sample screening strategy in this embodiment is as follows:

[0082] (1) High-importance old samples ( ), with a retention ratio of 40%;

[0083] (2) Medium-importance samples ( ), with a retention ratio of 30%;

[0084] (3) The retention ratio of newly generated technology samples is 30%.

[0085] Step 8: Construct an elastic feature adaptation mechanism.

[0086] Calculate the similarity between new and old task features based on the cosine similarity method of feature vectors. Dynamically adjust the depth of the shared layer and the knowledge distillation weight according to the task similarity to optimize the knowledge transfer and retention in the incremental learning process.

[0087] Adopt an elastic feature representation strategy to dynamically adjust the depth of the feature sharing layer between different versions. Its mathematical expression is:

[0088] , where is the depth of the shared layer, is the total depth of the network, is the similarity between new and old tasks (between 0 and 1), which is automatically estimated through feature correlation analysis.

[0089] Step 9: Version management and optimization.

[0090] After each incremental learning is completed, the model version is automatically iterated. Model pruning and quantization operations are performed regularly to optimize the model inference efficiency. Configure a performance rollback mechanism to ensure system stability. Compress the model parameter quantity through techniques such as progressive channel pruning to improve the system response efficiency.

[0091] Step 10: Comprehensive performance analysis.

[0092] Conduct a comprehensive quantitative evaluation of the test results, and calculate indicators such as accuracy, precision, recall, and F1 score. Pay special attention to the performance comparison before and after incremental learning, and evaluate the detection ability and cross-technology generalization ability of the model on different generation technologies. Continuously optimize the model structure and training strategy according to the performance analysis results to improve the overall performance of the system.

[0093] This embodiment constructs an end-to-end and continuously evolving AI-generated face image authenticity discrimination framework. This framework optimizes the multi-scale feature extraction ability through a composite scaling strategy, combines a dynamic adaptive enhancement mechanism to improve the model's adaptability to data distribution changes, and uses dynamic regularization design and mixed-precision training to achieve an efficient and stable optimization process. The introduction of the incremental learning mechanism enables the system to continuously adapt to new generation technologies, significantly reducing the model maintenance cost and enhancing the sustainability of the system in practical applications. Through the collaborative analysis of frequency-domain and spatial-domain features, the system can jointly capture the frequency-domain artifacts and spatial-domain texture anomalies of the generated images, thereby achieving robust detection of diverse generation technologies. This technical paradigm based on composite architecture optimization, dynamic strategy collaboration, and incremental knowledge inheritance provides a new solution for generated image detection. While ensuring high discrimination accuracy, it significantly improves the model's computational efficiency, cross-technology generalization ability, and continuous adaptability, providing reliable and sustainable technical support for the digital content security field.

[0094] Embodiment 2:

[0095] To further prove the detection performance of the method of the present invention for AI-generated face images, this embodiment adopts five core indicators: accuracy, precision, recall, F1 score, and cross-entropy loss (BCEWithLogitsLoss).

[0096] Accuracy is used to measure the overall prediction accuracy of the model, defined as the ratio of the number of correctly classified samples to the total number of samples. The calculation formula is:

[0097] , where (True Positive) represents the number of AI-generated images correctly identified, (True Negative) is the number of real images correctly identified, (False Positive) and (False Negative) respectively represent the number of real images misjudged as generated images and the number of generated images misjudged as real images. Accuracy reflects the comprehensive discrimination ability of the model on global samples.

[0098] Precision focuses on the prediction reliability of the model for the "AI-generated image" category, defined as the ratio of correctly identified generated images to all images predicted as generated images:

[0099] , A high precision rate indicates that the model has a low false positive rate when determining generated images, and is suitable for scenarios sensitive to false positives (such as content review).

[0100] Recall measures the detection ability of the model for generated images, calculated as the proportion of correctly identified generated images to the total number of actual generated images:

[0101] , A high recall rate means that the model can effectively reduce missed detections and is suitable for scenarios sensitive to missed detections (such as anti-fraud detection).

[0102] The F1 Score is the harmonic mean of precision and recall, used to comprehensively evaluate the classification balance of the model:

[0103] , This metric is more valuable when the class distribution is imbalanced (such as when the proportion of generated images is significantly lower than that of real images), avoiding the one-sidedness of a single metric.

[0104] Binary Cross Entropy Loss directly reflects the matching degree between the model's predicted probability and the true label, and its definition is:

[0105] , where represents the true label (0 for real images, 1 for generated images), is the generation probability output by the model, is the total number of samples. The lower the loss value, the closer the model's prediction result is to the true distribution, and the better the convergence.

[0106] Through the collaborative analysis of the above five metrics, the performance of the model in terms of detection accuracy, classification balance, and training stability can be comprehensively evaluated, providing a quantitative basis for model optimization and deployment in actual application scenarios.

[0107] The experimental environment of this embodiment is shown in Table 1:

[0108] Table 1 Detailed parameters for network model training

[0109]

[0110] The results of the comparative experiment are shown in Table 2:

[0111] Table 2 Comparison of core performance metrics

[0112]

[0113] Experimental results show that the method of the present invention demonstrates significant performance advantages on the DeepFaceGen benchmark test set. Compared with the ResNet-50 baseline model, the accuracy rate is increased by 13.20 percentage points to 95.50%, and the F1 score is increased by 13.78 percentage points to 95.38%, verifying the collaborative effectiveness of the compound scaling strategy and the adaptive enhancement mechanism. It is worth noting that the false positive rate of the model is as low as 1.60% (corresponding to a precision of 98.40%), indicating its extremely high reliability in real image discrimination. This feature is crucial for applications in highly sensitive scenarios such as judicial forensics and financial identity authentication, and can effectively avoid systemic risks caused by misjudgment.

[0114] In response to the challenge of generating technology diversity, the model still maintains an average accuracy rate of 93.7% on an independent test set containing 5 untrained generation methods (including the latest technologies released in 2024), as shown in Table 3.

[0115] Table 3 Detection performance of different generation technologies

[0116]

[0117] Among them, the accuracy rates of 93.5% and 87.6% are respectively achieved for Stable Diffusion v3.5 and the Transformer-based generation method (Tech-X). Feature visualization analysis shows that the model of the present invention captures the common forgery traces of different generation technologies through the synergistic effect of frequency domain artifact detection (abnormal high-frequency noise distribution) and spatial domain texture continuity analysis (such as skin micro-texture consistency). In particular, the recall rate of the model of the present invention for the Tech-X technology reaches 84.3%, which is 15.8 percentage points higher than that of the baseline model, indicating that its ability to control the risk of missed detection of new generation technologies has been significantly enhanced. This advantage stems from the generalization and capture ability of the frequency domain feature module for unseen artifact patterns. For example, the low-frequency components separated by discrete wavelet transform (DWT) can effectively identify structural distortions, while the high-frequency components are sensitive to local noise anomalies.

[0118] Table 4 Incremental learning performance evaluation

[0119]

[0120] The introduction of the incremental learning module has significantly improved the model's adaptability to new generation technologies. As shown in Table 4, in the case of only using 10% of the training data, incremental learning has increased the detection accuracy of the model for the latest Tech-X technology from 76.3% of the basic model to 87.6%, with an increase of 11.3 percentage points. More importantly, this adaptation process only takes 4 hours, which is 12 times more efficient than the traditional retraining method (48 hours), and the detection performance for known generation technologies only drops by 0.3 percentage points, effectively solving the problem of catastrophic forgetting.

[0121] Table 5 Comparison Table of Ablation Experiment Results

[0122]

[0123] By systematically conducting ablation experiments to quantify the contributions of each module, as shown in Table 5, the removal of the dynamic Dropout mechanism has led to a 4.7% decrease in accuracy in the small-sample scenario (10,000 training data) (90.8% vs. 95.5%), verifying the effectiveness of its strategy of balancing model capacity and generalization through probability decay; after disabling the compound scaling strategy, the number of model parameters has increased to 28M and the inference speed has decreased by 32%, but the accuracy has only decreased by 3.4% (92.1% vs. 95.5%), indicating that this strategy has significantly optimized the computational efficiency while maintaining performance; while turning off the frequency domain analysis module has caused the detection accuracy of the diffusion model to decrease by 9.2% (86.3% vs. 95.5%), highlighting the irreplaceability of frequency domain features in identifying high-fidelity generated images. The removal of the incremental learning module has led to a significant reduction in the adaptation efficiency for new generation technologies, with the accuracy decreasing by 11.3 percentage points (76.3% vs. 87.6%).

[0124] Table 6 Comparison Table of Performance with Mainstream Methods

[0125]

[0126] As shown in Table 6, compared with the current mainstream methods, the detection model constructed by the present invention shows significant advantages in both performance and generalization. For example, compared with the traditional method based on frequency domain Fourier transform (average accuracy of 58.7%), the accuracy of the model of the present invention has increased by 36.8 percentage points; compared with the latest multi-modal fusion method (such as FusionNet), the F1 score has increased by 7.2 percentage points (95.38% vs. 88.18%), and the number of parameters has decreased by 41%. More importantly, the advantages of the model of the present invention in the adaptation efficiency of new generation technologies are particularly obvious, shortening the adaptation cycle from 48 - 72 hours of traditional methods to 4 hours, enabling the detection technology to quickly keep up with the iterative upgrade of generation technologies.

[0127] The synergy between knowledge distillation and the sample memory bank is the key to achieving rapid adaptation. As shown in Table 7, by comparing different memory strategies, it is found that the selective memory mechanism based on sample importance improves the adaptation accuracy by 3.7 percentage points compared to random sampling (87.6% vs. 83.9%), and the adaptation speed is increased by 35% under the same storage overhead (20% of the training set). This verifies the innovative advantages of the detection model of the present invention in identifying key samples and effectively transferring knowledge.

[0128] Table 7 Comparison Table of Memory Strategies

[0129]

[0130] The experimental results show that the method of the present invention provides a new technical paradigm for the generated image detection task, achieving a better balance among accuracy, efficiency, and generalization. In particular, the introduction of the incremental learning mechanism enables the model to quickly adapt to new generation technologies, realizing the co-evolution of detection technology and generation technology. This breakthrough result has important academic value and engineering application potential, providing a technical foundation for sustainable development in the field of digital content security.

[0131] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.

[0132] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for distinguishing the authenticity of AI-generated facial images, characterized in that: The construction of the discriminant model includes the following steps: Step 1: Build a mixed dataset containing real face images and various AI-generated face images, use a stratified random sampling strategy to divide the training set, validation set, and test set, and build a core sample memory library; Step 2: Implement an adaptive dynamic enhancement strategy on the images of the training set, apply high-intensity enhancement in the early stage, gradually reduce the enhancement intensity as the training progresses, adopt a differentiated enhancement strategy for key samples in the memory library, and extract frequency domain features as auxiliary input channels; Step 3: Use the pre-trained deep convolutional network as the feature extraction backbone, coordinate the expansion ratio of network depth, width and resolution through a composite scaling strategy, build an adaptive classification head and embed a lightweight attention unit; Step 4: Design the teacher-student knowledge distillation framework, configure temperature parameters and soft label weights, set up feature distillation layers to capture intermediate layer feature representations and pass them in the incremental learning phase; Step 5: Execute the training process, use automatic mixed precision training technology, configure dynamic learning rate scheduling strategy and combine regularization technology to prevent overfitting; Step 6: Set up a performance monitoring module. When the model performance drops beyond a preset threshold or a new AI generation technology is detected, the incremental learning process is automatically triggered. In incremental learning, new data and memory samples are jointly trained, and the total loss consists of classification loss and knowledge distillation loss. Step 7: Calculate the sample importance score based on the gradient information, identify key samples and update the memory library; Step 8: Calculate the similarity of new and old task features, and dynamically adjust the shared layer depth and knowledge distillation weight; Step 9: After each incremental learning is completed, the model version is automatically iterated, and model pruning and quantization operations are performed regularly; Step 10: Evaluate the model test results and continuously optimize the model structure and training strategy based on the evaluation results; The compound scaling strategy is expressed as: in, represents the global scaling factor, , , Respectively represent the expansion ratio of the number of channels, number of layers and input resolution; The adaptive dynamic enhancement strategy of step 2 includes: When epoch < 10, random cropping is performed with a scaling ratio of 0.2 to 0.8, ±30% color jitter, and 15° perspective transformation, expressed as: in, represents the original input image, represents a random cropping transformation, Indicates color dithering transformation, Represents perspective transformation; When 10≤epoch<30, perform ±15% color dithering and 10° rotation; When epoch ≥ 30, random horizontal flipping and ±5% brightness adjustment are performed.

2. The method for distinguishing the authenticity of AI-generated facial images according to claim 1, characterized in that: Step 5 uses the AdamW optimizer and the cosine annealing learning rate scheduling strategy for training optimization, and the expression is: in, , , Indicates the restart cycle, Indicates the current iteration number.

3. The method for distinguishing the authenticity of AI-generated facial images according to claim 1, characterized in that: During the training process of step 5, the input image is subjected to a fast Fourier transform to extract frequency domain features, which are then fused with the spatial domain features through a cross-modal attention mechanism: in, represents discrete wavelet transform, For channel splicing operations, represents the frequency domain feature map, Represents spatial domain characteristics.

4. The method for distinguishing authenticity of AI-generated facial images according to claim 1, characterized in that: The shared layer depth in step 8 is expressed as: in, is the total depth of the network, is the similarity between the new and old tasks.

5. The method for distinguishing authenticity of AI-generated facial images according to claim 1, characterized in that: The memory bank sample screening strategy of step 7 is: When the importance score satisfies When , it is an old sample with high importance, and the retention ratio is 40%; When the importance score satisfies , which is a medium-importance sample, the retention ratio is 30%; The sample retention rate for new AI generation technology is 30%.

6. The AI-generated facial image authenticity determination method according to claim 1, characterized in that: The model classifier adopts a two-level fully connected layer structure and embeds a dynamic regularization mechanism. The first level contains a 512-dimensional fully connected layer with ReLU activation. The second level is reduced by a 128-dimensional fully connected layer and then outputs the probability through the Sigmoid function. The Dropout probability With training rounds Linear attenuation, the formula is: Among them, the Dropout probability The initial value of is set to 0.

5.

7. An AI-generated facial image authenticity determination device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the AI-generated facial image authenticity determination method as described in any one of claims 1 to 6 when executing the computer program.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for distinguishing the authenticity of an AI-generated facial image as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Generating neural networks

    CN117744762A

  • Image forgery detection method based on self-supervised contrast learning

    CN119693354A