AI (Artificial Intelligence) generated face image authenticity judgment method and device and storage medium
By constructing a hybrid data set, adaptive enhancement strategy and frequency domain feature extraction, combined with deep convolution networks and knowledge distillation frameworks, an incremental learning mechanism was introduced, which solved the problem that existing technology is difficult to detect face images generated by high-fidelity AI, and achieved high-precision, strong generalization and good adaptability detection effects.
Patent Information
- Application Number
- CN202510466361.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art is difficult to effectively detect forged traces in face images generated by high-fidelity AI, especially when facing complex images generated by diffusion models, detection reliability is significantly reduced.
By constructing a mixed data set containing real face images and multiple AI-generated face images, adopting adaptive dynamic enhancement strategies and frequency domain feature extraction, combining pre-trained deep convolutional networks and composite scaling strategies, a teacher-student knowledge distillation framework is designed, and an incremental learning mechanism is introduced to achieve the continuous adaptation of the model to the new generation technology.
It significantly improves the detection accuracy of AI-generated face images and the generalization ability of models, enhances the adaptability and continuous learning ability to new generation technologies, and reduces the computational complexity.
Smart Images

Figure CN119992630A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device and storage medium for distinguishing the authenticity of AI-generated facial images, and belongs to the field of computer vision and artificial intelligence security technology. Background Art
[0002] Currently, the rapid development of technologies such as generative adversarial networks (GANs) and diffusion models (such as the stable diffusion model) has made AI-generated facial images close to the real level in terms of visual fidelity, and traditional detection methods are facing severe challenges. Methods based on frequency domain analysis or local texture features are difficult to effectively capture subtle traces of forgery in high-fidelity generated images, especially when dealing with complex images generated by diffusion models, and the detection reliability is significantly reduced.
[0003] Although existing deep learning detection solutions have improved accuracy through complex network structures, they still have significant limitations: large models have huge parameters and cannot meet real-time detection needs; fixed-mode data enhancement strategies lack dynamic adaptability, resulting in insufficient generalization of models in small sample scenarios; and poor transferability across generative technologies. For example, detection systems trained for specific generative models will experience a sharp drop in performance when faced with new generative technologies. In addition, existing methods generally lack the ability to continuously learn and are unable to adapt to rapidly iterating generative technologies without complete retraining, which severely limits the adaptability and sustainable development of detection systems in actual environments. Although multimodal fusion methods can theoretically improve detection capabilities, their complex design significantly increases computational costs, restricting their actual deployment value. Summary of the invention
[0004] In order to improve the accuracy of AI-generated facial image authenticity discrimination, while balancing model efficiency, detection accuracy, generalization ability across generation technologies, and continuous adaptability to new generation technologies, the present invention provides an AI-generated facial image authenticity discrimination method, device, and storage medium, and the technical solution is as follows: The present invention provides a method for distinguishing the authenticity of a face image generated by AI, and the construction of a discrimination model includes the following steps: Step 1: Build a mixed dataset containing real face images and multiple AI-generated face images, use a stratified random sampling strategy to divide the training set, validation set, and test set, and build a core sample memory library; Step 2: Implement an adaptive dynamic enhancement strategy on the images of the training set, apply high-intensity enhancement in the early stage, gradually reduce the enhancement intensity as the training progresses, adopt a differentiated enhancement strategy for key samples in the memory library, and extract frequency domain features as auxiliary input channels; Step 3: Use the pre-trained deep convolutional network as the feature extraction backbone, coordinate the expansion ratio of network depth, width and resolution through a composite scaling strategy, build an adaptive classification head and embed a lightweight attention unit; Step 4: Design the teacher-student knowledge distillation framework, configure temperature parameters and soft label weights, set up feature distillation layers to capture intermediate layer feature representations and pass them in the incremental learning phase; Step 5: Execute the training process, use automatic mixed precision training technology, configure dynamic learning rate scheduling strategy and combine regularization technology to prevent overfitting; Step 6: Set up a performance monitoring module. When the model performance drops beyond a preset threshold or a new AI generation technology is detected, the incremental learning process is automatically triggered. In incremental learning, new data and memory samples are jointly trained. The total loss consists of classification loss and knowledge distillation loss. The total loss is expressed as: ,in is the total loss function, The weight factor for knowledge preservation, is the cross entropy loss for the new task, is the knowledge distillation loss, which is expressed as: ,in, is the predicted output of the teacher model, is the prediction output of the student model; Step 7: Calculate the sample importance score based on the gradient information, identify key samples and update the memory library; the importance score calculation formula is: ,in represents the gradient of the sample to the model parameters, is the Frobenius norm, represents the input sample, represents the model prediction output, represents the true label, represents the model parameters; Step 8: Calculate the similarity of new and old task features, and dynamically adjust the shared layer depth and knowledge distillation weight; Step 9: After each incremental learning is completed, the model version is automatically iterated, and model pruning and quantization operations are performed regularly; Step 10: Evaluate the model test results and continuously optimize the model structure and training strategy based on the evaluation results.
[0005] Optionally, the composite scaling strategy is expressed as: ,in, represents the global scaling factor, , , They represent the expansion ratio of the number of channels, number of layers, and input resolution respectively.
[0006] Optionally, the adaptive dynamic enhancement strategy of step 2 includes: When epoch < 10, random cropping is performed with a scaling ratio of 0.2 to 0.8, ±30% color jitter, and 15° perspective transformation, expressed as: ,in, represents the original input image, represents a random cropping transformation, Indicates color dithering transformation, Represents perspective transformation; When 10≤epoch<30, perform ±15% color dithering and 10° rotation; When epoch ≥ 30, random horizontal flipping and ±5% brightness adjustment are performed.
[0007] Optionally, the step 5 uses the AdamW optimizer and the cosine annealing learning rate scheduling strategy for training optimization, and the expression is: ,in, , , Indicates the restart cycle, Indicates the current iteration number.
[0008] Optionally, during the training process of step 5, the input image is subjected to a fast Fourier transform to extract frequency domain features, which are then fused with the spatial domain features through a cross-modal attention mechanism: ,in, represents discrete wavelet transform, For channel splicing operations, represents the frequency domain feature map, Represents spatial domain characteristics.
[0009] Optionally, the shared layer depth in step 8 is expressed as: ,in, is the total depth of the network, is the similarity between the new and old tasks.
[0010] Optionally, the memory bank sample screening strategy in step 7 is: When the importance score satisfies When , it is an old sample with high importance, and the retention ratio is 40%; When the importance score satisfies , which is a medium-importance sample, the retention ratio is 30%; The sample retention rate for new AI generation technology is 30%.
[0011] Optionally, the model classifier adopts a two-level fully connected layer structure and embeds a dynamic regularization mechanism. The first level contains a 512-dimensional fully connected layer with ReLU activation. The second level is reduced by a 128-dimensional fully connected layer and then outputs the probability through the Sigmoid function. Dropout probability With training rounds Linear attenuation, the formula is: , where the Dropout probability The initial value of is set to 0.5.
[0012] The present invention provides an AI-generated facial image authenticity discrimination device, comprising a memory and a processor; The memory is used to store computer programs; The processor is used to implement the AI-generated facial image authenticity determination method as described in any of the above items when executing the computer program.
[0013] The present invention provides a computer-readable storage medium, characterized in that a computer program is stored on the storage medium. When the computer program is executed by a processor, the AI-generated facial image authenticity determination method as described in any of the above items is implemented.
[0014] The beneficial effects of the present invention are: The present invention proposes an efficient, lightweight and highly generalized AI-generated face image detection solution. It optimizes the depth, width and resolution of the network architecture through a composite scaling strategy, dynamically adjusts the training intensity in combination with adaptive data enhancement technology, and introduces an incremental learning mechanism to achieve continuous adaptation and knowledge accumulation of the model to new generation technologies. While reducing computational complexity, it significantly improves the detection robustness of diversified generation technologies. Verification based on authoritative test sets (such as DeepFaceGen) shows that the present invention is superior to existing methods in detection accuracy, generalization ability and continuous adaptability, and has a better model efficiency balance characteristic, providing reliable technical support for digital content security verification. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 This is a workflow diagram of the AI-generated face image authenticity discrimination model of the present invention.
[0017] Figure 2This is a structural diagram of the AI-generated face image authenticity discrimination model of the present invention.
[0018] Figure 3 It is a flow chart for constructing an AI-generated face image authenticity discrimination model of the present invention. DETAILED DESCRIPTION
[0019] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0020] Embodiment 1: This embodiment provides a method for distinguishing the authenticity of AI-generated facial images. Figure 3 As shown, the following steps are included: Step 1: Dataset preparation and preprocessing.
[0021] Experimental verification is conducted based on synthetic face datasets from various sources, including real face images and face images generated by various AI generation technologies (GAN, diffusion model, etc.), forming a mixed data set. A stratified random sampling strategy is used to divide the training set, validation set, and test set to ensure a balanced distribution of samples of all types.
[0022] An additional core sample memory (CoreMemory) is constructed to store key samples in the incremental learning process, and the initial capacity is set to 20% of the original training set.
[0023] Step 2: Data enhancement and standardization.
[0024] An adaptive dynamic enhancement strategy is implemented for the training set images. High-intensity enhancement (random cropping, color jitter, perspective transformation) is applied in the early stage, and the enhancement intensity is gradually reduced as the training progresses. A differentiated enhancement strategy is adopted for key samples in the memory library to prevent excessive feature perturbation. All images are uniformly scaled to a standard resolution, and Fourier transform (FFT) is performed to extract frequency domain features as an auxiliary input channel to enhance feature expression.
[0025] Specifically, under the adaptive dynamic enhancement strategy, this embodiment applies high-intensity enhancement operations in the early stage of training (epoch<10), including random large-scale cropping (scaling ratio 0.2~0.8), ±30% color jitter and 15° perspective transformation, which can be mathematically expressed as: ,in, represents the original input image, represents a random cropping transformation, Indicates color dithering transformation, Represents a perspective transformation.
[0026] In the middle of training (10≤epoch<30), the training is transitioned to medium intensity, retaining ±15% color jitter and 10° rotation; in the late training (epoch≥30), only basic enhancements (random horizontal flip, ±5% brightness adjustment) are performed. This gradual adjustment strategy gradually shifts from emphasizing data diversity to feature stability by simulating the evolution of the model's learning state.
[0027] Step 3: Network initialization and configuration.
[0028] A pre-trained deep convolutional network is used as the feature extraction backbone. The expansion ratio of network depth, width and resolution is coordinated through a composite scaling strategy. An adaptive classification head is built at the end of the backbone network, and lightweight attention units are embedded in the key layers to enhance cross-layer feature associations. The pre-trained weights are loaded, the shallow convolution parameters are frozen to retain the general feature extraction capability, and the classification head parameters are specifically initialized.
[0029] Specifically, the compound scaling strategy is expressed as: ,in, represents the global scaling factor, which is determined through experimental optimization; , , Respectively represent the expansion ratio of the number of channels, the number of layers and the input resolution. The optimal ratio is The input resolution is first expanded to 256×256 to enhance the ability to capture detailed features, and high-frequency information is retained through bilinear interpolation. Then the network width and depth are gradually increased to ensure efficient extraction of multi-scale features.
[0030] Step 4: Construction of knowledge distillation framework.
[0031] Design the teacher-student knowledge distillation framework, configure temperature parameters and soft label weights. In the initial training phase, only the student model participates in the training; during incremental learning, the previously trained model is set as the teacher model to guide the learning of the new model. Set up the feature distillation layer to capture the intermediate layer feature representation and pass it in the incremental learning phase.
[0032] Step 5: Model training process.
[0033] The training process is performed on a high-performance GPU platform, and mixed precision training technology is used to improve efficiency. A dynamic learning rate scheduling strategy is configured, combined with appropriate regularization technology to prevent overfitting. The training parameters are designed in layers to ensure that the model learns efficiently while maintaining stability.
[0034] The classifier design of this embodiment adopts a two-level fully connected layer structure and embeds a dynamic regularization mechanism to balance feature learning and generalization performance. The first level contains a 512-dimensional fully connected layer with ReLU activation, and the second level is reduced by a 128-dimensional fully connected layer and then outputs the probability through a Sigmoid function. To avoid overfitting, the Dropout probability With training rounds Linear attenuation, the formula is: , Dropout probability The initial value of is set to 0.5, and is reduced by 0.1 every 10 rounds to a minimum of 0.2, thereby suppressing the risk of overfitting in the early stage of training and gradually releasing the model capacity in the later stage.
[0035] In terms of training optimization, this embodiment uses the AdamW optimizer (initial learning rate ) and the cosine annealing learning rate scheduling strategy, which is mathematically expressed as: ,in , , restart cycle Combined with the automatic mixed precision (AMP) technology, the forward calculation of the backbone network uses FP16 precision, the classification head maintains FP32 precision to ensure numerical stability, and the gradient scaling factor is dynamically adjusted according to the gradient amplitude to avoid calculation overflow.
[0036] Step 6: Incremental learning trigger mechanism.
[0037] Set up a performance monitoring module. When the accuracy of the test set drops below the preset threshold (3%) or a new AI generation technology emerges, incremental training is started and the model version is automatically iterated.
[0038] After incremental learning is started, the new data and memory samples are jointly trained. The total loss is composed of classification loss and knowledge distillation loss. The total loss is expressed as: ,in is the total loss function, The weight factor for knowledge preservation, is the cross entropy loss for the new task, is the knowledge distillation loss, which is expressed as: ,in, is the predicted output of the teacher model, is the predicted output of the student model.
[0039] The incremental learning process adopts a progressive fine-tuning strategy. First, the backbone network parameters are frozen to train only the newly added classification head, then the deep feature extraction module is unfrozen for fine-tuning, and finally the whole network is jointly optimized. The learning rate adopts a hierarchical design, with the learning rate of the newly added layer being the baseline value and the learning rate of the old layer being 0.1 times the baseline value, effectively preventing over-adaptation to new data and forgetting old knowledge.
[0040] Step 7: Sample importance assessment.
[0041] After training is completed, the importance scores of all samples are calculated based on the gradient information to identify key samples. The memory bank is updated according to the importance scores, and old samples with high and medium importance and new generation technology samples are retained in proportion. The memory bank capacity is dynamically managed and appropriately expanded as the model complexity increases.
[0042] The importance score of a sample is calculated as: ,in represents the gradient of the sample to the model parameters, is the Frobenius norm, represents the input sample, represents the model prediction output, represents the true label, Represents model parameters.
[0043] The sample screening strategy in this embodiment is as follows: (1) High-importance old samples ( ), the retention ratio is 40%; (2) Medium importance samples ( ), the retention ratio is 30%; (3) The retention ratio of newly generated technology samples is 30%.
[0044] Step 8: Build a flexible feature adaptation mechanism.
[0045] Calculate the similarity of new and old task features based on the feature vector cosine similarity method. Dynamically adjust the shared layer depth and knowledge distillation weight according to task similarity to optimize knowledge transfer and retention in the incremental learning process.
[0046] The elastic feature representation strategy is adopted to dynamically adjust the feature sharing layer depth between different versions. Its mathematical expression is: ,in is the shared layer depth, is the total depth of the network, is the similarity between the new and old tasks (between 0 and 1), which is automatically estimated through feature correlation analysis.
[0047] Step 9: Version management and optimization.
[0048] After each incremental learning is completed, the model version is automatically iterated, and model pruning and quantization operations are performed regularly to optimize model reasoning efficiency. Configure performance rollback mechanism to ensure system stability. Use technologies such as progressive channel pruning to compress the model parameters and improve system response efficiency.
[0049] Step 10: Comprehensive performance analysis.
[0050] Conduct a comprehensive quantitative evaluation of the test results, and calculate indicators such as accuracy, precision, recall, and F1 score. Pay special attention to the performance comparison before and after incremental learning, and evaluate the model's detection ability on different generation technologies and cross-technology generalization ability. Continuously optimize the model structure and training strategy based on the performance analysis results to improve the overall performance of the system.
[0051] This embodiment constructs an end-to-end and sustainably evolving AI-generated face image authenticity discrimination framework. The framework optimizes the multi-scale feature extraction capability through a composite scaling strategy, combines a dynamic adaptive enhancement mechanism to improve the model's adaptability to changes in data distribution, and uses dynamic regularization design and mixed precision training to achieve an efficient and stable optimization process. The introduction of the incremental learning mechanism enables the system to continuously adapt to new generation technologies, significantly reduces model maintenance costs, and improves the sustainability of the system in practical applications. Through the collaborative analysis of frequency-space domain features, the system can jointly capture the frequency domain artifacts and spatial domain texture anomalies of the generated image, thereby achieving robust detection of diversified generation technologies. This technical paradigm based on composite architecture optimization, dynamic strategy collaboration, and incremental knowledge inheritance provides a new solution for generated image detection. While ensuring high discrimination accuracy, it significantly improves the model's computational efficiency, cross-technology generalization capabilities, and continuous adaptability, providing reliable and sustainable technical support for the field of digital content security.
[0052] Embodiment 2: In order to further demonstrate the detection performance of the method of the present invention for AI-generated facial images, this embodiment adopts five core indicators: accuracy, precision, recall, F1 score and cross entropy loss (BCEWithLogitsLoss).
[0053] Accuracy is used to measure the accuracy of the overall prediction of the model. It is defined as the ratio of correctly classified samples to the total number of samples. The calculation formula is: ,in, (True Positive) indicates the number of AI-generated images that were correctly identified, (True Negative) is the number of true images that are correctly identified. (False Positive) and (False Negative) indicates the number of real images that are mistakenly identified as generated images and the number of generated images that are mistakenly identified as real images. The accuracy reflects the comprehensive discrimination ability of the model on global samples.
[0054] Precision focuses on the model’s prediction reliability for the “AI generated images” category, which is defined as the proportion of correctly identified generated images to all predicted generated images: ,High precision indicates that the model has a low false alarm rate when determining the generated image, and is suitable for scenarios that are sensitive to false alarms (such as content review).
[0055] Recall measures the model's ability to detect generated images, calculated as the ratio of correctly identified generated images to the total number of actually generated images: ,High recall rate means that the model can effectively reduce missed detections and is suitable for ,scenarios that are sensitive to missed detections (such as anti-fraud detection).
[0056] The F1 score is the harmonic mean of precision and recall, which is used to comprehensively evaluate the classification balance of the model: This indicator is more valuable for reference when the category distribution is unbalanced (for example, the proportion of generated images is significantly lower than that of real images), avoiding the one-sidedness of a single indicator.
[0057] The cross entropy loss (BCEWithLogitsLoss) directly reflects the degree of match between the model's predicted probability and the true label, which is defined as: ,in, represents the true label (0 for real image, 1 for generated image), is the generation probability of the model output, is the total number of samples. The lower the loss value, the closer the model prediction result is to the true distribution and the better the convergence.
[0058] Through the collaborative analysis of the above five indicators, the performance of the model in terms of detection accuracy, classification balance and training stability can be comprehensively evaluated, providing a quantitative basis for model optimization and deployment in actual application scenarios.
[0059] The experimental environment of this embodiment is shown in Table 1: Table 1 Network model training detailed parameters
[0060] The comparative experimental results are shown in Table 2: Table 2 Comparison of core performance indicators
[0061] Experimental results show that the proposed method exhibits significant performance advantages on the DeepFaceGen benchmark test set. Compared with the ResNet-50 baseline model, the accuracy rate increased by 13.20 percentage points to 95.50%, and the F1 score increased by 13.78 percentage points to 95.38%, verifying the synergistic effectiveness of the compound scaling strategy and the adaptive enhancement mechanism. It is worth noting that the false positive rate of the model is as low as 1.60% (corresponding to an accuracy of 98.40%), indicating that it has extremely high reliability in real image discrimination. This feature is crucial for applications in highly sensitive scenarios such as judicial evidence collection and financial identity authentication, and can effectively avoid systemic risks caused by misjudgment.
[0062] In response to the challenge of diversity in generation techniques, the model still maintains an average accuracy of 93.7% on an independent test set containing 5 untrained generation methods (including the latest techniques released in 2024), as shown in Table 3.
[0063] Table 3 Detection performance of different generation technologies
[0064] Among them, the accuracy rates of Stable Diffusion v3.5 and Transformer-based generation method (Tech-X) were 93.5% and 87.6%, respectively. Feature visualization analysis shows that the model of the present invention captures the common forgery traces of different generation technologies through the synergy of frequency domain artifact detection (abnormal high-frequency noise distribution) and spatial domain texture continuity analysis (such as skin microtexture consistency). In particular, the recall rate of the model of the present invention for Tech-X technology reached 84.3%, an increase of 15.8 percentage points over the baseline model, indicating that its ability to control the risk of missed detection of new generation technologies has been significantly enhanced. This advantage stems from the generalization capture ability of the frequency domain feature module for unseen artifact patterns. For example, the low-frequency components separated by discrete wavelet transform (DWT) can effectively identify structural distortions, while the high-frequency components are sensitive to local noise anomalies.
[0065] Table 4 Incremental learning performance evaluation
[0066] The introduction of the incremental learning module significantly improves the model's adaptability to new generation technologies. As shown in Table 4, with only 10% of the training data, incremental learning increases the model's detection accuracy for the latest Tech-X technology from 76.3% of the basic model to 87.6%, an increase of 11.3 percentage points. More importantly, this adaptation process takes only 4 hours, which is 12 times more efficient than the traditional retraining method (48 hours). At the same time, the detection performance of known generation technologies only drops by 0.3 percentage points, effectively solving the problem of catastrophic forgetting.
[0067] Table 5. Comparison of ablation experiment results
[0068] Through systematic ablation experiments, the contribution of each module is quantified. As shown in Table 5, the removal of the dynamic Dropout mechanism leads to a 4.7% decrease in accuracy (90.8% vs. 95.5%) in the small sample scenario (10,000 training data), which verifies the effectiveness of balancing model capacity and generalization through the probability decay strategy; after disabling the compound scaling strategy, the model parameters increase to 28M and the inference speed decreases by 32%, but the accuracy only decreases by 3.4% (92.1% vs. 95.5%), indicating that the strategy significantly optimizes the computational efficiency while maintaining performance; and disabling the frequency domain analysis module reduces the diffusion model detection accuracy by 9.2% (86.3% vs. 95.5%), highlighting the irreplaceable role of frequency domain features in identifying high-fidelity generated images. The removal of the incremental learning module leads to a significant reduction in the efficiency of adaptation to new generation techniques, with an accuracy drop of 11.3 percentage points (76.3% vs. 87.6%).
[0069] Table 6 Performance comparison with mainstream methods
[0070] As shown in Table 6, compared with the current mainstream methods, the detection model constructed by the present invention shows significant advantages in both performance and generalization. For example, compared with the traditional method based on frequency domain Fourier transform (average accuracy of 58.7%), the accuracy of the model of the present invention is improved by 36.8 percentage points; compared with the latest multimodal fusion method (such as FusionNet), the F1 score is improved by 7.2 percentage points (95.38% vs. 88.18%), and the number of parameters is reduced by 41%. More importantly, the model of the present invention has an obvious advantage in the efficiency of adapting to new generation technologies, shortening the adaptation cycle from 48-72 hours of traditional methods to 4 hours, enabling detection technology to quickly follow the iterative upgrade of generation technology.
[0071] The synergy between knowledge distillation and sample memory is the key to achieving rapid adaptation. As shown in Table 7, by comparing different memory strategies, it is found that the selective memory mechanism based on sample importance improves the adaptation accuracy by 3.7 percentage points compared with random sampling (87.6% vs. 83.9%), and with the same storage overhead (20% of the training set), the adaptation speed is increased by 35%. This verifies the innovative advantages of the detection model of the present invention in identifying key samples and effectively transferring knowledge.
[0072] Table 7 Comparison of memory strategies
[0073] Experimental results show that the proposed method provides a new technical paradigm for the task of generating image detection, achieving a better balance between accuracy, efficiency and generalization. In particular, the introduction of the incremental learning mechanism enables the model to quickly adapt to new generation technologies, achieving the co-evolution of detection technology and generation technology. This breakthrough has important academic value and engineering application potential, and provides a technical foundation for sustainable development in the field of digital content security.
[0074] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.
[0075] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for distinguishing the authenticity of AI-generated facial images, characterized in that: The construction of the discriminant model includes the following steps: Step 1: Build a mixed dataset containing real face images and multiple AI-generated face images, use a stratified random sampling strategy to divide the training set, validation set, and test set, and build a core sample memory library; Step 2: Implement an adaptive dynamic enhancement strategy on the images of the training set, apply high-intensity enhancement in the early stage, gradually reduce the enhancement intensity as the training progresses, adopt a differentiated enhancement strategy for key samples in the memory library, and extract frequency domain features as auxiliary input channels; Step 3: Use the pre-trained deep convolutional network as the feature extraction backbone, coordinate the expansion ratio of network depth, width and resolution through a composite scaling strategy, build an adaptive classification head and embed a lightweight attention unit; Step 4: Design the teacher-student knowledge distillation framework, configure temperature parameters and soft label weights, set up feature distillation layers to capture intermediate layer feature representations and pass them in the incremental learning phase; Step 5: Execute the training process, use automatic mixed precision training technology, configure dynamic learning rate scheduling strategy and combine regularization technology to prevent overfitting; Step 6: Set up a performance monitoring module. When the model performance drops beyond a preset threshold or a new AI generation technology is detected, the incremental learning process is automatically triggered. In incremental learning, new data and memory samples are jointly trained, and the total loss consists of classification loss and knowledge distillation loss. Step 7: Calculate the sample importance score based on the gradient information, identify key samples and update the memory library; Step 8: Calculate the similarity of new and old task features, and dynamically adjust the shared layer depth and knowledge distillation weight; Step 9: After each incremental learning is completed, the model version is automatically iterated, and model pruning and quantization operations are performed regularly; Step 10: Evaluate the model test results and continuously optimize the model structure and training strategy based on the evaluation results.
2. The method for distinguishing the authenticity of AI-generated facial images according to claim 1, characterized in that: The compound scaling strategy is expressed as: in, represents the global scaling factor, , , They represent the expansion ratios of the number of channels, number of layers, and input resolution respectively.
3. The method for distinguishing the authenticity of AI-generated facial images according to claim 1, characterized in that: The adaptive dynamic enhancement strategy of step 2 includes: When epoch < 10, random cropping is performed with a scaling ratio of 0.2 to 0.8, ±30% color jitter, and 15° perspective transformation, expressed as: in, represents the original input image, represents a random cropping transformation, Indicates color dithering transformation, Represents perspective transformation; When 10≤epoch<30, perform ±15% color dithering and 10° rotation; When epoch ≥ 30, random horizontal flipping and ±5% brightness adjustment are performed.
4. The method for distinguishing authenticity of AI-generated facial images according to claim 1, characterized in that: Step 5 uses the AdamW optimizer and the cosine annealing learning rate scheduling strategy for training optimization, and the expression is: in, , , Indicates the restart cycle, Indicates the current iteration number.
5. The method for distinguishing authenticity of AI-generated facial images according to claim 1, characterized in that: During the training process of step 5, the input image is subjected to a fast Fourier transform to extract frequency domain features, which are then fused with the spatial domain features through a cross-modal attention mechanism: in, represents discrete wavelet transform, For channel splicing operations, represents the frequency domain feature map, Represents spatial domain characteristics.
6. The AI-generated facial image authenticity determination method according to claim 1, characterized in that: The shared layer depth in step 8 is expressed as: in, is the total depth of the network, is the similarity between the new and old tasks.
7. The AI-generated facial image authenticity determination method according to claim 1, characterized in that: The memory bank sample screening strategy of step 7 is: When the importance score satisfies When , it is an old sample with high importance, and the retention ratio is 40%; When the importance score satisfies , which is a medium-importance sample, the retention ratio is 30%; The sample retention rate for new AI generation technology is 30%.
8. The method for distinguishing authenticity of AI-generated facial images according to claim 1, characterized in that: The model classifier adopts a two-level fully connected layer structure and embeds a dynamic regularization mechanism. The first level contains a 512-dimensional fully connected layer with ReLU activation. The second level is reduced by a 128-dimensional fully connected layer and then outputs the probability through the Sigmoid function. The Dropout probability With training rounds Linear attenuation, the formula is: Among them, the Dropout probability The initial value of is set to 0.
5.
9. An AI-generated facial image authenticity determination device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is used to implement the AI-generated facial image authenticity determination method as described in any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for distinguishing the authenticity of an AI-generated facial image as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Generating neural networks
CN117744762A
Image forgery detection method based on self-supervised contrast learning
CN119693354A
Deep-neural-network-based class-incremental learning method for mobile phone radiation source spectrogram
WO2024119422A1
Cited By
New goods category discovery method based on asymmetric enhancement and confidence collaborative learning
CN120236146A
New Goods Category Discovery Method Based on Asymmetric Enhancement and Confidence Collaborative Learning
CN120236146B
Identification method and system for artificial intelligence generated image and related equipment
CN120298810A
Training method and system for optimizing machine readable area recognition model
CN120976670A
AI generated image detection method based on multi-agent collaboration
CN121033632A