AI generated image general detection method and system based on high and low level feature fusion, and computer equipment

By fusing high- and low-level features and using the DINOv2:ViT-L/14 model to extract semantic and noise features, the problems of cross-model generalization and low computational efficiency of AI-generated image detection methods were solved, achieving real-time and accurate image authenticity detection and improving information discrimination capabilities.

CN120635647AActive Publication Date: 2025-09-12INNER MONGOLIA UNIV OF TECH

Patent Information

Application Number
CN202510744178.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Existing AI-generated image detection methods are difficult to generalize across models, have low computational efficiency, and are unable to meet real-time detection needs.

Method used

A method based on high- and low-level feature fusion is adopted to extract the semantic features of the image through the DINOv2:ViT-L/14 model, and feature splicing is performed in combination with noise features. A linear multi-layer perceptron is used for detection to output the true and false image results.

Benefits of technology

It achieves the generalization of unknown generation models, simplifies the detection steps, can meet the real-time detection needs, improves the detection efficiency, and enhances the public's ability to discern forged information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635647A_ABST
    Figure CN120635647A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI image authenticity detection, and discloses an AI generated image general detection method and system based on high-level and low-level feature fusion, and computer equipment, and the method comprises the steps: extracting high-level features (semantic features) and low-level features (noise features) through an extraction strategy and a fusion mechanism of the high-level features (semantic features) and the low-level features (noise features) of an image; according to the method, a fusion type and learnable general feature is provided for a detection model, generalization of an unknown generation model can be effectively realized, an image in an unknown generation mode or a mixed generation mode can be effectively generalized, the detection steps are simplified and general, the real-time detection requirement can be met, the consumed time is short, and the detection efficiency can be improved; by detecting the authenticity of the to-be-detected image, the authenticity of the image widely spread in the social media can be effectively deduced, a general evidence obtaining scheme for coping with AIGC counterfeiting is provided, and the identification ability and vigilance of the public for counterfeiting information can be improved. Meanwhile, the method can also be applied to other fields such as large model training sample screening and multi-modal information detection, and the application prospect is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of AI image authenticity detection, and in particular to a general detection method, system and computer device for AI-generated images based on the fusion of high- and low-level features. Background Art

[0002] In recent years, with the rapid development of deep learning technology, generative models have made significant progress in generating realistic images. Commercial platforms for cultural image processing (CGI) can generate high-quality, multi-scene, and semantically consistent realistic images. Furthermore, the low barrier to entry for these CGI applications makes it practically anyone to create large quantities of fake images. Due to their visual comprehensibility and cognitive plausibility, generated images are particularly convincing in many scenarios, raising numerous security and privacy concerns.

[0003] The existing general detection methods for generated images can be roughly divided into four categories: (1) Data-driven detection: Construct a training set containing a large number of true and false samples, conduct large-scale training, and drive the model to learn the characteristic differences between true and false samples; (2) Detection based on feature statistical analysis: distinguishing generated images from real images by analyzing the statistical or physical feature differences in the images; (3) Detection based on feature reconstruction: The image to be tested is mapped back to the latent space of the generative model, and feature reconstruction is performed. The authenticity of the image is judged based on the feature differences before and after reconstruction. (4) Detection based on data distribution difference: By comparing the distribution differences between generated images and real images in the latent space or feature space, the detection model is trained to learn to distinguish between real and fake images.

[0004] However, these detection methods all rely on datasets trained on specific generative models (such as GANs or diffusion models), making them difficult to generalize effectively to images generated using unknown or mixed methods. Furthermore, some deep learning-based detection methods are computationally expensive and complex, making them difficult to meet real-time detection requirements and resulting in low computational efficiency. Summary of the Invention

[0005] This application provides a general detection method for AI-generated images based on the fusion of high- and low-level features, aiming to solve the technical problems in the existing detection methods that are difficult to generalize across models and have low computational efficiency.

[0006] This application provides a general detection method for AI-generated images based on the fusion of high- and low-level features, including: Acquire the image to be tested; Extracting noise features of the image to be measured; Extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; Performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature; The fusion features are input into the detection model, and the detection results are output, where the detection results include AI-generated images and real images.

[0007] Preferably, the step of extracting the noise characteristics of the image to be measured includes: Randomly extracting local blocks with the same resolution from the image to be tested; Calculate the texture richness measure value of each local patch in each channel; Filter out the local images corresponding to the texture richness measurement values ​​that meet the preset range as texture-poor blocks; Performing high-pass filtering on the texture-poor image block on each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; A noise token of the target image in the texture-poor block is extracted according to the high-frequency noise component, and the noise token is used as a noise feature of the target image.

[0008] Preferably, the step of calculating the texture richness measure value of each local block in each channel includes: Get the total number of pixel rows and columns of each local tile; Get the grayscale value of a single pixel in a single channel of each local block; Based on the grayscale value of a single pixel in a single channel, the total number of rows and columns of pixels in each local block, the sum of the absolute differences between the grayscale of each pixel in a single channel and the domain pixel of each local block is calculated and summed to obtain the texture richness measurement value.

[0009] Preferably, the step of extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model includes: Importing the DINOv2:ViT-L / 14 model and pre-trained parameters, loading the pre-trained parameters into the DINOv2:ViT-L / 14 model, and freezing the pre-trained parameters; Extracting high-level semantic features and multiple local semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; Calculating an average value of the quantity dimensions of the plurality of local semantics to obtain an average local semantic feature; The high-level semantic features and the average local semantic features are used as semantic features of the image to be tested.

[0010] Preferably, the step of performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature comprises: Acquiring the semantic features, wherein the semantic features include high-level semantic features and average local semantic features; Obtaining a noise token corresponding to the noise feature, and adjusting the shape of the noise token based on the shape of the high-level semantic feature so that the two have the same shape; In the channel of the noise token, the dimensions of the noise token, the high-level semantic feature and the average local semantic feature are spliced ​​into a tensor corresponding to a preset value to obtain a high- and low-level fusion feature; The high-level and low-level fusion features are linearly flattened to obtain the fusion features.

[0011] Preferably, the step of inputting the fusion feature into the detection model and outputting the detection result includes: Obtaining training data and adding random data augmentation to the training data, wherein the random data augmentation includes randomly flipping the training images and adding noise; The detection model is trained as a true-false two-category detection model based on a binary cross entropy loss function, wherein the detection model includes a linear multilayer perceptron composed of multiple fully connected layers, and the linear multilayers include ReLU activation function layers and Dropout layers; The fusion feature is input into the authenticity two-class detection model, and the detection result is output.

[0012] This application also provides a general detection system for AI-generated images based on high- and low-level feature fusion, including: A first acquisition module is used to acquire an image to be tested; A noise feature extraction module, used to extract the noise features of the image to be measured; A semantic feature extraction module, configured to extract semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; A splicing module, configured to perform feature splicing on the noise features and semantic features to obtain fusion features; The detection module is used to input the fusion features into the detection model and output the detection results, wherein the detection results include AI-generated images and real images.

[0013] Preferably, the noise feature extraction module includes: A local image block extraction unit, configured to randomly extract local image blocks with the same resolution from the image to be tested; A computing unit, configured to calculate a texture richness measure value for each local tile in each channel; a screening unit, configured to screen out local images corresponding to texture richness measurement values ​​that meet a preset range as texture-poor blocks; a high-pass filtering unit, configured to perform high-pass filtering on the texture-poor image block on each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; A noise token extraction unit is configured to extract noise tokens of the target image in the texture-poor block according to the high-frequency noise component, and use the noise tokens as noise features of the target image.

[0014] The present application also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0015] The present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0016] The beneficial effects of the present application are as follows: the present application provides a fusion-type, learnable universal feature for the detection model through the extraction strategy and fusion mechanism of high-level image features (semantic features) and low-level features (noise features), which can more effectively realize the generalization of unknown generation models and effectively generalize to images of unknown generation methods or mixed generation methods. The detection steps are simplified and more universal, which can meet the needs of real-time detection, takes less time, and can improve detection efficiency. By detecting the authenticity of the image to be tested, the authenticity of the image widely circulated in social media can be effectively inferred, providing a universal forensic solution to deal with AIGC forgery, which helps to improve the public's ability to distinguish and be vigilant about forged information. At the same time, the present invention can also be applied to other fields such as large model training sample screening and multimodal information detection, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a schematic diagram of a method flow chart according to an embodiment of the present application.

[0018] Figure 2 This is a schematic diagram of the internal structure of a computer device according to an embodiment of the present application.

[0019] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0021] like Figure 1 、 Figure 2As shown, this application provides a general detection method for AI-generated images based on high- and low-level feature fusion, including: S1. Acquire the image to be tested; S2, extracting noise features of the image to be measured; S3. Extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; S4, performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature; S5. Input the fusion features into a detection model and output a detection result, wherein the detection result includes an AI-generated image and a real image.

[0022] As described in the above steps S1-S5, the present application obtains the noise features and semantic features of the image to be tested and fuses them. This fusion feature is more universal and can generalize the detection model to unknown generative models. It is not limited to the detection of fake faces, and effectively realizes the universal detection of generated images. According to the prior art, the high-frequency components such as low-level noise patterns and texture details of the image show more obvious differences in the texture-poor areas of the real image and the generated image. Therefore, the noise features in the texture-poor areas are more worthy of attention. Based on this, the present application extracts the noise features of the image to be tested. The noise feature extraction process mainly targets the high-frequency components such as noise patterns and texture details in the image. The semantic feature extraction process uses the pre-trained DINOv2:ViT-L / 14 network to extract and output high-level semantic features in the image to be tested. DINOv2 is a large-scale self-supervised model based on the Vision Transformer (ViT) architecture of Meta AI. It can learn universal visual features from any image collection without fine-tuning. It is suitable for a variety of downstream visual tasks including image classification and is currently one of the top visual feature extraction models. Through the pre-trained DINOv2:ViT-L / 14 model, it is possible to extract high-level semantic features for forensics from generated images that are unknown in the training process. Since the DINOv2 model itself has been fully trained on a large-scale integrated image data pool through advanced training engineering, it has the ability to learn, abstract and represent features from shallow to deep layers. Therefore, the pre-trained parameters of DINOv2 have the ability to embed samples of any scene image. Theoretically, DINOv2 has a variety of capabilities required to meet the general detection of AI-generated images. This application abandons the previous method of using only ProGAN as a training sample, and chooses a training method that combines ProGAN and ADM dual-generated samples, so that the detection model can better take into account the high-level semantic features and low-level noise features of the image to be tested, breaking the feature learning limitations brought about by the single-generated sample training mode.

[0023] The feature alignment and splicing process is used to align and splice the low-level noise features and high-level semantic features extracted from the original image, so that they are organically combined to form a fusion feature that can be directly accepted by the linear classifier.

[0024] The detection model is a linear multi-layer perceptron (MLP) consisting of multiple fully connected layers. Relu activation layers and Dropout (random dropout) layers are included between the linear layers for nonlinear feature mapping and to prevent model overfitting, respectively. The detection model receives the flattened fusion features and trains a binary classification model using BCELoss (binary cross entropy loss). The model ultimately outputs a probability value between 0 and 1 as the detection result. In the training set, real images are labeled "0" and generated images are labeled "1." If the probability value output by the detection model is greater than or equal to 0.5, the image is classified as a generated image; otherwise, it is classified as a real image. This application provides a fusion-type, learnable universal feature for the detection model through the extraction strategy and fusion mechanism of high-level image features (semantic features) and low-level features (noise features). It can more effectively generalize to unknown generation models and effectively generalize to images with unknown generation methods or mixed generation methods. The detection steps are simplified and more universal, which can meet the needs of real-time detection, takes less time, and can improve detection efficiency. By testing the authenticity of the image to be tested, the authenticity of the image widely circulated on social media can be effectively inferred, providing a universal forensic solution to counter AIGC (Artificial Intelligence Generated Content) forgery, which helps to improve the public's ability to discern and be vigilant against forged information. At the same time, the present invention can also be applied to other fields such as large model training sample screening and multimodal information detection, and has broad application prospects.

[0025] In one embodiment, the step S2 of extracting the noise characteristics of the image to be measured includes: S21, randomly extracting local blocks with the same resolution from the image to be tested; S22, calculating the texture richness measurement value of each local block in each channel; S23, screening out the local images corresponding to the texture richness measurement values ​​that meet the preset range as texture-poor blocks; S24, performing high-pass filtering on the texture-poor image block in each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; S25 , extracting noise tokens of the target image in the texture-poor block according to the high-frequency noise component, and using the noise tokens as noise features of the target image.

[0026] As described in steps S21-S25 above, high-frequency components such as low-level noise patterns and texture details in images exhibit more pronounced differences in texture-poor regions between real and generated images. Therefore, noise features in texture-poor regions are of particular interest. This embodiment first randomly extracts local patches of the same resolution (e.g., 32×32) from the image to be tested. Within each channel of the local patches, a texture richness measure for each channel is calculated, and a certain number of local patches with relatively low channel texture richness measures are retained. The retained texture-poor patches are then high-pass filtered using a Gaussian high-pass filter on each channel, extracting the high-frequency noise components within each texture-poor patch as noise features for these texture-poor patches. Finally, the noise features of these texture-poor patches are averaged across the corresponding channels to serve as noise tokens for the texture-poor regions of the original image. The resulting low-level noise features are equal to the resolution of each local patch and maintain channel independence (i.e., the low-level noise features are feature vectors of the form [3, 32, 32]).

[0027] In one embodiment, the step S22 of calculating the texture richness measure value of each local tile in each channel includes: S221, obtaining the total number of pixel rows or columns of each local image block; S222, obtaining the grayscale value of a single pixel in a single channel of each local image block; S223. Based on the grayscale value of a single pixel in a single channel and the total number of rows or columns of pixels in each local block, calculate the sum of the absolute differences between the grayscale of each pixel in each local block in a single channel and the grayscale of the domain pixel, and sum them up to obtain a texture richness measurement value.

[0028] As described in steps S221-S223 above, the specific formula for calculating the texture richness measurement value is: ; Wherein, the I div represents the texture richness measure of a single local tile within a single channel, M represents the total number of rows or columns of pixels in the local tile (the tile length and width are the same), x represents the grayscale value of a single pixel in the local tile within a single channel, and i and j represent the current row and column values ​​of the pixel in the local tile, respectively. By calculating this texture richness measure, the noise feature extraction process can focus more on texture-poor areas in the image, providing a realistic and accurate basis for subsequent image detection results. A simple Gaussian high-pass filter effectively filters out irrelevant low-frequency information in texture-poor areas in the frequency domain, effectively preserving noise features.

[0029] In one embodiment, step S3 of extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model includes: S31. Import the DINOv2:ViT-L / 14 model and pre-training parameters, load the pre-training parameters into the DINOv2:ViT-L / 14 model, and freeze the pre-training parameters; S32, extracting high-level semantic features and multiple local semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; S33, calculating an average value of the quantity dimensions of the plurality of local semantics to obtain an average local semantic feature; S34: Using the high-level semantic features, the multiple local semantic features, and the average local semantic features as semantic features of the image to be tested.

[0030] As described in steps S31-S34 above, the DINOv2:ViT-L / 14 model is imported and its official pre-trained parameters are loaded into the model. The model parameters are frozen (no fine-tuning is performed) and used as the semantic feature extraction module. Because the ViT-L / 14 network architecture is used, the hidden layer dimension in DINOv2:ViT-L / 14 is 1024, and the size of each semantic feature is strictly limited to 14×14. After receiving a test image whose resolution has been pre-processed to a maximum multiple of 14 (smaller than the original image resolution), the DINOv2:ViT-L / 14 model primarily outputs two types of features: classification tokens (high-level semantic features) and local block tokens (local semantic features). Classification tokens contain the global top-level semantics of the image and have a shape of [1, 1024]. Local block tokens contain all the local semantic information corresponding to N local patches and have a shape of [N, 1024]. The number N is calculated by dividing the resolution of the test image by the resolution of the local patch. To prevent the dimensionality of the output features from being too high and to balance the feature representations of all local block tokens, the output local block tokens are averaged in the number dimension to obtain an average local block token with a shape of [1, 1024], i.e., the average local semantic feature. This application extracts semantic features based on the DINOv2:ViT-L / 14 model, which can capture subtle differences in features between generated and real images and extract more diverse local semantic features. This allows the detection model to fully focus on local grayscale fluctuation anomalies and global consistency defects. The features output by the ViT-L / 14 network structure are typically high-dimensional. The detection model (MLP, Multilayer Perceptron) can map them to a space more suitable for classification through nonlinear transformations, thereby improving separation.

[0031] In one embodiment, the step S4 of performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature includes: S41, acquiring the semantic features, wherein the semantic features include high-level semantic features and average local semantic features; S42, obtaining a noise token corresponding to the noise feature, and adjusting the shape of the noise token based on the shape of the high-level semantic feature so that the shapes of the noise token and the noise token are consistent; S43, in the channel of the noise token, concatenating the dimensions of the noise token, the high-level semantic feature, and the average local semantic feature into a tensor corresponding to a preset value to obtain a high- and low-level fused feature; S44. Perform linear tensor flattening on the high- and low-level fusion features to obtain fusion features.

[0032] As described in steps S41-S44 above, low-level noise features and high-level semantic features are output through the noise feature extraction module and the semantic feature extraction module. All features are then aligned and spliced, and adjusted to a shape that can be directly accepted by the linear layer. Specifically, the shape of the noise token output by the noise feature extraction module is first adjusted from [3, 32, 32] to [3,1024] so that its pixel dimension is consistent with the hidden layer dimension of the high-level semantic feature. Secondly, the noise token, classification token and average local block token are spliced ​​into a tensor of shape [5, 1024] in the channel dimension of the noise token to form a high- and low-level fusion feature of the image to be tested. Finally, the fused features are flattened into a linear tensor of the shape of

[5120] to facilitate acceptance by the detection model.

[0033] In one embodiment, the step S5 of inputting the fusion features into the detection model and outputting the detection results includes: S51, obtaining training data, and adding random data enhancement to the training data, wherein the random data enhancement includes randomly flipping the training image and adding noise; S52. Training the detection model as a true-false two-category detection model based on a binary cross entropy loss function, wherein the detection model includes a linear multilayer perceptron composed of multiple fully connected layers, and the linear multilayers include ReLU activation function layers and Dropout layers; S53: Input the fusion feature into the authenticity two-class detection model and output the detection result.

[0034] As described in steps S51-S53 above, the detection model is a linear multilayer perceptron (MLP) consisting of multiple fully connected layers. The linear layers include Relu activation function layers and Dropout layers, which are used for nonlinear feature mapping and to prevent model overfitting, respectively. The detection model receives the flattened fused features and trains a binary true / false classification detection model using BCELoss. Furthermore, to reduce interference from post-processed images and improve the generalization of the detection model, random data augmentation is added before training begins. The training data undergoes random preprocessing such as image flipping and noise addition. During testing, the detection model receives a flattened fused feature tensor and, through inter-layer mapping, ultimately outputs a probability value between 0 and 1 as the detection result. In the dataset, real images are labeled "0" and generated images are labeled "1." If the probability value output by the detection model is greater than or equal to 0.5, the image under test is classified as a generated image; otherwise, it is classified as a real image.

[0035] This application is based on the strategy of dual-branch feature splicing, splitting the low-level noise features and DINOv2 for extracting high-level features into two parallel branches that do not interfere with each other. The low-level noise features and high-level semantic features come from two different processing methods of the same image, which are processed separately and finally combined into a fusion feature through feature splicing. The feature engineering in this scheme is first executed in parallel, and then sequential execution begins after feature splicing. Experiments have shown that this scheme can effectively utilize low-level noise features and high-level features, thereby making the generated detection results more accurate.

[0036] This application also provides a general detection system for AI-generated images based on high- and low-level feature fusion, including: A first acquisition module is used to acquire an image to be tested; A noise feature extraction module, used to extract the noise features of the image to be measured; A semantic feature extraction module, configured to extract semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; A splicing module, configured to perform feature splicing on the noise features and semantic features to obtain fusion features; The detection module is used to input the fusion features into the detection model and output the detection results, wherein the detection results include AI-generated images and real images.

[0037] In one embodiment, the noise feature extraction module includes: A local image block extraction unit, configured to randomly extract local image blocks with the same resolution from the image to be tested; A computing unit, configured to calculate a texture richness measure value for each local tile in each channel; a screening unit, configured to screen out local images corresponding to texture richness measurement values ​​that meet a preset range as texture-poor blocks; a high-pass filtering unit, configured to perform high-pass filtering on the texture-poor image block on each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; A noise token extraction unit is configured to extract noise tokens of the target image in the texture-poor block according to the high-frequency noise component, and use the noise tokens as noise features of the target image.

[0038] It should be noted that each module and unit in the general detection system for AI-generated images based on fusion of high- and low-level features corresponds one-to-one to the steps in the general detection method for AI-generated images based on fusion of high- and low-level features.

[0039] like Figure 2 As shown, the present application also provides a computer device, which can be a server, and its internal structure can be as shown in FIG. Figure 2 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store all data required for the process of the general detection method for AI-generated images based on the fusion of high and low-level features. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the general detection method for AI-generated images based on the fusion of high and low-level features is implemented.

[0040] Those skilled in the art will understand that Figure 2 The structure shown in is merely a block diagram of a portion of the structure related to the present application solution and does not constitute a limitation on the computer device to which the present application solution is applied.

[0041] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it implements any of the above-mentioned general detection methods for AI-generated images based on the fusion of high- and low-level features.

[0042] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM).

[0043] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0044] The above description is only a preferred embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A general detection method for AI-generated images based on high- and low-level feature fusion, characterized in that: include: Acquire the image to be tested; Extracting noise features of the image to be measured; Extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; Performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature; The fusion features are input into the detection model, and the detection results are output, where the detection results include AI-generated images and real images.

2. The general detection method for AI-generated images based on high- and low-level feature fusion according to claim 1 is characterized in that: The step of extracting the noise characteristics of the image to be measured includes: Randomly extracting local blocks with the same resolution from the image to be tested; Calculate the texture richness measure value of each local patch in each channel; Filter out the local images corresponding to the texture richness measurement values ​​that meet the preset range as texture-poor blocks; Performing high-pass filtering on the texture-poor image block on each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; A noise token of the target image in the texture-poor block is extracted according to the high-frequency noise component, and the noise token is used as a noise feature of the target image.

3. The general detection method for AI-generated images based on high- and low-level feature fusion according to claim 1 is characterized in that: The step of calculating the texture richness measurement value of each local tile in each channel includes: Get the total number of pixel rows and columns of each local tile; Get the grayscale value of a single pixel in a single channel of each local block; Based on the grayscale value of a single pixel in a single channel, the total number of rows and columns of pixels in each local block, the sum of the absolute differences between the grayscale of each pixel in a single channel and the domain pixel of each local block is calculated and summed to obtain the texture richness measurement value.

4. The general detection method for AI-generated images based on high- and low-level feature fusion according to claim 1 is characterized in that: The step of extracting semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model includes: Importing the DINOv2:ViT-L / 14 model and pre-trained parameters, loading the pre-trained parameters into the DINOv2:ViT-L / 14 model, and freezing the pre-trained parameters; Extracting high-level semantic features and multiple local semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; Calculating an average value of the quantity dimensions of the plurality of local semantics to obtain an average local semantic feature; The high-level semantic features and the average local semantic features are used as semantic features of the image to be tested.

5. The general detection method for AI-generated images based on high- and low-level feature fusion according to claim 1 is characterized in that: The step of performing feature splicing on the noise feature and the semantic feature to obtain a fusion feature includes: Acquiring the semantic features, wherein the semantic features include high-level semantic features and average local semantic features; Obtaining a noise token corresponding to the noise feature, and adjusting the shape of the noise token based on the shape of the high-level semantic feature so that the two have the same shape; In the channel of the noise token, the dimensions of the noise token, the high-level semantic feature and the average local semantic feature are spliced ​​into a tensor corresponding to a preset value to obtain a high- and low-level fusion feature; The high-level and low-level fusion features are linearly flattened to obtain the fusion features.

6. The general detection method for AI-generated images based on high- and low-level feature fusion according to claim 1 is characterized in that: The step of inputting the fusion feature into the detection model and outputting the detection result includes: Obtaining training data and adding random data augmentation to the training data, wherein the random data augmentation includes randomly flipping the training images and adding noise; The detection model is trained as a true-false two-category detection model based on a binary cross entropy loss function, wherein the detection model includes a linear multilayer perceptron composed of multiple fully connected layers, and the linear multilayers include ReLU activation function layers and Dropout layers; The fusion feature is input into the authenticity two-class detection model, and the detection result is output.

7. A general detection system for AI-generated images based on high- and low-level feature fusion, characterized by: include: A first acquisition module is used to acquire an image to be tested; A noise feature extraction module, used to extract the noise features of the image to be measured; A semantic feature extraction module, configured to extract semantic features of the image to be tested based on the DINOv2:ViT-L / 14 model; A splicing module, configured to perform feature splicing on the noise features and semantic features to obtain fusion features; The detection module is used to input the fusion features into the detection model and output the detection results, wherein the detection results include AI-generated images and real images.

8. The AI-generated image universal detection system based on high- and low-level feature fusion according to claim 7 is characterized in that: The noise feature extraction module includes: A local image block extraction unit, configured to randomly extract local image blocks with the same resolution from the image to be tested; A computing unit, configured to calculate a texture richness measure value for each local tile in each channel; a screening unit, configured to screen out local images corresponding to texture richness measurement values ​​that meet a preset range as texture-poor blocks; a high-pass filtering unit, configured to perform high-pass filtering on the texture-poor image block on each channel based on a Gaussian high-pass filter to obtain a high-frequency noise component; A noise token extraction unit is configured to extract noise tokens of the target image in the texture-poor block according to the high-frequency noise component, and use the noise tokens as noise features of the target image.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Deep counterfeit image detection method and system fused with noise perception

    CN114677372A

  • Generalized deeply-forged image detection method and system based on noise perception

    CN118196865A

  • Freezing ViT feature fusion network-based AI generated image detection method

    CN118570599A

  • Document image tampering positioning method

    CN118587411A

  • Computer Vision Systems and Methods for Blind Localization of Image Forgery

    US20210004648A1

Cited By

  • Low-light image degradation simulation method and system based on conditional shift diffusion trajectory

    CN122156000A