Synthetic image detection method and system based on digital imaging alignment traces

CN122841933APending Publication Date: 2026-09-29UESTC (SHENZHEN) ADVANCED RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610728610.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供基于数字成像对齐痕迹的合成图像检测方法及系统,以解决上述背景技术中提到的现有的误判率高、缺少基于物理来源的检测依据、现有零样本检测方法泛化能力不足等问题

Benefits of technology

[0015]由上述技术方案可知,本发明与现有技术相比至少具备以下优点和积极效果:本发明将检测任务转换为“图像是否具有真实RAW-RGB物理成像一致性”的判断问题,通过提取反映物理采集过程的对齐痕迹进行真实性鉴别;能有效克服现有方法在跨物理通道场景下的误判问题。在训练阶段仅使用真实RAW图像及其经过模拟成像流程转换的RGB图像,不需要任何人工智能生成图像作为训练样本;通过构建RAW-RGB共享特征空间并约束对齐痕迹的内容不变性和流程可分性,使训练后的模型能够零样本地检测来自任意生成模型合成的图像。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841933A_ABST
    Figure CN122841933A_ABST
Patent Text Reader

Abstract

The application discloses a synthetic image detection method and system based on trace alignment of digital imaging, comprising: constructing a training data set by using a real RAW image and a corresponding RGB image generated through a plurality of predefined RAW-RGB image processing procedures; constructing a multi-branch network comprising an RGB visual branch, a RAW topological branch and a RAW visual branch, and obtaining an alignment trace extraction model through joint optimization training; inputting a to-be-detected RGB image into the model, extracting an alignment trace reflecting consistency of the RAW-RGB image processing procedure, and judging whether the image is an artificial intelligence generated image based on the alignment trace. The application converts the detection problem into a judgment of whether the image has a real physical imaging source, does not need to rely on prior samples of a specific generated model, supports zero-sample clustering detection, zero-sample similarity detection and few-sample linear classification detection, and has good generalization ability for unknown generated models and cross-domain scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a method and system for detecting synthetic images based on digital imaging alignment traces. Background Technology

[0002] With the rapid development of generative artificial intelligence (AI) technology, images generated based on generative adversarial networks (GANs), diffusion models, and large-scale text-based image models are continuously improving in terms of visual quality, semantic consistency, and realism, leading to a gradual reduction in the appearance difference between AI-generated images and real-world images. How to effectively distinguish between real-world images and AI-generated images has become a crucial issue in digital image forensics and content security. Existing image forensics techniques typically input the image to be detected into a detection model, learning texture artifacts, frequency domain anomalies, upsampling traces, compression residuals, or features specific to the generative model in the RGB image to classify real and synthetic images. Some methods attempt to employ zero-shot or few-shot detection strategies to reduce reliance on training samples for specific generative models and improve the generalization detection capability against unknown generative models.

[0003] Existing detection methods primarily rely on statistical differences in the RGB space, making it difficult to characterize the physical origin of an image. These methods rely mainly on surface statistical features of digital images, rather than the imaging process of a real image from the physical world to the digital space. When synthetic images undergo physical remapping processes such as printing, scanning, or screen capture, they re-enter the physical-to-digital acquisition process, leading to misclassification by existing detectors. Current detection tasks typically classify images directly as real or synthetic, but the classification of AI-generated images after physical remapping is ambiguous. Existing technologies lack detection criteria based on physical origin. Existing zero-shot detection methods have insufficient generalization ability for unknown generation models and cross-domain scenarios; detection performance tends to degrade when the image to be detected comes from an unknown generation model, different generation paradigms, or different post-processing conditions. Existing methods lack alignment trace modeling mechanisms for the RAW-RGB imaging process; the formation of a real image involves the correlation between the RAW signal, the camera's internal processing flow, and the final RGB output. Existing detection technologies typically only focus on the final RGB image, failing to fully utilize the parameter correlations and structural constraints present in the RAW-RGB conversion process. Summary of the Invention

[0004] The purpose of this invention is to provide a synthetic image detection method and system based on digital imaging alignment traces, in order to solve the problems mentioned in the background art, such as high false positive rate, lack of detection basis based on physical source, and insufficient generalization ability of existing zero-shot detection methods.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: According to one aspect of the present invention, a method for detecting synthetic images based on digital imaging alignment traces is provided, the method comprising: A training dataset was constructed using real RAW images and corresponding RGB images generated through various predefined RAW-RGB image processing workflows. A multi-branch network containing an RGB visual branch, a RAW topology branch, and a RAW visual branch is constructed. An alignment trace extraction model is obtained by joint optimization training using the training dataset. The alignment trace extraction model is used to extract alignment traces from the input RGB image that reflect the consistency of the RAW-RGB image processing flow. The RGB image to be detected is input into the alignment trace extraction model to extract alignment traces, and the image is determined to be an artificial intelligence-generated image based on the alignment traces.

[0006] Based on the aforementioned scheme, the RAW-RGB image processing flow includes at least a demosaic operation, and further includes one or more of white balance, color correction, and tone mapping; wherein, white balance and tone mapping each appear at most once in the same RAW-RGB image processing flow; the RAW-RGB image processing flow also includes optional image post-processing operations, which include one or more of compression, scaling, blurring, or noise addition.

[0007] Based on the aforementioned scheme, the RGB visual branch includes: inputting the RGB image into a pre-trained visual feature extraction network to obtain a spatial feature representation; inputting the spatial feature representation into a hierarchical entropy attention module, and using learnable initial latent features as query features to output alignment traces; The hierarchical entropy attention module is used to calculate the physical entropy of each local region in the spatial feature representation, divide each local region into multiple entropy intervals according to the physical entropy, sample local features hierarchically from each entropy interval, use the query feature as the query vector, use the local feature as the key value vector, and perform weighted fusion through the attention mechanism to generate the alignment trace.

[0008] Based on the aforementioned scheme, the RAW topology branch includes: representing each RAW-RGB image processing flow as a directed weighted topology graph, where nodes represent processing operations, directed edges represent operation order, and edge weights represent operation parameters; inputting the directed weighted topology graph into a graph neural network to extract graph-level features; and aligning the graph-level features with the alignment traces generated by the RGB visual branch in a shared embedding space using a loss function.

[0009] Based on the aforementioned scheme, for the same image processing workflow, the loss function minimizes the distance between the graph-level features of the workflow and the alignment traces of the workflow; for different image processing workflows, the loss function maximizes the distance between the graph-level features of the workflow and the alignment traces of other workflows.

[0010] Based on the aforementioned scheme, the RAW visual branch includes: encoding the RAW image and the corresponding RGB image using a pre-trained autoencoder or variational autoencoder respectively to obtain the RAW latent distribution and the RGB latent distribution; using the alignment traces generated by the RGB visual branch as conditional information, modulating the mean of the RAW latent distribution through an attention mechanism; and using KL divergence constraints to ensure consistency between the modulated RAW latent distribution and the RGB latent distribution.

[0011] Based on the aforementioned scheme, the joint optimization is achieved by minimizing the total loss function, which is a weighted sum of the loss of the RAW topology branch and the loss of the RAW vision branch; wherein the loss of the RAW topology branch is used to constrain the consistency between the alignment traces and the topology of the image processing flow, and the loss of the RAW vision branch is used to constrain the consistency of the distribution of RAW images and RGB images in the latent space.

[0012] Based on the aforementioned scheme, the step of determining whether an image is an AI-generated image based on the alignment traces employs a zero-shot detection method, including: Multiple alignment marks are extracted from multiple images to be detected. An unsupervised clustering algorithm is used to cluster the multiple alignment marks. Based on the clustering results, real images and artificial intelligence generated images are distinguished. Alternatively, the similarity or distance between the alignment marks of the image to be detected and the pre-stored alignment marks of the real image reference can be calculated, and the determination can be made based on the relationship between the similarity or distance and a preset threshold.

[0013] Based on the aforementioned scheme, the method of determining whether an image is an AI-generated image based on alignment traces, using a few-sample detection approach, includes: freezing the parameters of a trained alignment trace extraction model; adding a linear classifier to the output of the alignment trace extraction model; training the linear classifier using a small number of samples labeled with AI-generated images or real images; inputting the image to be detected into the alignment trace extraction model to obtain alignment traces, and then inputting the alignment traces into the trained linear classifier to obtain a classification result.

[0014] According to another aspect of the present invention, a synthetic image detection system based on digital imaging alignment traces is provided, the system comprising: The data generation module is used to acquire real RAW images and convert them into corresponding RGB images through various predefined RAW-RGB image processing procedures to build a training dataset. The feature extraction model, which includes an RGB visual branch, a RAW topology branch, and a RAW visual branch, is obtained through joint optimization training. It is used to extract alignment traces from the input RGB image that reflect the consistency of the RAW-RGB image processing flow. The detection module is used to acquire alignment traces extracted by the feature extraction model from the RGB image to be detected, and to determine whether the image is an artificial intelligence generated image based on the alignment traces.

[0015] As can be seen from the above technical solution, the present invention has at least the following advantages and positive effects compared with the prior art: The present invention transforms the detection task into a judgment problem of "whether the image has true RAW-RGB physical imaging consistency", and performs authenticity identification by extracting alignment traces that reflect the physical acquisition process; it can effectively overcome the misjudgment problem of existing methods in cross-physical channel scenarios. In the training phase, only real RAW images and their RGB images converted through simulated imaging processes are used, without requiring any artificial intelligence-generated images as training samples; by constructing a RAW-RGB shared feature space and constraining the content invariance and process separability of alignment traces, the trained model can detect images synthesized from arbitrary generative models with zero samples.

[0016] The alignment traces designed in this invention target the imaging process. Their content invariance ensures that images with different content under the same processing flow produce similar features, while their process separability ensures that images with the same content under different processing flows produce distinguishable features. The hierarchical entropy attention module further extracts features from texture-rich and uniform regions in a balanced manner through physical entropy-driven methods, avoiding the detection results being dominated by specific semantic regions. This invention maintains stable detection performance on images in different scenes and with different object categories. The RAW-RGB image processing flow is represented as a directed weighted topological graph. Structural features of the flow are extracted through a graph neural network and aligned with RGB side alignment traces using contrastive loss, explicitly modeling the RAW-RGB imaging flow and enhancing the robustness and interpretability of the features. This invention provides a unified feature extraction model that can flexibly adapt to different detection data conditions and is compatible with both zero-sample and few-sample application scenarios.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention. It is obvious that the drawings described below are merely some embodiments of the invention, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram of the synthetic image detection method based on digital imaging alignment traces according to the present invention; Figure 2 This is a flowchart illustrating the alignment trace modeling steps of the present invention. Figure 3 This is a flowchart illustrating the detection application steps of the present invention. Detailed Implementation

[0019] To more clearly illustrate the purpose, technical solutions, and advantages of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein. On the contrary, these embodiments are provided so that the present invention will be more comprehensive and complete, and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0020] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a full understanding of embodiments of the invention. However, those skilled in the art will recognize that the technical solutions of the invention can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of the invention.

[0021] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0022] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0023] The present invention will now be described in detail with reference to specific embodiments.

[0024] Example 1

[0025] like Figure 1 As shown, this embodiment provides a synthetic image detection method based on digital imaging alignment traces. The specific steps of this method are as follows: S1: Construct a training dataset using real RAW images and corresponding RGB images generated through various predefined RAW-RGB image processing procedures.

[0026] During the training phase, the first step is data generation to construct the training dataset needed to train the alignment trace extraction model. Raw RAW images are acquired from at least one physical image acquisition device (e.g., a digital camera, smartphone, industrial camera, or other imaging device with RAW output capability). RAW images are raw pixel data directly from the image sensor, unprocessed by an image signal processor (ISP), typically stored in Bayer pattern or other color filter arrays, preserving complete physical lighting information, linear photometric response characteristics, and sensor noise characteristics. The acquired RAW images constitute the original image set R = {Ri}, where Ri represents the i-th RAW image sample. To ensure the generalization ability of the trained feature extraction model, the RAW image set should cover various shooting scenes, lighting conditions, camera models, and sensor parameter configurations.

[0027] Multiple image processing flows for converting RAW images to RGB images are predefined, forming a flow set P={P j}, P j This represents the j-th RAW-RGB image processing flow; each processing flow P jThis simulates the complete or partial imaging chain from RAW signal to final RGB image in a real camera, including but not limited to the camera's internal ISP core operations and optional image post-processing operations. The ISP core operations include at least demosaicing, and may further include one or more of white balance, color correction, and tone mapping. Specifically, demosaicing is used to interpolate a single-channel RAW array into a three-channel RGB image, employing bilinear interpolation, edge-adaptive interpolation, or deep learning-based methods; white balance is used to eliminate the influence of light source color temperature, employing gray-world method, perfect reflection method, or methods based on camera preset parameters; color correction is used to convert the image from the sensor color space to a standard color space (such as sRGB, Adobe RGB), typically achieved through a 3×3 or higher-dimensional color correction matrix; tone mapping is used to compress a high dynamic range linear signal to the limited dynamic range that the display device can render, employing S-curves, logarithmic transformation, or perceptual methods. Image post-processing operations include, but are not limited to: image compression (such as JPEG compression, with adjustable quality factor), image scaling (such as upsampling or downsampling, with interpolation methods such as nearest neighbor, bilinear, and bicubic), image blurring (such as Gaussian blur and mean blur), noise addition (such as Gaussian noise and Poisson noise), sharpening, and contrast adjustment. Each processing step P... j The operation sequence, algorithm type, and parameter values ​​can all be configured according to actual needs to simulate image changes that may occur under different camera models, shooting settings, and transmission or storage conditions.

[0028] In a preferred implementation, each image processing flow P j It must include at least a depigmentation operation; white balance and tone mapping operations must each appear at most once in a single pass; post-processing operations must contain a maximum of three steps, and these steps must be arranged in a logical image processing pipeline order (e.g., depigmentation first, then white balance, then color correction, then tone mapping, and finally post-processing such as compression or scaling). This setup simulates a typical imaging chain of a real camera while avoiding over-complexity that would increase the difficulty of training.

[0029] For each RAW image in the RAW image set R... i It is then sequentially input into each image processing process P in the process set P. j Execute all operations defined in the process flow to obtain the corresponding RGB output image. Each pair ( ) constitute a training sample pair, where The ground truth as a physical source. The RGB observation images are used as the learning target; all training samples constitute the training image set. In this process, the same RAW image undergoes different processing steps to generate multiple RGB images with significantly different appearances, while different RAW images undergo the same processing step to generate RGB images with different content but similar imaging process characteristics. The training dataset constructed in this way enables subsequent feature extraction models to learn alignment traces that are related to the RAW-RGB conversion process but independent of the specific image content.

[0030] To enhance the diversity of training data and improve the model's robustness to common degradations in real-world scenes, random data augmentation operations can be performed after generating RGB images. These operations include, but are not limited to, random cropping, horizontal flipping, rotation, color jittering, and adding slight sensor noise. The generated RGB images and their corresponding RAW images are stored as image pairs for subsequent multi-branch joint training.

[0031] Through a data generation step, a high-quality dataset can be constructed to train an alignment trace extraction model using only physically acquired RAW images and their RGB images obtained through simulated imaging processes, without relying on artificial intelligence to generate image samples. This dataset contains the complete transformation correlation of real images from the physical world to digital space, laying a data foundation for subsequent determination of whether images to be detected share the same physical origin.

[0032] S2: Construct a multi-branch network containing an RGB visual branch, a RAW topology branch, and a RAW visual branch. Use the training dataset to train an alignment trace extraction model through joint optimization. The alignment trace extraction model is used to extract alignment traces from the input RGB image that reflect the consistency of the RAW-RGB image processing flow. After completing the data generation step and obtaining a training dataset containing RAW images and corresponding RGB image pairs, the alignment trace modeling step is performed; a feature extraction model is trained to extract RAW-RGB alignment traces from any input RGB image, characterizing whether the image originates from a real physical imaging process. For example... Figure 2 As shown, the multi-branch network includes an RGB visual branch, a RAW topology branch, and a RAW visual branch. The RGB visual branch is used to extract visual features from the RGB image and generate alignment traces. The RAW topology branch is used to structure the imaging processing flow and align it with the alignment traces. The RAW visual branch is used to constrain the potential distribution consistency between the RAW and RGB images.

[0033] The RGB vision branch includes: inputting RGB images into a pre-trained visual feature extraction network to obtain spatial feature representations; inputting spatial feature representations into a hierarchical entropy attention module, and using learnable initial latent features as query features to output alignment traces; the hierarchical entropy attention module is used to calculate the physical entropy of each local region in the spatial feature representation, divide each local region into multiple entropy intervals according to the physical entropy, hierarchically sample local features from each entropy interval, use the query features as query vectors, use local features as key-value vectors, and perform weighted fusion through an attention mechanism to generate alignment traces.

[0034] Specifically, the RGB images in the training dataset The input is fed into a pre-trained visual feature extraction network; this network can employ a convolutional neural network (e.g., ResNet, EfficientNet) or a visual Transformer (e.g., ViT, DINO) architecture, whose parameters are frozen during training to preserve the general visual representation capabilities learned on large-scale natural images. For the input RGB image, the visual feature extraction network outputs a multi-scale spatial feature map. Where K represents the number of local spatial regions, h k This represents the visual feature vector of the k-th local region. These spatial features contain information about low-level visual attributes of the image, such as texture, edges, and color distribution, but also contain information related to higher-level content, such as object category and scene semantics. To extract clean features relevant to the subsequent imaging process, these raw visual features need to be selectively filtered.

[0035] After the RGB visual feature extraction network, a hierarchical entropy attention module is set up to filter out feature components that are related to the physical imaging process but weakly related to the semantic content of the image from the spatial feature map. The initial potential features that are randomly initialized and updated during training are used as query features to perform attention fusion on the hierarchically sampled spatial features to generate RGB side alignment traces.

[0036] Physical entropy is calculated for each local region of the image. Physical entropy is obtained by weighted fusion of two parts: the first part is gradient change information, measured by the entropy value of the histogram of pixel gradient magnitude distribution within the region; the second part is luminance-chrominance distribution information, obtained by converting the image patch to a luminance-chrominance separated color space (such as Lab or YCbCr), statistically analyzing the probability distributions of the luminance and chrominance components, calculating their respective Shannon entropies, and then combining them. A higher physical entropy value indicates richer underlying imaging details (such as edges, texture transitions, and illumination changes), and these regions typically contain more parameter traces related to the RAW-RGB imaging process. Uniform regions with lower physical entropy (such as the sky or white walls) primarily reflect semantic content rather than imaging process information.

[0037] Based on the calculated physical entropy values ​​of each local region, all spatial features are divided into multiple entropy intervals. For example, the features are divided into L equal-frequency intervals or equal-width intervals according to the entropy values ​​from low to high, with each interval corresponding to a different physical entropy level.

[0038] For hierarchical sampling, a predetermined number of local feature vectors are sampled from each entropy interval to ensure that the model can pay balanced attention to image regions with different entropy levels during training, and to avoid the detection features being dominated by a single texture-rich region or texture-poor region.

[0039] The sampled local feature vector is input into a multi-head attention layer or a cross-attention layer. In this attention layer, the initial latent features F are... 0j (Can be set as a zero vector or a learnable embedding vector) as the query, using the sampled local features as the key and value, calculate the attention weights and perform a weighted summation of the local features, and output the aggregated feature vector, i.e., the RGB side alignment trace. This attention mechanism enables the model to adaptively assign higher weights to local regions that are more relevant to the physical imaging process, thereby enhancing the ability of alignment traces to represent the imaging process.

[0040] The RAW topology branch includes: representing each RAW-RGB image processing flow as a directed weighted topology graph, where nodes represent processing operations, directed edges represent the operation order, and edge weights represent operation parameters; inputting the directed weighted topology graph into a graph neural network to extract graph-level features; and aligning the graph-level features with the alignment traces generated by the RGB visual branch in a shared embedding space by comparing loss functions.

[0041] Specifically, for the j-th RAW-RGB image processing flow P defined in the data generation step... j Transform it into a directed weighted topological graph This topology graph is used to explicitly model the types of operations, execution order, and parameter dependencies in the complete imaging chain from RAW to RGB. Each node in the node set V represents an operation in the processing flow, including but not limited to desacrifice nodes, white balance nodes, color correction nodes, tone mapping nodes, and optional post-processing nodes such as compression, scaling, blurring, and noise addition. Each directed edge in the directed edge set E represents the order and flow dependency between two operations; for example, an edge from a desacrifice node to a white balance node indicates that the white balance operation is performed after the desacrifice operation. Edge weight function. Each directed edge is assigned a weight vector, which consists of normalized parameters of adjacent operation nodes, such as the RGB channel gain coefficient in white balance operation, matrix elements in color correction matrix, control point coordinates in tone mapping curve, and quality factor in compression operation.

[0042] In a preferred implementation, the directed weighted topology graph is constructed according to the actual camera ISP pipeline sequence, that is, the directed edges start from the demosaic node, connect to the subsequent operation nodes in sequence, and finally reach the last post-processing node; each node may contain an algorithm type identifier for the operation in addition to the type identifier (e.g., the demosaic algorithm is bilinear or edge adaptive); the edge weight is formed by concatenating the core parameters of adjacent operations after standardization.

[0043] The directed weighted topology Q j The input is fed into a Graph Neural Network (GNN), such as a Graph Convolutional Network (GCN), a Graph Attention Network (GAT), or a Graph Isomorphic Network (GIN). The GNN aggregates the neighbor information of each node and updates the node feature representation through a message-passing mechanism. After multiple layers of graph convolution, global pooling (such as summation pooling, average pooling, or attention pooling) is performed on the features of all nodes in the entire graph to obtain a graph-level feature vector. The dimension of this feature vector is the same as the RGB side alignment trace F. j The dimensions are the same. Furthermore, a contrastive learning strategy is employed, using cross-modal contrastive alignment loss as the loss for the RAW topology branch. Used to transform graph-level features Alignment mark F with RGB side j Alignment is performed within a shared embedding space to constrain consistency between the two. Specifically, for the same RAW-RGB processing flow P j The corresponding graph features and RGB side alignment traces are treated as positive sample pairs, and their corresponding graph features are minimized. Alignment mark F with RGB side j The distance between them; for different processing flows P j With P kThe feature pairs corresponding to (j≠k) are considered as negative sample pairs, and the maximum value is maximized. With F k The distance between them. The loss function used can be contrast loss (InfoNCE loss), triplet loss, or cross-modal matching loss. Through this constraint, the finally extracted alignment traces implicitly encode the structural logic and parameter information of the RAW-RGB processing flow. That is, the alignment traces not only depend on the apparent statistics of the RGB images, but are also supervised by the topology of the imaging flow.

[0044] To further constrain alignment traces from a visual statistical distribution perspective, a RAW visual branch is set up. The RAW visual branch includes: encoding the RAW image and the corresponding RGB image using a pre-trained autoencoder or variational autoencoder respectively to obtain the RAW latent distribution and the RGB latent distribution; using the alignment traces generated by the RGB visual branch as conditional information, modulating the mean of the RAW latent distribution through an attention mechanism; and using KL divergence to constrain the consistency between the modulated RAW latent distribution and the RGB latent distribution.

[0045] Specifically, a pre-trained autoencoder or variational autoencoder (VAE) is used to process the RAW image R... i and the corresponding RGB image Encoding is performed to obtain the distribution parameters of both in the latent space. Taking a variational autoencoder as an example, the encoder network... Receive RAW image R i Output the Gaussian parameters of its latent distribution: mean vector Sum of logarithmic variance vector Encoder network Receive RGB images Output the mean vector of its potential distribution. Sum of logarithmic variance vector Recorded as , .

[0046] To correlate the RAW distribution with the RGB side-alignment traces, an attention modulation mechanism is introduced. (The RGB side-alignment traces are then used as a reference.) As conditional information, the latent mean of the RAW image is modulated by an attention module (e.g., a cross-attention-based modulation layer). :calculate The attention module will Treat it as a query, Treating them as keys and values, the output is the modulated mean vector. This allows the mean of the RAW latent distribution to be adaptively shifted based on the alignment traces of the current imaging process.

[0047] Kullback-Leibler divergence is used as a measure of distribution difference to constrain the consistency between the latent distribution of the modulated RAW image and the latent distribution of the RGB image. The loss function for the RAW visual branch is constructed as follows: ; This represents the KL divergence; the loss function makes the latent distribution of the RAW image modulated with alignment traces approximate the latent distribution of the RGB image; unlike directly using pixel-level reconstruction errors, the distribution consistency constraint can tolerate a certain degree of noise, slight spatial shifts, and local pixel variations, allowing the model to learn a more robust RAW-RGB visual correspondence. Meanwhile, the modulation process introduces alignment traces... When this loss is backpropagated, it forces the alignment traces to carry information that can bridge the distribution differences between RAW and RGB, thereby enhancing the ability of the alignment traces to encode the visual statistical features of physical imaging.

[0048] The loss of the above RAW topology branches Loss of RAW visual branches By performing a weighted combination, we obtain the total loss function: ; in and These are preset regularization weights used to balance the impact of structural and distributional constraints on model training. By minimizing the total loss function, the parameters of the feature extraction model (including the parameters of the hierarchical entropy attention module, the graph neural network parameters, and the optional modulation module parameters) are iteratively updated using stochastic gradient descent or its variant optimizer.

[0049] During the joint optimization process, the model simultaneously learns three alignment relationships: 1) the correspondence between RGB side alignment traces and the corresponding imaging process topology (through...). ); 2) The correlation between RGB side alignment traces and RAW-RGB visual distribution consistency (through Attention modulation and KL divergence implementation in the model); 3) Mapping function from the RGB image itself to the alignment trace (implemented through global forward propagation).

[0050] By aligning the traces and modeling the process, a fully trained feature extraction model is obtained. This model can accept any RGB image as input and output a low-dimensional embedding vector, namely the RAW-RGB alignment trace. This alignment trace characterizes whether the image has undergone the complete imaging process from RAW signal to RGB image in the real physical world. Formally, let the RGB image to be detected be I, then the alignment trace F = f(I), where... It is a feature extraction function obtained through training, which outputs a feature embedding of dimension d. This alignment trace has content invariance (images with different content under the same image processing flow produce similar alignment traces) and process separability (images with the same content under different image processing flows produce distinguishable alignment traces), thus providing effective feature basis for subsequent detection stages to determine whether an image has a real physical source.

[0051] S3: Input the RGB image to be detected into the alignment trace extraction model, extract the alignment traces, and determine whether the image is an artificial intelligence generated image based on the alignment traces.

[0052] like Figure 3 As shown, after completing step S2, aligning trace modeling, and obtaining a fully trained feature extraction model, the detection application step is executed; the feature extraction model is used to determine the authenticity of the RGB image to be detected, judging whether it originates from the real physical world RAW signal acquisition process or is directly synthesized by an artificial intelligence generation model; the detection application step supports zero-sample detection mode and few-sample detection mode.

[0053] Specifically, the RGB image to be detected is input into the feature extraction model obtained during the training phase. The feature extraction model processes the input image via forward propagation, sequentially passing it through a pre-trained visual feature extraction network and a hierarchical entropy attention module in the RGB visual branch, and outputs the corresponding RAW-RGB alignment trace feature vector. The position of the alignment trace feature vector in the embedding space reflects the degree of correlation of the RAW-RGB imaging process contained in the input image. If the image is acquired from a real camera, its alignment trace will fall in the same distribution area as the real image in the training stage. If the image is directly synthesized by an artificial intelligence generation model (such as a generative adversarial network, a diffusion model, or an autoregressive model), its alignment trace will deviate significantly from the real distribution due to the lack of a real RAW signal source and physical imaging process.

[0054] Without relying on artificial intelligence to generate prior image annotation data, an unsupervised clustering method is used for detection, including: for a set of RGB images to be detected. Each image is input into the feature extraction model to obtain the corresponding alignment trace feature set. Unsupervised clustering algorithms (such as K-Means, spectral clustering, DBSCAN, or hierarchical clustering) are used to cluster the alignment trace feature set. The number of clusters can be preset to 2 (corresponding to the two categories of "real images" and "artificial intelligence generated images"), or it can be determined adaptively. Since the RAW-RGB alignment traces of real images are constrained to have process separability during the training phase, while artificial intelligence generated images lack the same physical constraints, the two naturally form separate clusters in the alignment trace space. For each cluster obtained after clustering, its category can be determined by any of the following methods: 1) If there are a few known real reference images, the clusters of their alignment traces are marked as real class, and the remaining clusters are marked as generated class; 2) In the absence of any prior information, the similarity between the alignment traces within the cluster and the distribution of alignment traces in real images during the training phase can be used for discrimination; 3) Optionally, the visual quality indicators of the image itself (such as noise level, color consistency, etc.) can be combined to assist in the distinction.

[0055] For scenarios involving a single image to be detected or where a batch set cannot be formed, zero-shot detection is performed using a similarity-based matching method, including: (1) constructing a reference distribution. At least one real RAW image is pre-selected from the training dataset and converted into an RGB image using a standard image processing procedure (such as a processing procedure using typical camera ISP parameters), or any real source RGB image from the training phase is directly used and input into the feature extraction model to obtain reference alignment traces. To improve robustness, multiple real RGB images from different scenes can be selected, and the mean value of their corresponding alignment traces can be calculated. Covariance Matrix (2) Calculate the distance or similarity for the alignment marks of the image to be detected. Calculate the alignment trace with the reference. Or refer to the measurement values ​​between distributions; Euclidean distance, cosine similarity or Mahalanobis distance can be used as the measurement method. (3) Use a threshold to make a judgment. Pre-set a judgment threshold. (This can be determined by the distribution of alignment marks on the validation set based on real images, for example, by taking the mean ± several times the standard deviation); when the metric meets the judgment condition, the image to be detected is judged as a real image; otherwise, it is judged as an AI-generated image. For distance metrics, the judgment condition is d≤τ; for similarity metrics, the judgment condition is sim≥τ. (4) Optionally, a few-sample detection is used. When there is a small amount of labeled data, the parameters of the trained alignment mark extraction model are frozen, and a linear classifier is added to the output of the alignment mark extraction model; a small number of samples labeled with AI-generated images or real images are used to train the linear classifier; the image to be detected is input into the alignment mark extraction model to obtain alignment marks, and then the alignment marks are input into the trained linear classifier to obtain the classification result.

[0056] The output of the detection application step includes a category determination result (real image or AI-generated image) for each image to be detected, as well as an optional confidence score. This detection application step enables effective authenticity detection even without prior AI-generated image samples (zero-sample mode), and can further improve detection performance when a small number of labeled samples are available (few-sample mode), thus making it suitable for diverse data constraints in real-world forensic scenarios.

[0057] Example 2 This embodiment exemplarily presents a synthetic image detection system based on digital imaging alignment traces, including a data generation module, a feature extraction model, and a detection module.

[0058] The data generation module acquires real RAW images and converts them into corresponding RGB images through various predefined RAW-RGB image processing workflows, constructing the RAW-RGB image pair dataset required for training. The input to this module is a set of real RAW images output from at least one physical image acquisition device (e.g., a digital camera or smartphone). Internally, the data generation module stores or dynamically generates various predefined RAW-RGB image processing workflows. Each workflow includes at least a demosaicing operation and optionally includes core ISP operations such as white balance, color correction, and tone mapping, as well as optional post-processing operations such as compression, scaling, blurring, or noise addition. For each input RAW image, the data generation module sequentially calls multiple image processing workflows to convert the RAW image into a corresponding RGB output image, forming a training dataset consisting of (RAW image, RGB image) pairs. The output of this module is a data stream or data storage file that can be used for training the feature extraction model.

[0059] The feature extraction model, comprising an RGB visual branch, a RAW topology branch, and a RAW visual branch, is trained through joint optimization. It is used to extract alignment traces reflecting the consistency of the RAW-RGB image processing flow from input RGB images. The model accepts any RGB image as input and outputs a low-dimensional embedding vector, i.e., the RAW-RGB alignment trace. For the RGB visual branch, a pre-trained visual feature extraction network is used to acquire local spatial features of the input image. Learnable initial latent features are used as query features, initialized during training and jointly optimized with the model. A hierarchical entropy attention module performs physical entropy-driven hierarchical sampling and attention fusion on the local spatial features, outputting the RGB side alignment trace. This branch is used to extract low-level visual information related to the physical imaging process from the appearance of the RGB image. For the RAW topology branch, each RAW-RGB image processing flow defined in the data generation module is converted into a directed weighted topology graph. A graph neural network is used to extract the structural features of the flow, and a contrastive loss is used to align these structural features with the alignment traces output by the RGB visual branch in a shared embedding space. This branch enables the alignment traces to implicitly encode the operational order and parameter dependencies of the imaging flow. For the RAW vision branch, a pre-trained autoencoder or variational autoencoder is used to encode the latent distribution of the RAW image and the corresponding RGB image, respectively. The alignment trace generated by the RGB vision branch is used as conditional information to modulate the RAW latent distribution. The consistency between the modulated RAW distribution and the RGB distribution is constrained by KL divergence. This branch enables the alignment trace to simultaneously carry the visual statistical correlation between RAW and RGB.

[0060] The three branches are jointly trained through optimization (minimizing the weighted sum of topological contrast loss and visual distribution loss) to obtain an alignment trace extraction model that can be used independently for inference. During the detection phase, only the forward path of this model (i.e., the RGB visual branch part, or the whole model but only performing the mapping from RGB image to alignment trace) needs to be retained, and there is no need to use the RAW topology branch and RAW visual branch again.

[0061] The detection module acquires alignment traces extracted from the RGB image to be detected by the feature extraction model and determines whether the image is an AI-generated image based on these alignment traces. This module supports zero-shot and few-shot detection modes. For zero-shot detection, the module incorporates an unsupervised clustering algorithm (such as K-Means) or a similarity measurement algorithm (such as cosine similarity or Euclidean distance). For a batch of images to be detected, the module first calls the feature extraction model to acquire alignment traces one by one, then performs clustering or calculates the distance between each alignment trace and a pre-stored real reference alignment trace. Based on the clustering or distance threshold, it outputs the category label ("real image" or "AI-generated image") and the corresponding confidence score for each image. For few-shot detection, the module adds a trainable linear classifier (e.g., a single-layer fully connected network) after the feature extraction model. When a small number of samples labeled with AI-generated or real images are provided, the module freezes all parameters of the feature extraction model and only updates the weights and biases of the linear classifier. After training, for the image to be detected, the module extracts its corresponding alignment traces and inputs them into the linear classifier to directly obtain a binary classification result. The output of the detection module can be provided to users or downstream forensic analysis systems in the form of visual labels, data files, or API responses.

[0062] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims. It should be understood that the invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A synthetic image detection method based on digital imaging alignment traces, characterized in that, The method includes: A training dataset was constructed using real RAW images and corresponding RGB images generated through various predefined RAW-RGB image processing workflows. A multi-branch network containing an RGB visual branch, a RAW topology branch, and a RAW visual branch is constructed. An alignment trace extraction model is obtained by joint optimization training using the training dataset. The alignment trace extraction model is used to extract alignment traces from the input RGB image that reflect the consistency of the RAW-RGB image processing flow. The RGB image to be detected is input into the alignment trace extraction model to extract alignment traces, and the image is determined to be an artificial intelligence-generated image based on the alignment traces.

2. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The RAW-RGB image processing workflow includes at least a demosaicing operation, and further includes one or more of white balance, color correction, and tone mapping; wherein white balance and tone mapping each appear at most once in the same RAW-RGB image processing workflow; the RAW-RGB image processing workflow also includes optional image post-processing operations, which include one or more of compression, scaling, blurring, or noise addition.

3. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The RGB visual branch includes: inputting the RGB image into a pre-trained visual feature extraction network to obtain a spatial feature representation; inputting the spatial feature representation into a hierarchical entropy attention module, and using learnable initial latent features as query features to output alignment traces; The hierarchical entropy attention module is used to calculate the physical entropy of each local region in the spatial feature representation, divide each local region into multiple entropy intervals according to the physical entropy, sample local features hierarchically from each entropy interval, use the query feature as the query vector, use the local feature as the key value vector, and perform weighted fusion through the attention mechanism to generate the alignment trace.

4. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The RAW topology branch includes: representing each RAW-RGB image processing flow as a directed weighted topology graph, where nodes represent processing operations, directed edges represent operation order, and edge weights represent operation parameters; inputting the directed weighted topology graph into a graph neural network to extract graph-level features; and aligning the graph-level features with the alignment traces generated by the RGB visual branch in a shared embedding space using a loss function.

5. The synthetic image detection method based on digital imaging alignment traces according to claim 4, characterized in that, For the same image processing workflow, the loss function minimizes the distance between the graph-level features of the workflow and the alignment traces of the workflow; for different image processing workflows, the loss function maximizes the distance between the graph-level features of the workflow and the alignment traces of other workflows.

6. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The RAW visual branch includes: encoding the RAW image and the corresponding RGB image using a pre-trained autoencoder or variational autoencoder respectively to obtain the RAW latent distribution and the RGB latent distribution; using the alignment traces generated by the RGB visual branch as conditional information, modulating the mean of the RAW latent distribution through an attention mechanism; and using KL divergence constraints to ensure consistency between the modulated RAW latent distribution and the RGB latent distribution.

7. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The joint optimization is achieved by minimizing the total loss function, which is a weighted sum of the losses of the RAW topology branch and the RAW vision branch. The loss of the RAW topology branch is used to constrain the consistency between the alignment traces and the topology of the image processing flow, and the loss of the RAW vision branch is used to constrain the consistency of the distribution of RAW images and RGB images in the latent space.

8. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The step of determining whether an image is an AI-generated image based on the alignment traces employs a zero-shot detection method, including: Multiple alignment marks are extracted from multiple images to be detected. An unsupervised clustering algorithm is used to cluster the multiple alignment marks. Based on the clustering results, real images and artificial intelligence generated images are distinguished. Alternatively, the similarity or distance between the alignment marks of the image to be detected and the pre-stored alignment marks of the real image reference can be calculated, and the determination can be made based on the relationship between the similarity or distance and a preset threshold.

9. The synthetic image detection method based on digital imaging alignment traces according to claim 1, characterized in that, The method for determining whether an image is an AI-generated image based on alignment traces employs a few-shot detection approach, including: freezing the parameters of a trained alignment trace extraction model; adding a linear classifier to the output of the alignment trace extraction model; training the linear classifier using a small number of samples labeled with AI-generated or real images; inputting the image to be detected into the alignment trace extraction model to obtain alignment traces; and then inputting the alignment traces into the trained linear classifier to obtain a classification result.

10. A synthetic image detection system based on digital imaging alignment traces, used to implement the method as described in any one of claims 1-9, characterized in that, include: The data generation module is used to acquire real RAW images and convert them into corresponding RGB images through various predefined RAW-RGB image processing procedures to build a training dataset. The feature extraction model, which includes an RGB visual branch, a RAW topology branch, and a RAW visual branch, is obtained through joint optimization training. It is used to extract alignment traces from the input RGB image that reflect the consistency of the RAW-RGB image processing flow. The detection module is used to acquire alignment traces extracted by the feature extraction model from the RGB image to be detected, and to determine whether the image is an artificial intelligence generated image based on the alignment traces.