Digital image identification method and system based on artificial intelligence
By constructing a three-layer fusion neural network architecture, combining physical parameter extraction and semantic understanding, the problems of low accuracy and insufficient interpretability in AI-generated image identification are solved, achieving efficient and stable identification of AI-generated images.
Patent Information
- Application Number
- CN202511096136.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-08-06
AI Technical Summary
Existing technologies lack a unified framework for verifying physical laws and analyzing semantic logic when dealing with AI-generated images, resulting in low detection accuracy, lack of interpretability and adaptability, making it difficult to cope with rapidly evolving image generation technologies.
A three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer. Through a physical-semantic cross-attention mechanism and a Bayesian inference framework, high-accuracy identification of AI-generated images is achieved.
It improves the accuracy of identifying AI-generated images, can stably detect in high-quality generated images, provides interpretable identification results, and has adaptive capabilities, making it suitable for various image types and forgery techniques.
Smart Images

Figure CN120997645A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of digital image identification, more specifically, it relates to a digital image identification method and system based on artificial intelligence. BACKGROUND
[0002] Modern AI-generated images have reached a level of visual quality that is almost indistinguishable from real photos, posing a serious challenge to traditional image authenticity identification techniques.
[0003] Currently, the existing technologies in the field of image authenticity identification mainly fall into the following categories: Traditional methods based on pixel-level statistical features: Early image identification techniques mainly rely on JPEG compression trace analysis, pixel value statistical property detection, frequency domain feature analysis, etc. For example, by detecting double JPEG compression traces, CFA (Color Filter Array) interpolation traces or statistical distribution anomalies of pixel values in the image to identify image tampering. However, this method has limitations: first, modern generative adversarial networks can generate highly realistic images at the pixel level, and traditional statistical feature detection methods are easily ineffective; second, high-quality image post-processing techniques can effectively eliminate these statistical traces, making detection more difficult.
[0004] Identification methods based on physical models: This method analyzes the physical features such as light consistency, shadow direction, perspective relationship in the image to determine the authenticity of the image. Typical techniques include light source direction estimation, shadow geometry verification, mirror reflection analysis, etc. Although these methods have some effect in detecting image tampering that obviously violates physical laws, they face the following problems: current advanced image generation techniques can simulate basic physical laws to some extent, generating images that conform to lighting and geometric constraints; in addition, this method usually only focuses on local physical features, lacking comprehensive evaluation of global physical consistency.
[0005] Detection methods based on semantic consistency: This technology analyzes the semantic logical rationality of image content to identify fake images, including object relationship detection, scene rationality analysis, spatiotemporal consistency verification, etc. For example, detecting whether there are unreasonable object combinations, scene layouts that violate common sense, or temporal logic contradictions in the image. However, these methods also have limitations: with the development of large-scale language models and multi-modal AI technology, modern image generation systems can produce highly reasonable image content at the semantic level; at the same time, pure semantic methods often ignore the physical reality of the image, and cannot detect images that are logically reasonable in content but impossible in physical implementation.
[0006] The existing technologies generally have the following fundamental problems: Various detection methods usually work independently, lack a unified theoretical framework to organically combine physical verification and semantic understanding, and cannot fully utilize the complementary advantages of the two dimensions; most existing methods use deterministic binary classification judgment, lack a probabilistic reasoning mechanism, and cannot quantify the uncertainty of the identification results; the existing technology generally lacks explainability, cannot provide intuitive and understandable explanations for the identification results, and is difficult to obtain user trust and professional recognition; in the face of rapidly evolving image generation technology, the generalization ability and adaptability of existing methods are insufficient, and frequent updates of training data and model parameters are required.
[0007] These technical limitations show that it is urgent to develop a new type of image identification theory and method that can unify physical law verification and semantic logic analysis to cope with the increasingly complex AI-generated image detection challenges. SUMMARY
[0008] The present application provides a digital image identification method and system based on artificial intelligence, which solves the technical problems of low recognition accuracy of AI-generated images, lack of unified verification framework of physical and semantic, insufficient explainability of detection results and weak adaptability to new forgery techniques in related technologies.
[0009] The present application provides a digital image identification method based on artificial intelligence, comprising: Constructing a three-layer fusion neural network architecture, including a physical parameter extraction network layer, a semantic understanding network layer and a variational inference engine layer; Using the physical parameter extraction network layer to analyze the input image and generate a set of physical feature vectors; Applying the semantic understanding network layer to process the input image to produce semantic feature vectors and relationship graphs; Based on the physical feature vectors and the semantic feature vectors, the variational inference engine layer is used to construct a physical-semantic joint distribution; Implementing a physical-semantic cross-attention mechanism to enhance the bidirectional information flow between physical features and semantic features; Analyzing the deviation degree of the physical-semantic joint distribution from the pre-established standard distribution of natural images; Based on the deviation degree, generating the identification result and the explainability analysis report of the input image.
[0010] Further, the construction method of the physical parameter extraction network layer comprises: Using a deep convolutional neural network structure, including a feature extraction module and a physical parameter estimation module; Applying multi-scale convolution filters to the input image to extract image features of different scales; Fusing multi-scale features through a feature pyramid structure to form a comprehensive feature map; A specialized physical parameter estimation module is applied to infer lighting model parameters, surface reflectance properties, shadow consistency, and perspective relationships from the synthesized feature map. An uncertainty estimate of the physical parameters is calculated to provide confidence information for subsequent probabilistic reasoning.
[0011] Further, the construction method of the semantic understanding network layer includes: A Transformer-based visual model is used to extract semantic features of the input image; An object detection and segmentation module is applied to identify key objects and regions in the input image; An object relationship graph is constructed to represent the spatial and functional relationships between objects in the input image; An event reasoning module is applied to analyze events in the scene and their rationality.
[0012] Further, the construction method of the variational inference engine layer includes: The physical feature vector and the semantic feature vector are input into the variational autoencoder to map to a shared latent space; Conditional probability modeling is applied in the latent space to calculate physical compliance probability and semantic consistency probability; The conditional probabilities are integrated into a joint probability distribution through a Bayesian inference framework; The Monte Carlo integration method is applied to calculate the integral to obtain the authenticity probability estimate.
[0013] Further, the step of implementing the physical-semantic cross-attention mechanism includes: A physical-to-semantic attention map is constructed to guide the semantic understanding to focus on abnormal physical regions; A semantic-to-physical attention map is constructed to make the physical analysis preferentially process semantically important regions; Through alternating iterative optimization of the two attention maps, deep fusion of physical and semantic features is achieved; The fused feature representation is extracted for subsequent anomaly detection.
[0014] Further, the step of analyzing the degree of deviation of the physical-semantic joint distribution from the pre-established standard distribution of natural images includes: A standard distribution model based on large-scale real image data is established as a reference benchmark; The Kullbac Leibler divergence between the joint distribution of the input image and the standard distribution is calculated to quantify the degree of deviation; An adaptive threshold is set to judge the authenticity of the input image based on the degree of deviation; The key regions causing the deviation are located, and a visualized anomaly heat map is provided.
[0015] Further, the step of generating the identification result and the explainability analysis report of the input image includes: generating a probability score of the authenticity of the input image, quantifying the identification confidence; identifying potential abnormal areas and labeling them on the input image; providing an abnormal reason analysis to explain the violation of physical laws and semantic logic; generating a detailed identification report including various indicators and comprehensive evaluation results.
[0016] Further, the step of integrating conditional probabilities into joint probability distribution through the Bayesian inference framework specifically includes: temperature scaling calibration of the physical semantic joint distribution, adjusting the temperature parameter to improve the reliability of the probability estimation of the calibrated probability distribution; using a variational autoencoder to establish a multi-scale interaction strategy to realize the interaction of physical feature vectors and semantic feature vectors at different abstraction levels through cross-layer connection; applying a feature enhancement gate unit to control the strength of information interaction between physical feature vectors and semantic feature vectors.
[0017] Further, the collaborative working mechanism of the physical parameter extraction network layer and the semantic understanding network layer specifically includes: pre-classifying the input image through an image type classifier, and automatically adjusting the parameter configuration of the multi-scale convolution filter and the attention weight of the Transformer-based visual model according to the classification result; building an incremental learning module in the physical parameter extraction network layer and the semantic understanding network layer, so that the deep convolutional neural network structure and the Transformer-based visual model can continuously learn from new forgery technology samples and dynamically update the network weights; establishing a multi-modal fusion mechanism to fuse the illumination model parameters and surface reflection characteristics output by the physical parameter estimation module with the key object information identified by the object detection and segmentation module, and simultaneously integrating image metadata for comprehensive analysis.
[0018] The present application provides an artificial intelligence-based digital image identification system for executing the above-mentioned artificial intelligence-based digital image identification method, comprising: a three-layer fusion neural network architecture module for building a neural network architecture including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer; a physical analysis module for analyzing the input image using the physical parameter extraction network layer to generate a set of physical feature vectors; a semantic understanding module for applying the semantic understanding network layer to process the input image to generate semantic feature vectors and relationship graphs; a cross-attention module configured to implement a physical-semantic cross-attention mechanism to enhance bidirectional information flow between physical features and semantic features; a joint distribution construction module configured to construct a physical-semantic joint distribution based on the physical feature vector and the semantic feature vector using a variational inference engine layer; a distribution analysis module configured to analyze a deviation degree of the physical-semantic joint distribution of the input image from a pre-established standard distribution of natural images; a result output module configured to generate an identification result and an explainability analysis report of the input image based on the deviation degree.
[0019] The present application has the beneficial effects that: through the dual verification mechanism of the physical parameter extraction network layer and the semantic understanding network layer, the detection capability of AI generated images is improved, especially the identification accuracy of high-quality generated images is improved; The coupling relationship verified by the physical-semantic cross-attention mechanism can maintain stable detection performance even under targeted adversarial sample attacks as long as there is inconsistency in the physical-semantic relationship; The probability distribution analysis and abnormal heat map generation technology constructed by the variational inference engine layer can accurately locate the specific area of the conflict between physics and semantics in the forged image and provide detailed abnormal reason analysis, providing a reliable scientific basis for the identification result; Through the adaptive learning mechanism based on the Bayesian inference framework, the system can automatically optimize the physical-semantic coupling model as the sample processing progresses, gradually improving the recognition ability of unknown forgery techniques.
[0020] The complex multi-dimensional verification problem is converted into probability distribution similarity calculation through the "double helix" theoretical framework, which reduces the computational complexity of the algorithm, improves the processing efficiency, and makes real-time processing of high-resolution images possible; Through the physical-semantic joint distribution modeling technology, it is suitable for multiple image types and forgery techniques, and is not limited to specific image content or tampering means, and shows good stability in different application scenarios; Through the modular three-layer fusion neural network architecture design, the system is easy to integrate new physical models or semantic understanding technologies according to actual needs, allowing individual updating or replacement of specific components without affecting the overall architecture, providing good scalability for technology iteration; The present application realizes the paradigm shift from "finding forgery traces" to "verifying naturalness", providing a new theoretical basis and technical path for the field of digital image identification, and helping to cope with the increasingly complex image forgery challenges. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1is a flow chart of a digital image authentication method based on artificial intelligence in the present invention; Figure 2 is a scatter plot of the relationship between Kullback-Leibler divergence and the probability of real images; Figure 3 is a combination chart of the detection accuracy and false positive rate of the system under different threshold settings; Figure 4 is a column chart comparing the accuracy of the physical-semantic double helix theory, traditional physical verification and traditional semantic checking method in identifying ordinary GAN images and high-quality GAN images; Figure 5 is a bar chart of the authentication accuracy of the physical-semantic double helix theory and traditional methods on different types of images; Figure 6 is a radar chart of the performance comparison of the physical-semantic double helix theory and traditional methods in six dimensions. DETAILED DESCRIPTION
[0022] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is merely meant to provide a better understanding of the subject matter described herein and can be changed in function and arrangement without departing from the scope of the present description. Various examples can omit, substitute, or add various procedures or components as appropriate or desired. Also, some features described with respect to one example can be combined with features of other examples.
[0023] In at least one embodiment of the present invention, a digital image authentication method based on artificial intelligence is disclosed, as shown in Figure 1 comprising: Step 1, a three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer and a variational inference engine layer; Step 1.1, the physical parameter extraction network layer; A deep convolutional neural network structure is adopted, including a feature extraction module and a physical parameter estimation module, which is used to extract physical feature parameters such as illumination distribution, shadow direction and perspective relationship from the input image. Specifically, this network uses ResNet-50 as the backbone network, and adds a special physical parameter estimation branch on its basis, and each branch is optimized for a specific physical attribute. For example, the light estimation branch estimates the direction and intensity of the main light source by analyzing the brightness gradient and shadow distribution of the image; the perspective relationship branch estimates the camera parameters by detecting vanishing points and parallel lines. Optionally, in some implementations, DenseNet or EfficientNet can also be used as the backbone network to balance performance and computational efficiency in different scenarios.
[0024] The structure of the physical parameter extraction network layer can be further refined into four main components: The backbone feature network adopts the ResNet-50 structure, which contains 5 stages, and each stage contains 3, 4, 6, and 3 residual blocks, respectively. Each residual block contains three convolutional layers and a shortcut connection, effectively solving the gradient vanishing problem in deep networks. The network input is a 224x224x3 RGB image, which is processed through the initial 7x7 convolution and max pooling, and then through the residual blocks of each stage, finally outputting a 2048-channel feature map. Optionally, in specific application scenarios, different depth variants of ResNet (such as ResNet-34 or ResNet-101) can be used to balance the computational complexity and feature extraction capability.
[0025] The light estimation module is composed of a three-layer 1x1 convolutional network, which reduces the channel number from 2048 to 1024, 512, and 256, respectively, and then outputs the light source parameters through global average pooling and a fully connected layer. The light source parameters include the main light source direction (azimuth and elevation), light source intensity (normalized value), and ambient light intensity. This module also integrates a light consistency evaluation unit to detect potential inconsistencies by comparing the light parameter estimation results in different regions of the image.
[0026] The shadow analysis module uses the U-Net structure for pixel-level shadow segmentation, and then applies a gradient direction analysis algorithm to infer the light source direction. This module first outputs a shadow probability map, then locates the object and shadow intersection through shadow edge detection, and finally calculates the light source direction based on geometric relationships. In actual implementation, the results of this module and the light estimation module are cross-validated to improve the reliability of parameter estimation.
[0027] The perspective geometry module includes a line segment detection network and a vanishing point estimation network. The line segment detection network is based on a neural network implementation of the LSD (Line Segment Detector) algorithm, which can identify straight line segments in the image; the vanishing point estimation network determines the main vanishing point location through the RANSAC algorithm for grouping and voting on the detected line segments, and then infers the camera parameters and scene geometry. In some embodiments, this module can also introduce prior geometric constraints, such as the Manhattan World assumption, to further improve the accuracy of geometric parameter estimation.
[0028] In the training phase, the physical parameter extraction network uses a multi-task learning framework to simultaneously optimize multiple physical parameter estimation targets. The loss function consists of two parts: parameter estimation loss and physical consistency loss: ; wherein, Ltotal represents the total loss of the physical parameter extraction network; ∑ represents the summation symbol; Lparam represents the parameter estimation loss; Lconsistency represents the physical consistency loss. a loss weight coefficient of a physical parameter; an estimated loss of a physical parameter (such as an angle loss of light source direction, a mean square error of intensity, etc.); a weight coefficient of a physical consistency loss; a physical consistency constraint loss (such as the consistency of light source direction and shadow direction).
[0029] Through this multi-task learning manner, the network can learn the internal correlation between physical parameters, and improve the overall estimation accuracy.
[0030] Step 1.2, semantic understanding network layer; Based on the Transformer architecture, it contains a self-attention module and a semantic relationship inference module, which is used to identify objects in the image, understand the relationship between objects and infer the logic of the event. This network adopts the VisionTransformer (ViT) structure, which captures long-distance semantic dependencies by dividing the image into a series of blocks (patches) and calculating the self-attention relationship between these blocks. In specific implementation, the input image is first divided into 16x16 pixel blocks, and then processed by a 12-layer Transformer encoder to finally generate object recognition results and relationship graphs. It should be understood that in another embodiment, SWINTransformer or a structure based on the combination of convolution and Transformer can also be used to better capture local and global information.
[0031] The structure of the semantic understanding network layer can be further refined into the following five key parts: Image blocking and embedding: The input image is first divided into N fixed-size blocks (patches), each with a size of 16x16 pixels. For an input image with a resolution of 224x224, a total of 196 blocks are generated. Then, each block is mapped to an embedding vector with a dimension of 768 through a linear projection layer. In addition, position encoding (position embedding) is added to retain spatial position information, and a learnable classification token (class token) is added, which will be used for final image-level prediction. Optionally, in some implementations, image blocks of different sizes (such as 8x8 or 32x32) can be used, or a two-dimensional position encoding method can be used to adapt to different image size and resolution requirements.
[0032] Transformer encoder stack: composed of 12 standard Transformer encoder blocks, each block contains a multi-head self-attention (MSA, Multi-Head Self-Attention) module and a feed-forward neural network (FFN, Feed-Forward Network) module. In self-attention calculation, each block uses 12 attention heads, each head has a dimension of 64, and the total calculation formula is: ; where, represents the attention output; represents the query matrix; represents the key matrix; represents the value matrix; represents the transpose operator; represents the softmax normalization operation; represents the dimension of the key vector, used to scale the dot product result to prevent gradient disappearance.
[0033] The feed-forward neural network is composed of two fully connected layers, using GELU activation function in the middle, the dimension expands from 768 to 3072 and then back to 768. After each module, layer normalization (LayerNorm) and residual connection are applied to ensure training stability and information flow.
[0034] Object detection and segmentation module: based on the output features of the Transformer encoder, add a special object detection head and instance segmentation head. The detection head uses a design similar to DETR (Detection Transformer), using learnable object queries and encoder output features for cross-attention calculation, directly predicting bounding boxes and classes. In specific implementation, 100 object queries are used, each query interacts with image features through a 6-layer Transformer decoder, and finally outputs object position (bounding box coordinates) and class probability. The segmentation head is based on the detection result, using dynamic convolution to generate a mask for each object, realizing pixel-level segmentation. It should be noted that unlike traditional object detection methods, this module does not require non-maximum suppression post-processing, simplifying the algorithm process and improving the speed.
[0035] Relationship graph construction module: based on the detected objects, construct a scene graph (SceneGraph). The relationship graph uses a graph structure: ; where represents the scene graph, represents the node set, represents the edge set.
[0036] Relationship prediction employs a bidirectional attention mechanism: first, relationship features are calculated from the features and location information of the two objects; then, the relationship category is predicted using a multilayer perceptron. In practice, to control computational complexity, only one attention is considered for each object. The most relevant other objects ( Instead of calculating all possible object pairs, this module uses a knowledge graph (=10) to guide the relationship prediction process by distilling knowledge. Furthermore, to enhance the accuracy of relationship prediction, this module introduces an external knowledge graph as prior knowledge.
[0037] Event Logic Reasoning Module: This module performs high-order event logic reasoning based on the relationship graph. It employs a Graph Convolutional Network (GCN) architecture, using three layers of graph convolutions for message passing. The convolution operation is defined as follows: ; in, Indicates the first The node feature matrix of the layer; Represents the ReLU activation function; This represents the degree matrix after adding self-loops; This represents the adjacency matrix after adding the self-loop; Represents the negative 1 / 2 power (normalization) of the degree matrix; Indicates the first The node feature matrix of the layer; Indicates the first The learnable weight matrix of the layer.
[0038] Through iterative graph convolution operations, information propagates within the relationship graph, allowing each node to gradually aggregate contextual information, forming a more global scene understanding. Ultimately, this module outputs scene event type predictions and abnormal event detection results. Optionally, in some implementations, a temporal reasoning mechanism can be introduced to analyze the event evolution logic across multiple frames, further enhancing semantic understanding capabilities.
[0039] The semantic understanding network layers are trained using a multi-stage strategy: first, the basic Transformer encoder is pre-trained on a large-scale image classification dataset; then, the object detection and segmentation modules are fine-tuned on object detection and instance segmentation datasets; finally, the relationship graph construction and event logic reasoning modules are trained on a scene graph dataset. The training loss function consists of a weighted sum of object classification loss, bounding box regression loss, instance segmentation loss, and relationship classification loss. ; in, This represents the total loss of the semantic understanding network layers; The weighting coefficients represent the object classification loss; Represents the object classification loss; a weight coefficient representing the bounding box regression loss; a weight coefficient representing the bounding box regression loss; a weight coefficient representing the instance segmentation loss; a weight coefficient representing the instance segmentation loss; a weight coefficient representing the relation classification loss; a weight coefficient representing the relation classification loss.
[0040] Step 1.3, Variational Inference Engine Layer; Integrating Variational Autoencoder and Bayesian Inference Network for Building Joint Probability Distribution Model of Physical Parameters and Semantic Features and Performing Explainable Inference. The engine consists of an encoder network, a latent variable space, and a decoder network. The encoder network maps physical and semantic features to the latent variable space, and the decoder reconstructs the original features from the latent representation. During training, network parameters are optimized by minimizing reconstruction error and Kullback-Leibler divergence, enabling the latent space to effectively represent the probability distribution of physical-semantic relationships. In addition, the dimension of the latent space can be adjusted for different application scenarios, balancing expressiveness and computational complexity.
[0041] The structure of the variational inference engine layer can be broken down into the following main components: Feature Integration Preprocessing: This component receives the output features of the physical parameter extraction network and the semantic understanding network, and performs preliminary feature normalization and alignment. The physical feature vector (dimension 512) and the semantic feature vector (dimension 768) are first transformed by independent fully connected layers to a common feature space (dimension 256). To ensure information integrity, residual connections and layer normalization operations are introduced: ; ; where, represents the normalized physical feature vector; represents the layer normalization operation; represents the original physical feature vector; represents the fully connected layer transformation of the physical feature; represents the normalized semantic feature vector; represents the original semantic feature vector; represents the fully connected layer transformation of the semantic feature.
[0042] Conditional Variational Encoder: This component adopts a Conditional Variational Autoencoder (CVAE) architecture to encode the normalized physical and semantic features into a probability distribution rather than a deterministic point. The encoder structure consists of three fully connected layers, each followed by batch normalization and LeakyReLU activation functions. The final output is the posterior distribution parameters of the latent variable : mean vector and log-variance vector , both with a dimension of 128.
[0043] The sampling process uses the reparameterization trick to ensure that gradients can be propagated through the random sampling operation: ; where represents the sampled latent variable; represents the mean vector of the posterior distribution; represents the standard deviation vector of the posterior distribution, , represents the log-variance vector of the posterior distribution, output by the encoder network; represents a noise vector sampled from a standard normal distribution ; represents a multivariate normal distribution with mean 0 and covariance identity matrix; represents a distribution relationship.
[0044] In specific implementations, the encoder is designed to accept conditional inputs, i.e., physical parameters θ (such as lighting conditions, camera parameters, etc.), so that the posterior distribution becomes a conditional probability , enhancing the model's adaptability to different physical conditions.
[0045] Prior Network: To more accurately model the prior distribution of physical parameters, a learnable prior network is introduced instead of simply assuming a standard normal distribution. This network receives physical parameters as input and outputs the parameters of the prior distribution : prior mean vector and log-variance vector . The network structure consists of two fully connected layers with ReLU activation functions in between. By learning the prior distribution, the model can capture the natural distribution characteristics of the latent variable under different physical conditions, improving modeling accuracy. Optionally, to further enhance the expressiveness of the prior distribution, Normalizing Flow technology can be used to enhance the complexity of the distribution through a series of reversible transformations.
[0046] Physical-Semantic Probability Decoder: This component reconstructs the physical compliance probability and semantic consistency probability from latent variables z. The decoder contains four parallel branches: Physical Reconstruction Branch: A three-layer fully connected network that reconstructs the physical features; Physical Compliance Branch: Outputs the physical compliance probability , computed by comparing the reconstructed physical features with the expected physical parameters; Semantic Reconstruction Branch: A three-layer fully connected network that reconstructs the semantic features; Semantic Consistency Branch: Outputs the semantic consistency probability , computed by evaluating the reconstructed semantic relationships against common sense knowledge.
[0047] The outputs of each branch are normalized to the 0-1 interval by a sigmoid function, representing the corresponding probability values. It is important to note that the compliance probability calculation not only considers reconstruction error, but also includes hard constraint checks of physical laws, such as shadow direction consistency with light sources, reflection angle relationship with surface normals, etc.
[0048] Probability Integration Module: Implements the integration operation in the Bayesian inference framework to calculate the posterior probability of image authenticity. This module uses the Monte Carlo method to approximate the integration: ; where, represents the image is the posterior probability of the image being a real image; represents approximately equal to; represents the summation operator; represents the number of samples; represents the physical compliance probability of image under the th physical parameter ; represents the semantic consistency probability of image under the th physical parameter ; represents the th physical parameter sampled from the prior distribution .
[0049] To improve sampling efficiency, the module implements an importance sampling strategy, which preferentially samples from high-probability regions of the parameter space. In actual deployment, the number of samples can be flexibly adjusted according to the computing resources and accuracy requirements, usually set to 1000 on high-performance devices, and can be reduced to 100-200 on resource-constrained devices, while using variance reduction techniques to maintain estimation accuracy.
[0050] The training of the variational inference engine layer is in an end-to-end manner, and the loss function is composed of four parts: ; Among them, represents the total loss of the variational inference engine layer; represents the weight coefficient of the reconstruction loss; represents the reconstruction loss (measuring the information retention ability of the encoding-decoding process); represents the weight coefficient of the KL divergence loss; represents the KL divergence loss (making the posterior distribution close to the prior distribution); represents the weight coefficient of the physical compliance loss; represents the physical compliance supervision loss; represents the weight coefficient of the semantic consistency loss; represents the semantic consistency supervision loss.
[0051] Weight coefficient , , Determined by the validation set tuning to balance the importance of different learning goals. At the beginning of training, a larger and a smaller , gradually increase the value of , realize the annealing-like KL divergence training strategy, avoid the posterior distribution collapse to the prior distribution too early.
[0052] Step 1.4, inter-layer interaction connection; Build a bidirectional attention connection mechanism to realize information interaction between the physical layer and the semantic layer, forming a "double helix" structure. Specifically, two sets of attention modules are designed: one set of mapping from physical features to semantic features, and the other set of mapping from semantic features to physical features. Each set of mapping is realized through a multi-head attention mechanism, allowing fine-grained information exchange between different features. This bidirectional connection is not simply a feature concatenation, but a dynamic, content-based selective information transfer, ensuring that physical analysis focuses on semantic-related areas, while semantic understanding is guided by physical constraints. It should be emphasized that the "double helix" structure of the present application is an innovative information flow design, which is different from traditional unidirectional or parallel processing architecture.
[0053] The realization of inter-layer interaction connection can be divided into the following core parts in detail: Bidirectional cross-attention mechanism: This mechanism is the core of the "double helix" structure, realizing the deep interaction between physical and semantic features. The attention module from physics to semantics and the attention module from semantics to physics together constitute a complete interaction system. Each attention module adopts a multi-head attention structure, and the specific calculation is as follows: Physical-to-semantic attention computation: ; ; ; ; wherein, denotes the query matrix of semantic features; denotes the semantic query weight matrix; denotes the semantic feature vector; denotes the key matrix of physical features; denotes the physical key weight matrix; denotes the physical feature vector; denotes the value matrix of physical features; denotes the physical value weight matrix; denotes the physical-to-semantic attention output; denotes the softmax normalization operation; denotes the dimension of the key vector in the physical-to-semantic attention module.
[0054] Semantic-to-physical attention computation: ; wherein, denotes the query matrix of physical features; denotes the physical query weight matrix; denotes the physical feature vector; denotes the key matrix of semantic features; denotes the semantic key weight matrix; denotes the semantic feature vector; denotes the value matrix of semantic features; denotes the semantic value weight matrix; denotes the semantic-to-physical attention output; denotes the softmax normalization operation; denotes the dimension of the key vector in the semantic-to-physical attention module.
[0055] In actual implementation, each attention module uses 8 heads, each with a dimension of 32, to capture feature dependency relationships of different dimensions. Optionally, in some embodiments, the number of attention heads can be dynamically adjusted according to the complexity of the input image to balance computational efficiency and expressiveness.
[0056] Iterative refinement layer: To further enhance the depth of fusion of physical and semantic information, an iterative refinement layer is introduced to realize multiple rounds of information exchange. In the first round of iteration, the information update formula is: ; where, denotes the physical feature vector after the i-th iteration; denotes the physical feature vector after the i-th iteration; denotes the learning rate parameter, controlling the information update speed; denotes the feed-forward neural network (two fully connected layers + ReLU) of the physical feature; denotes the semantic-to-physical attention output after the i-th iteration; denotes the semantic feature vector after the i-th iteration; denotes the semantic feature vector after the i-th iteration; denotes the feed-forward neural network of the semantic feature; denotes the physical-to-semantic attention output after the i-th iteration. Feature enhancement gating unit: To avoid the interference of irrelevant information, a feature enhancement gating unit is adopted to control the strength of information interaction. The gating mechanism dynamically adjusts the influence degree of cross-attention output on the original feature through learnable parameters: ; ;
[0057] ; ; ; ; ; where, denotes the physical feature gating coefficient; denotes the sigmoid activation function; denotes the physical gating weight matrix; denotes the concatenation of the physical feature and the semantic-to-physical attention output; denotes the physical gating bias; denotes the enhanced physical feature; denotes element-wise multiplication; denotes the semantic-to-physical attention output; denotes the original physical feature; denotes the semantic feature gating coefficient; denotes the semantic gating weight matrix; denotes the concatenation of the semantic feature and the physical-to-semantic attention output; denotes the semantic gating bias; denotes the enhanced semantic feature; denotes the physical-to-semantic attention output; denotes the original semantic feature.
[0058] Multi-scale interaction strategy: To capture the physical-semantic associations at different abstraction levels, interaction connections are established at multiple levels of the network. Specifically, cross connections are established between the low-level, mid-level, and high-level features of the physical parameter extraction network and the semantic understanding network, forming a multi-level "double helix" structure: Low-level connection: focusing on the association between local texture and basic semantic elements; Mid-level connection: focusing on the association between object parts and regional semantics; High-level connection: focusing on the association between global physical parameters and scene semantics.
[0059] Different levels of connection use similar attention mechanisms, but the parameters are not shared to adapt to the characteristics of features at different abstraction levels. Multi-scale interaction greatly enriches the way of fusing physical and semantic information, enabling the system to capture information associations at both microscopic and macroscopic levels. Therefore, this design improves the system's ability to handle complex scenes, especially for images that contain multiple physical phenomena and rich semantic content.
[0060] Attention visualization and explanation module: To enhance the explainability of the system, an attention visualization module is designed to display the attention relationship between physical and semantic features in real time. This module maps the attention weights back to the original image space to generate heat maps showing the dependence strength between different features. The visualization of attention weights not only helps to understand the decision-making process of the system, but also intuitively shows the patterns of physical-semantic associations, providing important clues for subsequent analysis. In specific implementation, the system remaps the attention weights at the feature level to the pixel level through upsampling and bilinear interpolation, and uses color coding to represent the dependence strength, making it easier for humans to understand. Optionally, attention visualization can also be used as an intermediate debugging tool to help developers optimize system parameters and identify potential problems.
[0061] In summary, the inter-layer interaction connection mechanism realizes the deep fusion of physical and semantic information through the above carefully designed components, forming a true "double helix" structure. This structure not only enables physical analysis and semantic understanding to guide and enhance each other, but also treats conflicts as important discriminative signals when they are discovered, which is one of the core innovations of the present application that distinguishes it from existing technologies.
[0062] Step 2, use the physical parameter extraction network layer to analyze the input image and generate a set of physical feature vectors; Step 2.1, apply multi-scale convolution filters to the input image to extract image features at different scales; In an embodiment according to the present application, the input image is processed using five different sizes of convolution kernels (3x3, 5x5, 7x7, 9x9, 11x11), each of which contains 64 filters for capturing features at different levels from detailed textures to macro structures. Optionally, the filter parameters can be automatically adjusted according to the image resolution to ensure that features of appropriate granularity can be extracted at different resolutions. It should be noted that in other embodiments, dilated convolution techniques can also be used to obtain different receptive fields by adjusting the dilation rate without increasing the number of parameters or reducing the resolution.
[0063] Step 2.2, fuse multi-scale features through a feature pyramid structure to form a comprehensive feature map; This process adopts a bidirectional feature fusion strategy from top to bottom and bottom to top to ensure effective integration of information at different scales. Specifically, a bottom-up feature pyramid is first constructed by downsampling, and a top-down feature pyramid is then constructed by upsampling and lateral connection, finally generating a comprehensive feature map containing multi-scale information. It should be understood that this structure can both preserve detailed information and capture global semantics. In addition, in some embodiments, an attention mechanism can be introduced to assist the fusion process of the feature pyramid, adaptively adjusting the importance weights of features at different scales to further improve the fusion effect.
[0064] Step 2.3, apply a dedicated physical parameter estimation module to infer physical parameters from the comprehensive feature map; The physical parameters include illumination model parameters, surface reflection characteristics, shadow consistency, and perspective relationships, etc. According to an embodiment of the present application, the module contains multiple parallel sub-networks, each of which is responsible for the estimation of a specific physical parameter. For example, the illumination model sub-network outputs parameters such as light source direction, intensity, and color temperature; the reflection characteristic sub-network estimates the diffuse and specular reflection properties of the material; the perspective relationship sub-network infers the camera intrinsic and extrinsic parameters. In addition, in some embodiments, a physical consistency verification sub-network can also be included to check whether the basic physical constraint relationships between different physical parameters are satisfied. Optionally, these sub-networks can adopt a multi-task learning framework, sharing the bottom feature extraction layer, not only reducing the number of parameters, but also improving the overall estimation accuracy through knowledge transfer between tasks.
[0065] Step 2.4, calculate the uncertainty estimation of the physical parameters to provide confidence information for subsequent probabilistic reasoning; In a preferred embodiment according to the present application, the uncertainty estimation is achieved through the Monte Carlo Dropout technique, i.e. keeping the Dropout layer activations during the inference phase, performing multiple forward propagations on the same input, and statistically quantifying the output variability as the uncertainty measure. It can be seen that this uncertainty estimation method does not require additional training process and can be directly integrated into existing networks. In addition, in other embodiments, ensemble learning methods or Bayesian neural networks can also be used to obtain uncertainty estimation, although these methods may require more computational resources, but in some specific scenarios can provide more accurate uncertainty quantification.
[0066] Step 3, applying the semantic understanding network layer to process the input image, generating a semantic feature vector and a relationship graph; Step 3.1, using a pre-trained visual Transformer model to extract semantic features of the image; According to an embodiment of the present application, the model is pre-trained on a large-scale image dataset and has strong semantic understanding ability. In a preferred embodiment, the ViT-Base variant is used, which contains 12 Transformer blocks, each with 12 attention heads, a hidden layer dimension of 768, and a total parameter amount of about 86 million. It should be noted that for different application scenarios, smaller ViT-Small or larger ViT-Large variants can be selected to balance performance and computational resource requirements. In addition, in environments with limited computing resources, lightweight Transformer architectures such as MobileViT can also be considered, or knowledge distillation techniques can be used to transfer the knowledge of large models to small models, reducing the computational burden while maintaining performance.
[0067] Step 3.2, applying an object detection and segmentation module to identify key objects and regions in the image; According to an embodiment of the present application, the module is based on the MaskR-CNN architecture, but replaces the backbone network with a Transformer to enhance feature extraction capability. The processing flow includes four stages of generating region proposals, object classification, bounding box regression, and instance segmentation, finally outputting the class, position and pixel-level mask of each identified object. Optionally, in some embodiments, a keypoint detection branch can also be added for more fine-grained object structure analysis. In addition, for scenarios that require real-time processing, single-stage detectors such as YOLOX or DETR can be used, which reduce computational complexity and speed up processing through end-to-end design.
[0068] Step 3.3, constructing an object relationship graph; The object relation graph represents the spatial and functional relations between objects in the image. In embodiments according to the present application, the relation graph is represented in a graph structure, where nodes correspond to detected objects and edges represent the relations between objects. Relation types include spatial relations (e.g. "above", "inside"), functional relations (e.g. "use", "contain") and semantic relations (e.g. "belong to the same category"). The graph is constructed by computing the similarity and positional relation between object feature vectors. In addition, to enhance the relation reasoning capability, an external knowledge graph can be introduced as a source of prior knowledge to guide the relation prediction process. It should be understood that in other possible implementations, a global relation modeling method based on attention mechanism can also be used, without explicitly constructing a graph structure, directly capturing the interaction between objects through a self-attention mechanism.
[0069] Step 3.4, applying event reasoning module to analyze the events that can occur in the scene and their rationality; According to a preferred embodiment of the present application, the module uses a graph convolution network (GCN) to process the object relation graph, and infers the interaction between objects and the events that can occur in the scene through a message passing mechanism. The event rationality evaluation is based on a pre-trained event knowledge base, which contains the event types that can be formed by common object combinations and their occurrence probabilities. It should be emphasized that the event rationality evaluation is an important innovation point of the present application, which goes beyond simple object recognition and relation understanding, and realizes higher level semantic understanding of the scene. Optionally, in some embodiments, a dynamic event reasoning model can also be constructed by combining the timing information (if applicable), not only considering the event rationality of the static scene, but also evaluating the naturalness of the event development sequence.
[0070] Step 4, based on the physical feature vector and the semantic feature vector, a physical-semantic joint distribution is constructed using a variational inference engine layer; Step 4.1, input the physical feature vector and the semantic feature vector into the variational autoencoder to map them to a shared latent space; According to embodiments of the present application, the variational autoencoder is composed of an encoder and a decoder, the encoder maps high-dimensional features into mean and variance parameters of low-dimensional latent variables, and the decoder reconstructs the original features from the latent variables. In this embodiment, the dimension of the latent space is set to 128, and both the encoder and the decoder use a multi-layer fully connected network. Optionally, a convolutional variational autoencoder structure can also be used to better preserve the spatial information of the features. In addition, in another embodiment, a hierarchical variational autoencoder (HVAE) architecture can be introduced to more effectively capture the feature relations at different abstraction levels by constructing a multi-level latent variable structure.
[0071] Step 4.2, applying conditional probability modeling in latent space, calculating physical compliance probability and semantic consistency probability; According to an embodiment of the present application, the conditional probability model adopts a conditional variational autoencoder (CVAE) structure, which allows the inference of probability distribution given specific conditions. The physical compliance probability is calculated by comparing the consistency of actual physical parameters and expected physical parameters; the semantic consistency probability is obtained by evaluating the degree of conformity of the semantic relationship of the image content with common sense knowledge. It should be understood that this double-checking mechanism can comprehensively evaluate the authenticity of the image from both the physical and semantic dimensions. Optionally, in some embodiments, an adversarial training strategy can also be introduced to enhance the expressiveness and robustness of the conditional probability model through a generator-discriminator structure.
[0072] Step 4.3, integrating the conditional probability into a joint probability distribution through a Bayesian inference framework; The mathematical expression is: ; where represents the posterior probability of the image being a real image; represents the integral over the physical parameter space ; represents the physical compliance probability of the image under the physical parameter ; represents the semantic consistency probability of the image under the physical parameter ; represents the prior distribution of the physical parameter ; represents the integral variable over the physical parameter .
[0073] This integral expression represents the way in which physical compliance and semantic consistency jointly affect the authenticity of the image over all possible physical parameter spaces. In addition, in practical applications, the weights of physical compliance and semantic consistency can be adjusted according to different scenarios to adapt to different types of images and forgery methods. For example, for landscape images, more emphasis can be placed on physical compliance verification; while for images of people or events, more attention can be paid to semantic consistency evaluation.
[0074] Step 4.4, applying the Monte Carlo integration method to approximate the integral, obtaining the final authenticity probability estimate; According to an embodiment of the present application, the approximate calculation is performed using the Monte Carlo method. Specifically, a plurality of parameter points are sampled from the prior distribution Then the function value at each sample point is calculated , and finally the average value is taken as the approximate result of the integral. In actual implementation, 1000 is taken to balance the calculation efficiency and the approximation accuracy. Alternatively, in an environment with limited calculation resources, a more efficient Monte Carlo variant method such as importance sampling can be used to reduce the number of required sample points. In addition, in other possible implementations, a variational inference technique can also be used to directly optimize the approximate expression of the posterior distribution, avoiding explicit integral calculation and further improving the calculation efficiency.
[0075] Step 5, a physical semantic cross-attention mechanism is implemented to enhance the bidirectional information flow between physical features and semantic features; Step 5.1, a physical-to-semantic attention mapping is constructed to guide the semantic understanding to focus on the physically abnormal regions; According to the preferred embodiment of the present application, the mapping generates an attention weight map by calculating the anomaly score of each position in the physical feature map, and then applies this weight map to the semantic features, so that the semantic analysis process pays more attention to the physically suspicious regions. In specific implementation, the anomaly score is calculated by comparing the deviation of the actual physical features from the expected physical features. In addition, in order to enhance the stability of the attention mechanism, a residual connection can be introduced to retain the original feature information while enhancing the representation of key regions. Another optional implementation is to use an adaptive threshold strategy to dynamically adjust the screening standard of abnormal regions according to the complexity and quality of the image, improving the adaptability of the mechanism to different types of images.
[0076] Step 5.2, a semantic-to-physical attention mapping is constructed to make the physical analysis preferentially process the semantically important regions; In an embodiment of the present application, the mapping generates attention weights for physical features based on the importance scores of semantic features, ensuring that the physical analysis process focuses on semantically key areas (such as main objects, interaction areas, etc.). The importance scores are calculated through object classification confidence, node centrality in the relationship graph, etc. It should be emphasized that this bidirectional attention mapping is one of the core innovations of the present application, which realizes the mutual guidance of physical and semantic analysis, forming a truly "double helix" information flow structure. Optionally, in some embodiments, the functional nature of the objects and the importance of the roles in the events can also be combined to assign different attention weights to different semantic entities, making the physical analysis more focused on the key elements in the scene.
[0077] Step 5.3, deep fusion of physical and semantic features is achieved by alternating iterative optimization of the two attention mappings; According to an embodiment of the present application, the process adopts an iterative refinement strategy, and in each iteration, the two attention mappings influence and update each other, gradually improving the quality of information interaction. Specifically, the physical-semantic attention mapping of the first round is influenced by the semantic-physical attention mapping of the first round, and vice versa. The number of iterations is set to 3, balancing accuracy and efficiency. Optionally, in some embodiments, the number of iterations can be dynamically adjusted to adaptively determine the depth of information exchange according to the complexity of the image. In addition, an early stopping mechanism can also be introduced to terminate the iteration in advance when the change in attention distribution is less than a preset threshold, further improving the computational efficiency.
[0078] Step 5.4, extract the fused feature representation for subsequent anomaly detection; In an embodiment according to the present application, the fused features are obtained by nonlinear combination of the attention-weighted physical features and semantic features. In specific implementation, a gated fusion unit is used, which includes a control gate that dynamically adjusts the contribution weights of different features according to the relevance of the input features, ensuring the flexibility and adaptability of the information flow. Therefore, this fusion mechanism can automatically adjust the importance of physical and semantic features according to the characteristics of different images, improving the versatility and robustness of the system. It should be understood that in other embodiments, a multi-head attention mechanism can also be used to achieve more fine-grained feature fusion, or a feature conversion module can be introduced to enhance the expressiveness of the features through nonlinear transformation.
[0079] Step 6, analyze the deviation of the physical-semantic joint distribution from the pre-established standard distribution of natural images; Step 6.1, establish a standard distribution model based on large-scale real image data as a reference benchmark; According to a preferred embodiment of the present application, the model is obtained by training on a dataset containing 1 million real images, covering diverse samples of different scenes, lighting conditions and shooting devices. The training process adopts an unsupervised learning method to learn the physical-semantic joint distribution features of real images through a variational autoencoder. In addition, in order to improve the representativeness of the model, images of different quality levels and sources are included in the dataset to ensure that the standard distribution can cover the diversity of real images. It should be noted that in some embodiments, specific standard distribution models can also be constructed according to specific application fields, such as news picture standard models, artistic work standard models, etc., to improve the identification accuracy for specific types of images.
[0080] Step 6.2, calculate the Kullback-Leibler divergence between the joint distribution of the test image and the standard distribution, and quantify the degree of deviation; In an embodiment of the present application, the Kullback-Leibler divergence is an effective tool for measuring the difference between two probability distributions, and its calculation formula is: wherein represents the Kullback-Leibler divergence of the distribution and ; represents the summation operator; represents the probability value of the test image in the dimension; represents the natural logarithm function; represents the probability value of the standard distribution in the dimension; represents the dimension number of the distribution.
[0081] In specific implementations, due to the complexity of the distribution form, the Monte Carlo method is used to approximate the divergence value. Alternatively, in some embodiments, other distribution measurement methods such as Jensen-Shannon divergence or Wasserstein distance can also be used to adapt to different types of distribution difference features. In addition, for environments with limited computing resources, low-rank approximation or feature dimension reduction techniques can be used to reduce the computational complexity while maintaining the effectiveness of the measurement.
[0082] As Figure 2 The Kullback-Leibler divergence of the image is shown in relation to the true image probability. The horizontal axis represents the KL divergence between the image distribution and the standard distribution, and the vertical axis represents the true image probability determined by the system. It can be seen that as the KL divergence increases, the probability of the image being determined as true shows a clear downward trend, and the two show a strong negative correlation. This relationship verifies the effectiveness of using KL divergence as a basis for image authenticity determination, provides a theoretical basis for threshold setting, and shows that the divergence calculation method can accurately reflect the degree of deviation of the image from the natural distribution.
[0083] Step 6.3, set adaptive threshold, determine the authenticity of the image based on the degree of deviation; According to the embodiments of the present application, the threshold setting adopts a dynamic strategy, which is adjusted according to the content complexity, quality and application scenario of the image.
[0084] Specifically, the threshold T is composed of a basic threshold and a scene adjustment factor : ; wherein represents the adaptive threshold finally used for judgment; represents the basic threshold, which is determined by grid search optimization on a verification set containing 50,000 labeled images, and the default value is 0.65; represents the scene adjustment factor, which is dynamically assigned by the scene classification module according to the image type, and the value range is 0.8-1.2. A higher value is set for complex scenes (such as crowd gathering, multi-object interaction), and a lower value is set for simple scenes (such as landscape, still life).
[0085] In practical applications, adjustable threshold options can also be provided according to different user requirements for accuracy and recall rate, for example, a stricter threshold can be set in news media applications to reduce the misjudgment rate, and a more lenient threshold can be set in rapid screening applications to improve the detection rate.
[0086] It should be understood that in other embodiments, a multi-threshold strategy can also be used, setting three thresholds , and the judgment result is divided into: determined true: KL divergence < T ; possibly true: ≤ KL divergence < T ; possibly fake: ≤ KL divergence < T ; determined fake: KL divergence ≥ T ; Provide more granular identification results, where 、 、 respectively represent the first, second, and third judgment thresholds.
[0087] As shown in Figure 3 , the detection accuracy (bar chart) and false positive rate (line chart) of the system under different threshold settings are shown. As the threshold increases from 0.1 to 1.0, the detection accuracy gradually decreases, while the false positive rate increases. This chart can help users find the best balance point between accuracy and false positive rate according to the needs of specific application scenarios. For example, in news media applications with low false positive rate requirements, a higher threshold (such as 0.8-0.9) can be selected; while in preliminary screening applications requiring high detection rate, a lower threshold (such as 0.3-0.5) can be selected, providing important parameter selection guidance for adaptive threshold setting.
[0088] Step 6.4, locate the key area causing deviation, provide visual abnormal heat map; In embodiments according to the present application, the heat map generation is based on feature importance analysis, using the gradient weighted class activation mapping (Grad-CAM++) technique to calculate the contribution of each pixel to the final decision. This technique calculates the gradient of the loss function with respect to the feature map through backpropagation, then uses these gradients as weights to perform weighted summation on the feature map, and finally obtains a heat map highlighting the abnormal area. It should be noted that visualizing the abnormal heat map not only helps users understand the identification results, but also provides important clues for subsequent forensics and analysis. Optionally, in some embodiments, saliency detection techniques can also be combined to further enhance the visual prominence of abnormal areas, or multi-level heat map fusion strategies can be used to simultaneously display abnormal features at different semantic levels.
[0089] Step 7, based on the degree of deviation, generate the identification results and explainability analysis report of the input image; Step 7.1, generate a probability score of image authenticity to quantify the identification confidence; According to the preferred embodiments of the present application, the probability score is based on the posterior probability calculated in step 4 and is calibrated to ensure the consistency and interpretability of the score. Calibration uses temperature scaling method to adjust the kurtosis of the probability distribution, making the score more reliable and intuitive. In addition, to facilitate user understanding, the system converts the probability score to a percentage form of 0-100 and provides corresponding confidence level descriptions. It should be noted that in other embodiments, integrated scoring strategies can also be used to combine the prediction results of multiple models to enhance the stability and reliability of the score, especially for the handling of boundary conditions.
[0090] Step 7.2, identify potential abnormal regions and mark on the original image; In an embodiment according to the present application, the marking process is based on the abnormal heat map generated in step 6.4. The heat map is converted into an abnormal region mask by thresholding, and then the mask is superimposed on the original image to mark the abnormal intensity with different colors. Optionally, the system can also automatically extract the boundary contour of the abnormal region and sort the regions according to the abnormal degree, highlighting the most suspicious part. In addition, in some embodiments, different color coding can also be used for different types of abnormalities (such as physical abnormalities and semantic abnormalities), making the marking result more informative and helping users quickly understand the nature of the abnormality.
[0091] Step 7.3, provide abnormal reason analysis; According to an embodiment of the present application, the abnormal reason analysis module interprets the abnormal region from two dimensions of physics and semantics. The physical dimension analysis includes inconsistent lighting, shadow error, perspective distortion, etc.; the semantic dimension analysis includes object relationship contradiction, unreasonable event, common sense violation, etc. The analysis results are output in the form of structured text, providing intuitive understanding for professionals and ordinary users. In addition, the system will automatically sort the abnormal reasons according to the severity, and prioritize the problems that have the greatest impact on the authenticity of the image. It should be understood that in other embodiments, an interactive abnormal analysis interface can also be introduced, allowing users to obtain more detailed analysis by clicking on a specific region, or to explore the results under different abnormal detection thresholds by adjusting parameters.
[0092] Step 7.4, generate a detailed identification report; In a preferred embodiment according to the present application, the report contains the following contents: image basic information, authenticity probability score, abnormal region visualization, physical analysis result, semantic analysis result, comprehensive evaluation conclusion and suggestion. The report adopts a hierarchical structure, and different users can view information of different detailed levels according to their needs. In addition, in practical applications, customized report templates can be provided according to user identity and application scenarios to meet the needs of professionals in different fields. Optionally, the report generation system can also support multiple output formats such as PDF, HTML or JSON to adapt to different use scenarios; at the same time, for applications that need batch processing, summary reports can also be generated to display the identification result statistics and abnormal pattern analysis of multiple images.
[0093] As Figure 4As shown, the accuracy of the physical-semantic double helix theory compared with traditional physical verification and traditional semantic inspection methods in identifying ordinary GAN images and high-quality GAN images is illustrated. The results show that the physical-semantic double helix theory maintains a high accuracy of over 90% on both types of images, while the accuracy of traditional methods decreases when dealing with high-quality GAN images. This comparison clearly demonstrates the advantages of the proposed method in handling complex AI-generated images, especially its stable performance when dealing with high-quality forged images.
[0094] like Figure 5 As shown, the identification accuracy of the physical-semantic double helix theory and traditional methods on different types of images (news images, portrait images, landscape images, indoor scenes, and complex events) is demonstrated. The physical-semantic double helix theory maintains high accuracy across all image types, exhibiting good versatility, while traditional methods perform worst on complex event images. This result proves that the proposed method has excellent cross-domain adaptability and stability, and can provide reliable image identification services for different application scenarios.
[0095] like Figure 6 As shown, the overall performance of the physical-semantic double helix theory and traditional methods is compared across six dimensions: accuracy, robustness, interpretability, adaptability, computational efficiency, and generality. The physical-semantic double helix theory outperforms traditional methods in all dimensions except computational efficiency, with its advantages being most pronounced in interpretability and adaptability. This comprehensive performance comparison clearly demonstrates the full technical advantages of the proposed method, providing strong performance support for its deployment and promotion in practical applications.
[0096] An artificial intelligence-based digital image authentication system, used to execute the aforementioned artificial intelligence-based digital image authentication method, includes: The three-layer fusion neural network architecture module is used to build a neural network architecture that includes a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer. The physics analysis module is used to extract the input image from the network layer analysis using physical parameters and generate a set of physical feature vectors. The semantic understanding module is used to process the input image using the semantic understanding network layer, generating semantic feature vectors and relationship graphs. The cross-attention module is used to implement the physical-semantic cross-attention mechanism to enhance the bidirectional information flow between physical features and semantic features; The joint distribution construction module is used to construct a physical-semantic joint distribution based on physical feature vectors and semantic feature vectors using the variational inference engine layer. The distribution analysis module is used to analyze the degree of deviation between the physical and semantic joint distribution of the input image and the pre-established standard distribution of natural images; a result output module configured to generate an identification result and an explainability analysis report of the input image based on the deviation degree.
[0097] The above describes the embodiments of the present application, but the embodiments are not limited to the specific embodiments described above, which are only illustrative but not restrictive. Those skilled in the art can make more equivalent embodiments under the inspiration of the embodiments, which are all within the protection scope of the embodiments.
Claims
1. An artificial intelligence-based digital image authentication method, characterized by, The application relates to a method for generating an authenticity verification report of an input image, comprising the following steps: A three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer and a variational inference engine layer; The physical parameter extraction network layer is used to analyze the input image and generate a set of physical feature vectors; The semantic understanding network layer is used to process the input image and generate semantic feature vectors and a relationship graph; Based on the physical feature vectors and the semantic feature vectors, the variational inference engine layer is used to construct a physical-semantic joint distribution; A physical-semantic cross-attention mechanism is implemented to enhance the bidirectional information flow between the physical features and the semantic features; The deviation degree of the physical-semantic joint distribution from a pre-established standard distribution of natural images is analyzed; Based on the deviation degree, an authenticity verification result and an explainability analysis report of the input image are generated.
2. The method of claim 1, wherein the method is based on artificial intelligence. The construction method of the physical parameter extraction network layer comprises the following steps: A deep convolutional neural network structure is adopted, including a feature extraction module and a physical parameter estimation module; Multi-scale convolution filters are applied to the input image to extract image features of different scales; Multi-scale features are fused through a feature pyramid structure to form a comprehensive feature map; A special physical parameter estimation module is applied to infer the illumination model parameters, surface reflection characteristics, shadow consistency and perspective relationship from the comprehensive feature map; The uncertainty estimation of the physical parameters is calculated to provide confidence information for subsequent probabilistic inference.
3. The method of claim 1, wherein the method further comprises: The construction method of the semantic understanding network layer comprises the following steps: A visual model based on Transformer is used to extract the semantic features of the input image; An object detection and segmentation module is applied to identify key objects and regions in the input image; An object relationship graph is constructed to represent the spatial and functional relationships between the objects in the input image; An event inference module is applied to analyze the events in the scene and their rationality.
4. The method of claim 1, wherein the method is based on artificial intelligence. The construction method of the variational inference engine layer comprises the following steps: The physical feature vectors and the semantic feature vectors are input into a variational autoencoder to be mapped to a shared latent space; Conditional probability modeling is applied in the latent space to calculate the physical compliance probability and the semantic consistency probability; The conditional probability is integrated into a joint probability distribution through a Bayesian inference framework; The authenticity probability estimate is obtained by applying the Monte Carlo integration method to calculate the integral.
5. The method of claim 1, wherein the method is based on artificial intelligence. The steps of implementing the physical-semantic cross-attention mechanism comprise the following steps: A physical-to-semantic attention mapping is constructed to guide the semantic understanding to focus on the abnormal areas of the physical features; A semantic-to-physical attention mapping is constructed to make the physical analysis preferentially process the important areas of the semantic features; The two attention mappings are alternately optimized to realize the deep fusion of the physical and semantic features; The fused feature representation is extracted for subsequent anomaly detection.
6. The method of claim 1, wherein the method further comprises: The steps of analyzing the deviation degree of the physical-semantic joint distribution from the pre-established standard distribution of natural images comprise the following steps: A standard distribution model based on large-scale real image data is established as a reference benchmark; The Kullbac Leibler divergence between the joint distribution of the input image and the standard distribution is calculated to quantify the deviation degree; An adaptive threshold is set to judge the authenticity of the input image based on the deviation degree; The key areas causing the deviation are located to provide a visual abnormality heat map.
7. The method of claim 1, wherein the method further comprises: The steps of generating the authenticity verification result and the explainability analysis report of the input image comprise the following steps: Generate a probability score of the input image authenticity, quantify the identification confidence; Identify potential abnormal areas and mark them on the input image; Provide abnormal reason analysis, explain the violation of physical laws and semantic logic; Generate a detailed identification report, including various indicators and comprehensive evaluation results.
8. The method of claim 4, wherein the method further comprises: The step of integrating conditional probabilities into joint probability distribution through Bayesian inference framework specifically includes: Temperature scaling calibration of physical semantic joint distribution, adjusting temperature parameters to improve the reliability of probability estimation of calibrated probability distribution; Using variational autoencoder to establish multi-scale interaction strategy, realizing the interaction of physical feature vectors and semantic feature vectors at different abstraction levels through cross-layer connection; Applying feature enhancement gate unit to control the strength of information interaction between physical feature vectors and semantic feature vectors. 9.The method of claim 1, wherein, The collaborative working mechanism of the physical parameter extraction network layer and the semantic understanding network layer specifically includes: Pre-classify the input image through the image type classifier, and automatically adjust the parameter configuration of the multi-scale convolution filter and the attention weight of the Transformer-based visual model according to the classification result; Build incremental learning modules in the physical parameter extraction network layer and the semantic understanding network layer, so that the deep convolutional neural network structure and the Transformer-based visual model can continuously learn from new forgery technology samples and dynamically update the network weights; Establish a multi-modal fusion mechanism to fuse the lighting model parameters and surface reflection characteristics output by the physical parameter estimation module with the key object information identified by the object detection and segmentation module, and simultaneously integrate image metadata for comprehensive analysis.
10. An artificial intelligence-based digital image authentication system, characterized by, An artificial intelligence-based digital image identification method for performing any one of claims 1-9, comprising: A three-layer fusion neural network architecture module for building a neural network architecture including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer; A physical analysis module for analyzing the input image using the physical parameter extraction network layer to generate a set of physical feature vectors; A semantic understanding module for processing the input image using the semantic understanding network layer to generate semantic feature vectors and relationship graphs; A cross-attention module for implementing a physical-semantic cross-attention mechanism to enhance bidirectional information flow between physical features and semantic features; A joint distribution construction module for constructing a physical-semantic joint distribution based on physical feature vectors and semantic feature vectors using the variational inference engine layer; A distribution analysis module for analyzing the deviation of the physical-semantic joint distribution of the input image from the pre-established natural image standard distribution; A result output module for generating identification results and explainability analysis reports of the input image based on the deviation.
Citation Information
Patent Citations
Forgery image detection method, electronic equipment and computer readable storage medium
CN115984178A
RGB-T image semantic segmentation method and device
CN116091765A
Generative artificial intelligence polymorphic sensitive puzzle detection method, device and equipment
CN118587516A
Multi-image forgery detection method and system based on cross-modal visual large language model
CN120355985A
Intelligent fruit quality grading system and method based on multi-mode variational auto-encoder
CN120375358A
Cited By
Posterior probability calibration method of semantic segmentation model based on multi-scale temperature scaling
CN121544890A
Unmanned aerial vehicle low-altitude route drawing method, equipment and medium
CN121877013A
Detection method and device of AI generated image, electronic equipment and storage medium
CN122156176A