A digital image identification method and system based on artificial intelligence
By constructing a three-layer fusion neural network architecture, combining physical parameter extraction and semantic understanding, the problems of low accuracy and insufficient interpretability in image identification in existing technologies are solved, and efficient, stable identification and detailed analysis of AI-generated images are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-03-13
AI Technical Summary
Existing image identification technologies cannot effectively combine physical law verification and semantic logic analysis, resulting in low recognition accuracy, lack of interpretability and adaptability of AI-generated images, making it difficult to cope with rapidly evolving image generation technologies.
A three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer. Through a physical-semantic cross-attention mechanism and a Bayesian inference framework, joint physical and semantic verification and interpretability analysis of images are achieved.
It improves the accuracy of identifying AI-generated images, can stably detect physical and semantic conflict areas of forged images in high-quality generated images, provides detailed analysis of the causes of anomalies, adapts to different forgery techniques, reduces computational complexity, and achieves real-time processing and good stability.
Smart Images

Figure CN120997645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital image authentication technology, and more specifically, to a digital image authentication method and system based on artificial intelligence. Background Technology
[0002] Modern AI-generated images have reached a level of visual quality that is almost indistinguishable from real photographs, posing a serious challenge to traditional image authentication techniques.
[0003] Currently, existing technologies in the field of image authentication can be mainly divided into the following categories:
[0004] Traditional methods based on pixel-level statistical features: Early image authentication techniques mainly relied on techniques such as JPEG compression trace analysis, pixel value statistical characteristic detection, and frequency domain feature analysis. For example, image tampering was identified by detecting double JPEG compression traces, CFA (Color Filter Array) interpolation traces, or statistical distribution anomalies in pixel values. However, these methods have limitations: First, modern generative adversarial networks can generate highly realistic images at the pixel level, making traditional statistical feature detection methods prone to failure; second, high-quality image post-processing techniques can effectively eliminate these statistical traces, making detection even more difficult.
[0005] Physically-based authentication methods determine image authenticity by analyzing physical features such as illumination consistency, shadow direction, and perspective relationships. Typical techniques include light source direction estimation, shadow geometry verification, and specular reflection analysis. While these methods are effective in detecting image tampering that clearly violates physical laws, they face the following challenges: current advanced image generation techniques can simulate basic physical laws to a certain extent, generating images that conform to illumination and geometric constraints; furthermore, these methods typically focus only on local physical features and lack a comprehensive assessment of global physical consistency.
[0006] Semantic consistency-based detection methods identify forged images by analyzing the semantic logic of image content, including object relationship detection, scene plausibility analysis, and spatiotemporal consistency verification. For example, they detect unreasonable object combinations, illogical scene layouts, or temporal logical contradictions within an image. However, these methods also have limitations: with the development of large-scale language models and multimodal AI technologies, modern image generation systems can produce image content that is highly semantically plausible; simultaneously, purely semantic methods often neglect the physical reality of the image, failing to detect images that are logically plausible but physically impossible.
[0007] Existing technologies generally suffer from the following fundamental problems:
[0008] Various detection methods typically operate independently, lacking a unified theoretical framework to organically combine physical verification and semantic understanding, thus failing to fully leverage the complementary advantages of the two dimensions. Most existing methods employ deterministic binary classification judgments, lacking probabilistic reasoning mechanisms and failing to quantify the uncertainty of identification results. Existing technologies generally lack interpretability, failing to provide intuitive and understandable explanations for identification results, making it difficult to gain user trust and professional recognition. Faced with rapidly evolving image generation technologies, existing methods suffer from insufficient generalization and adaptability, requiring frequent updates to training data and model parameters.
[0009] These technical limitations indicate an urgent need to develop a new image identification theory and method that can unify physical law verification and semantic logic analysis to address the increasingly complex challenges of AI-generated image detection. Summary of the Invention
[0010] This invention provides a digital image identification method and system based on artificial intelligence, which solves the technical problems in related technologies such as low accuracy of AI-generated image recognition, lack of a unified physical and semantic verification framework, insufficient interpretability of detection results, and weak adaptability to new forgery technologies.
[0011] This invention provides a digital image identification method based on artificial intelligence, comprising:
[0012] A three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer;
[0013] The physical parameters are used to extract network layer analysis of the input image and generate a set of physical feature vectors.
[0014] The semantic understanding network layer is applied to process the input image, generating semantic feature vectors and relationship graphs;
[0015] Based on physical feature vectors and semantic feature vectors, a joint physical-semantic distribution is constructed using a variational inference engine layer;
[0016] Implement a physical-semantic cross-attention mechanism to enhance bidirectional information flow between physical and semantic features;
[0017] Analyze the degree of deviation between the joint distribution of physical and semantic information and the pre-established standard distribution of natural images;
[0018] Based on the degree of deviation, an identification result and interpretability analysis report of the input image are generated.
[0019] Furthermore, the method for constructing the physical parameter extraction network layer includes:
[0020] It adopts a deep convolutional neural network structure, which includes a feature extraction module and a physical parameter estimation module;
[0021] Multi-scale convolutional filters are applied to the input image to extract image features at different scales;
[0022] Multi-scale features are fused using a feature pyramid structure to form a comprehensive feature map;
[0023] A dedicated physical parameter estimation module is used to infer lighting model parameters, surface reflection characteristics, shadow consistency, and perspective relationships from the comprehensive feature map.
[0024] Calculate the uncertainty estimate of the physical parameters to provide confidence information for subsequent probabilistic inference.
[0025] Furthermore, the method for constructing the semantic understanding network layer includes:
[0026] A Transformer-based visual model is used to extract semantic features from the input image;
[0027] The application object detection and segmentation module identifies key objects and regions in the input image;
[0028] Construct an object relationship graph to represent the spatial and functional relationships between objects in the input image;
[0029] The event reasoning module is used to analyze the events in the scenario and their rationality.
[0030] Furthermore, the method for constructing the variational inference engine layer includes:
[0031] The physical feature vector and semantic feature vector are input into the variational autoencoder and mapped to the shared latent space.
[0032] Conditional probability modeling is applied in the latent space to calculate the probability of physical compliance and the probability of semantic consistency;
[0033] Conditional probabilities are integrated into a joint probability distribution using a Bayesian inference framework;
[0034] The Monte Carlo integration method is applied to calculate the integral, and the probability estimate of the authenticity is obtained.
[0035] Furthermore, the steps for implementing the physical semantic cross-attention mechanism include:
[0036] Construct a physical-to-semantic attention mapping to guide semantic understanding to focus on physically anomalous regions;
[0037] Construct a semantic-to-physical attention mapping so that physical analysis prioritizes semantically important regions;
[0038] By iteratively optimizing two attention maps, a deep fusion of physical and semantic features is achieved.
[0039] The fused feature representation is extracted and used for subsequent anomaly detection.
[0040] Furthermore, the step of analyzing the deviation between the joint physical-semantic distribution and the pre-established standard distribution of natural images includes:
[0041] Establish a standard distribution model based on large-scale real image data as a reference benchmark;
[0042] Calculate the Kulbac-Leibler divergence between the joint distribution and the standard distribution of the input image, and the degree of quantization deviation;
[0043] Set an adaptive threshold to determine the authenticity of the input image based on the degree of deviation;
[0044] Locate the key areas causing the deviation and provide a visual heatmap of the anomalies.
[0045] Furthermore, the step of generating the identification results and interpretability analysis report of the input image includes:
[0046] Generate a probability score for the authenticity of the input image to quantify the confidence level of the identification.
[0047] Identify potential abnormal regions and annotate them on the input image;
[0048] Provide an analysis of the causes of the anomalies, explaining the discrepancies between physical laws and semantic logic;
[0049] Generate a detailed evaluation report, including various indicators and comprehensive assessment results.
[0050] Furthermore, the step of integrating conditional probabilities into a joint probability distribution using a Bayesian inference framework specifically includes:
[0051] Temperature scaling calibration is applied to the joint distribution of physical semantics, and the reliability of probability estimation is improved by adjusting the temperature parameter after calibration.
[0052] A multi-scale interaction strategy is established using variational autoencoders, and the interaction between physical feature vectors and semantic feature vectors at different levels of abstraction is realized through cross-layer connections;
[0053] The application of feature enhancement gating units controls the intensity of information interaction between physical feature vectors and semantic feature vectors.
[0054] Furthermore, the collaborative working mechanism between the physical parameter extraction network layer and the semantic understanding network layer specifically includes:
[0055] The input image is pre-classified by an image type classifier, and the parameter configuration of the multi-scale convolutional filter and the attention weight of the Transformer-based visual model are automatically adjusted according to the classification results.
[0056] Incremental learning modules are constructed in the physical parameter extraction network layer and the semantic understanding network layer, respectively, so that the deep convolutional neural network structure and the Transformer-based visual model can continuously learn from new forgery technology samples and dynamically update the network weights.
[0057] A multimodal fusion mechanism is established to fuse the illumination model parameters and surface reflection characteristics output by the physical parameter estimation module with the key object information identified by the object detection and segmentation module, while integrating image metadata for comprehensive analysis.
[0058] This invention provides an artificial intelligence-based digital image authentication system for performing the aforementioned artificial intelligence-based digital image authentication method, comprising:
[0059] The three-layer fusion neural network architecture module is used to build a neural network architecture that includes a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer.
[0060] The physics analysis module is used to extract the input image from the network layer analysis using physical parameters and generate a set of physical feature vectors.
[0061] The semantic understanding module is used to process the input image using the semantic understanding network layer, generating semantic feature vectors and relationship graphs.
[0062] The cross-attention module is used to implement the physical-semantic cross-attention mechanism to enhance the bidirectional information flow between physical features and semantic features;
[0063] The joint distribution construction module is used to construct a physical-semantic joint distribution based on physical feature vectors and semantic feature vectors using the variational inference engine layer.
[0064] The distribution analysis module is used to analyze the degree of deviation between the physical and semantic joint distribution of the input image and the pre-established standard distribution of natural images;
[0065] The results output module is used to generate identification results and interpretability analysis reports for the input image based on the degree of deviation.
[0066] The beneficial effects of this invention are as follows: by constructing a dual verification mechanism of physical parameter extraction network layer and semantic understanding network layer, the detection capability of AI-generated images is improved, especially the identification accuracy of high-quality generated images is improved;
[0067] The coupling relationship established through the physical-semantic cross-attention mechanism can maintain stable detection performance even under targeted adversarial attacks, as long as there is inconsistency in the physical-semantic relationship.
[0068] The probability distribution analysis and anomaly heatmap generation technology built through the variational inference engine layer can accurately locate specific areas of physical and semantic conflicts in forged images and provide detailed anomaly cause analysis, providing credible scientific basis for the identification results.
[0069] Through an adaptive learning mechanism based on a Bayesian inference framework, the system can automatically optimize the physical-semantic coupling model as samples are processed, thereby gradually improving its ability to identify unknown forgery techniques.
[0070] By transforming the complex multi-dimensional verification problem into probability distribution similarity calculation through the "double helix" theoretical framework, the computational complexity of the algorithm is reduced, the processing efficiency is improved, and real-time processing of high-resolution images becomes possible.
[0071] The physical-semantic joint distribution modeling technique is applicable to various image types and forgery techniques, and is not limited to specific image content or tampering methods. It exhibits good stability in different application scenarios.
[0072] Through a modular three-layer fusion neural network architecture design, the system can easily integrate new physical models or semantic understanding technologies according to actual needs, allowing for the individual updating or replacement of specific components without affecting the overall architecture, providing good scalability for technology iteration;
[0073] This invention realizes a paradigm shift from "finding traces of forgery" to "verifying naturalness," providing a brand-new theoretical foundation and technical approach for the field of digital image authentication, and helping to address the increasingly complex challenges of image forgery. Attached Figure Description
[0074] Figure 1 This is a flowchart of a digital image identification method based on artificial intelligence in this invention;
[0075] Figure 2 This is a scatter plot showing the relationship between the Kullback-Leibler divergence and the probability of the true image.
[0076] Figure 3 This is a graph showing the combination of the system's detection accuracy and false alarm rate under different threshold settings;
[0077] Figure 4 This is a bar chart comparing the accuracy of the physical-semantic double helix theory, traditional physical verification, and traditional semantic inspection methods in identifying ordinary GAN images and high-quality GAN images.
[0078] Figure 5 This is a bar chart showing the identification accuracy of the physical-semantic double helix theory and traditional methods on different types of images;
[0079] Figure 6 This is a radar chart comparing the performance of the physics-semantic double helix theory and traditional methods across six dimensions. Detailed Implementation
[0080] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0081] At least one embodiment of the present invention discloses a digital image identification method based on artificial intelligence, such as... Figure 1 As shown, it includes:
[0082] Step 1: Construct a three-layer fusion neural network architecture, including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer;
[0083] Step 1.1, physical parameter extraction of network layers;
[0084] A deep convolutional neural network structure is employed, including a feature extraction module and a physical parameter estimation module, to extract physical feature parameters such as illumination distribution, shadow direction, and perspective from the input image. Specifically, the network uses ResNet-50 as the backbone, with dedicated physical parameter estimation branches added on top of it. Each branch is optimized for specific physical properties. For example, the illumination estimation branch infers the direction and intensity of the main light source by analyzing the image brightness gradient and shadow distribution; the perspective branch estimates camera parameters by detecting vanishing points and parallel lines. Optionally, in some implementations, DenseNet or EfficientNet can also be used as the backbone to balance performance and computational efficiency in different scenarios.
[0085] The structure of the physical parameter extraction network layer can be further refined into four main components:
[0086] The backbone feature network adopts a ResNet-50 architecture, consisting of 5 stages, each containing 3, 4, 6, and 3 residual blocks respectively. Each residual block contains three convolutional layers and shortcut connections, effectively solving the gradient vanishing problem in deep networks. The network input is a 224×224×3 RGB image, which, after initial 7×7 convolutions and max pooling, is sequentially passed through the residual blocks of each stage, finally outputting a 2048-channel feature map. Optionally, in specific application scenarios, ResNet variants of different depths (such as ResNet-34 or ResNet-101) can be used to balance computational complexity and feature extraction capabilities.
[0087] The illumination estimation module consists of three 1×1 convolutional networks, with the number of channels reduced from 2048 to 1024, 512, and 256 respectively. It then outputs light source parameters through global average pooling and fully connected layers. These parameters include the main light source direction (azimuth and elevation angles), light source intensity (normalized value), and ambient light intensity. This module also integrates an illumination consistency evaluation unit, which detects potential inconsistencies by comparing the illumination parameter estimation results for different regions of the image.
[0088] The shadow analysis module employs a U-Net architecture for pixel-level shadow segmentation, then applies a gradient direction analysis algorithm to infer the light source direction. This module first outputs a shadow probability map, then locates the boundary between the object and the shadow through shadow edge detection, and finally calculates the light source direction based on geometric relationships. In the actual implementation, the results of this module are cross-validated with those of the illumination estimation module to improve the reliability of parameter estimation.
[0089] The perspective geometry module comprises a line segment detection network and a vanishing point estimation network. The line segment detection network is implemented using a neural network based on the LSD (LineSegmentDetector) algorithm, capable of identifying line segments in the image. The vanishing point estimation network uses the RANSAC algorithm to group and vote on detected line segments, determining the locations of major vanishing points and thus inferring camera parameters and scene geometry. In some implementations, this module can also introduce prior geometric constraints, such as the Manhattan world assumption, to further improve the accuracy of geometric parameter estimation.
[0090] During the training phase, the physics parameter extraction network employs a multi-task learning framework, simultaneously optimizing multiple physics parameter estimation objectives. The loss function consists of two parts: parameter estimation loss and physics consistency loss.
[0091] ;
[0092] in, This represents the total loss of the physical parameter extraction network; Indicates the summation symbol; Indicates the first Loss weighting coefficients for various physical parameters; Indicates the first The estimation loss of various physical parameters (such as the angular loss of the light source direction, the mean square error of the intensity, etc.); The weighting coefficients representing the physical consistency loss; This represents the loss of physical consistency constraints (such as the consistency between the direction of the light source and the direction of the shadow).
[0093] Through this multi-task learning approach, the network can learn the intrinsic relationships between physical parameters, thereby improving the overall estimation accuracy.
[0094] Step 1.2, Semantic Understanding Network Layer;
[0095] Based on the Transformer architecture, this network includes a self-attention module and a semantic relationship inference module to identify objects in images, understand the relationships between objects, and infer event logic. The network employs a VisionTransformer (ViT) structure, capturing long-distance semantic dependencies by segmenting the image into a series of patches and calculating the self-attention relationships between these patches. In its implementation, the input image is first segmented into 16×16 pixel patches, then processed by a 12-layer Transformer encoder to finally generate object recognition results and a relationship graph. It should be understood that in another embodiment, a SWINTransformer or a hybrid structure based on convolution and Transformer can also be used to better capture local and global information.
[0096] The structure of semantic understanding network layers can be further refined into the following five key parts:
[0097] Image segmentation and embedding: The input image is first segmented into N fixed-size patches, each 16×16 pixels. For a 224×224 resolution input image, a total of 196 patches are generated. Then, each patch is mapped to a 768-dimensional embedding vector through a linear projection layer. Furthermore, position embedding is added to preserve spatial location information, and a learnable class token is incorporated, which will be used for final image-level prediction. Optionally, in some implementations, image patches of different sizes (e.g., 8×8 or 32×32), or a two-dimensional position encoding method, can be used to accommodate different image sizes and resolution requirements.
[0098] The Transformer encoder stack consists of 12 standard Transformer encoder blocks, each containing a Multi-Head Self-Attention (MSA) module and a Feed-Forward Network (FFN) module. In the self-attention calculation, each block uses 12 attention heads, each with a dimension of 64. The overall calculation formula is:
[0099] ;
[0100] in, Indicates attention output; Represents the query matrix; Represents the key matrix; Represents a value matrix; This represents the transpose operator; This represents the softmax normalization operation; This represents the dimension of the key vector, used to scale the dot product result to prevent gradient vanishing.
[0101] The feedforward neural network consists of two fully connected layers, with the GELU activation function used in between. The dimension expands from 768 to 3072 and then back to 768. Layer normalization (LayerNorm) and residual connections are applied after each module to ensure training stability and information flow.
[0102] The object detection and segmentation module adds dedicated object detection and instance segmentation heads based on the output features of the Transformer encoder. The detection head adopts a design similar to DETR (DetectionTransformer), using learnable object queries and encoder output features for cross-attention calculation to directly predict bounding boxes and categories. Specifically, it uses 100 object queries, each interacting with image features through a 6-layer Transformer decoder, ultimately outputting the object location (bounding box coordinates) and category probability. The segmentation head, based on the detection results, uses dynamic convolution to generate a mask for each object, achieving pixel-level segmentation. It's worth noting that, unlike traditional object detection methods, this module does not require non-maximum suppression post-processing, simplifying the algorithm and improving speed.
[0103] Relationship Graph Construction Module: Based on the detected objects, constructs a scene relationship graph (SceneGraph). The relationship graph adopts a graph structure.
[0104] ;
[0105] in Represents the scene relationship diagram. Represents a set of nodes. Represents an edge set.
[0106] Relationship prediction employs a bidirectional attention mechanism: first, relationship features are calculated from the features and location information of the two objects; then, the relationship category is predicted using a multilayer perceptron. In practice, to control computational complexity, only one attention is considered for each object. The most relevant other objects ( Instead of calculating all possible object pairs, this module uses a knowledge graph (=10) to guide the relationship prediction process by distilling knowledge. Furthermore, to enhance the accuracy of relationship prediction, this module introduces an external knowledge graph as prior knowledge.
[0107] Event Logic Reasoning Module: This module performs high-order event logic reasoning based on the relationship graph. It employs a Graph Convolutional Network (GCN) architecture, using three layers of graph convolutions for message passing. The convolution operation is defined as follows:
[0108] ;
[0109] in, Indicates the first The node feature matrix of the layer; Represents the ReLU activation function; This represents the degree matrix after adding self-loops; This represents the adjacency matrix after adding the self-loop; Represents the negative 1 / 2 power (normalization) of the degree matrix; Indicates the first The node feature matrix of the layer; Indicates the first The learnable weight matrix of the layer.
[0110] Through iterative graph convolution operations, information propagates within the relationship graph, allowing each node to gradually aggregate contextual information, forming a more global scene understanding. Ultimately, this module outputs scene event type predictions and abnormal event detection results. Optionally, in some implementations, a temporal reasoning mechanism can be introduced to analyze the event evolution logic across multiple frames, further enhancing semantic understanding capabilities.
[0111] The semantic understanding network layers are trained using a multi-stage strategy: first, the basic Transformer encoder is pre-trained on a large-scale image classification dataset; then, the object detection and segmentation modules are fine-tuned on object detection and instance segmentation datasets; finally, the relationship graph construction and event logic reasoning modules are trained on a scene graph dataset. The training loss function consists of a weighted sum of object classification loss, bounding box regression loss, instance segmentation loss, and relationship classification loss.
[0112] ;
[0113] in, This represents the total loss of the semantic understanding network layers; The weighting coefficients representing the object classification loss; Represents the object classification loss; The weighting coefficients represent the bounding box regression loss; This represents the bounding box regression loss; The weight coefficients represent the instance segmentation loss. Indicates instance segmentation loss; The weighting coefficients representing the loss of relation classification; This represents the loss in relation classification.
[0114] Step 1.3, Variational Inference Engine Layer;
[0115] This paper integrates variational autoencoders and Bayesian inference networks to build a joint probability distribution model of physical parameters and semantic features, enabling interpretable inference. The engine consists of an encoder network, a latent variable space, and a decoder network. The encoder network maps physical and semantic features to the latent variable space, while the decoder reconstructs the original features from the latent representation. During training, the network parameters are optimized by minimizing the reconstruction error and Kullback-Leibler divergence, ensuring that the latent space effectively represents the probability distribution of physical-semantic relationships. Furthermore, the dimensionality of the latent space can be adjusted for different application scenarios, achieving a balance between expressive power and computational complexity.
[0116] The structure of the variational inference engine layer can be broken down into the following main components:
[0117] Feature integration preprocessing: This component receives the output features from the physical parameter extraction network and the semantic understanding network, and performs preliminary feature normalization and alignment. The physical feature vector (512-dimensional) and the semantic feature vector (768-dimensional) are first transformed into a common feature space (256-dimensional) through independent fully connected layers. To ensure information integrity, residual connections and layer normalization operations are introduced.
[0118] ;
[0119] ;
[0120] in, This represents the normalized physical feature vector; Presentation layer normalization operation; Represents the original physical feature vector; Fully connected layer transformations representing physical characteristics; This represents the normalized semantic feature vector; Represents the original semantic feature vector; Fully connected layer transformations representing semantic features.
[0121] Conditional Variational Encoder: This component employs a Conditional Variational Autoencoder (CVAE) architecture, encoding normalized physical and semantic features as probability distributions rather than deterministic points. The encoder structure consists of three fully connected layers, each followed by batch normalization and a LeakyReLU activation function. The final output is the latent variables. Posterior distribution parameters: mean vector Sum of logarithmic variance vector All dimensions are 128.
[0122] The sampling process uses reparameterization techniques to ensure that gradients can propagate through random sampling operations:
[0123] ;
[0124] in, This represents the latent variables obtained from sampling; Represents the mean vector of the posterior distribution; Denotes the standard deviation vector of the posterior distribution. , The log-variance vector representing the posterior distribution is output by the encoder network; Indicates the distribution from the standard normal distribution The sampled noise vector; This represents a multivariate normal distribution with a mean of 0 and a covariance of an identity matrix. This indicates that the distribution follows a certain relationship.
[0125] In its implementation, the encoder is designed to accept conditional input, namely physical parameters θ (such as lighting conditions, camera parameters, etc.), such that the posterior distribution becomes a conditional probability. This enhances the model's adaptability to different physical conditions.
[0126] Prior network: To more accurately model the prior distribution of physical parameters, a learnable prior network is introduced, instead of simply assuming a standard normal distribution. This network receives the physical parameters... As input, the output prior distribution Parameters: Prior mean vector Sum of logarithmic variance vector The network structure consists of two fully connected layers with a ReLU activation function in between. By learning the prior distribution, the model can capture the natural distribution characteristics of latent variables under different physical conditions, improving modeling accuracy. Optionally, to further enhance the expressive power of the prior distribution, the NormalizingFlow technique can be used to enhance the complexity of the distribution through a series of invertible transformations.
[0127] Physico-semantic probabilistic decoder: This component reconstructs the physical compliance probability and semantic consistency probability from the latent variable z. The decoder contains four parallel branches:
[0128] Physical Reconstruction Branch: A three-layer fully connected network reconstructs physical features;
[0129] Physical compliance branch: Outputs the probability of physical compliance. The calculation is performed by comparing the reconstructed physical characteristics with the expected physical parameters.
[0130] Semantic Reconstruction Branch: A three-layer fully connected network reconstructs semantic features;
[0131] Semantic consistency branch: Outputs the semantic consistency probability. The calculation is performed by evaluating the degree to which the reconstructed semantic relationships conform to common sense knowledge.
[0132] The output of each branch is normalized to the 0-1 range using the sigmoid function, representing the corresponding probability value. It should be noted that the calculation of compliance probability not only considers reconstruction error, but also includes hard constraint checks of physical laws, such as the consistency between shadow direction and light source, and the relationship between reflection angle and surface normal.
[0133] The probability integration module implements the integration operation in the Bayesian inference framework to calculate the posterior probability of image authenticity. This module uses the Monte Carlo method to approximate the integral calculation.
[0134] ;
[0135] in, Representing an image is the posterior probability of the real image; This indicates an approximate equality; This represents the summation operator; Indicates the number of samples; Indicates the first physical parameters Below, image The probability of physical compliance; Indicates the first physical parameters Below, image The probability of semantic consistency; Indicates from the prior distribution The first sampled One physical parameter.
[0136] To improve sampling efficiency, the module implements an importance sampling strategy, prioritizing sampling from high-probability regions in the parameter space. In actual deployment, the number of samples... It can be flexibly adjusted according to computing resources and accuracy requirements. It is usually set to 1000 on high-performance devices and reduced to 100-200 on resource-constrained devices. At the same time, variance reduction techniques are used to maintain estimation accuracy.
[0137] The variational inference engine layer is trained end-to-end, and the loss function consists of four parts:
[0138] ;
[0139] in, This represents the total loss of the variational inference engine layer; The weighting coefficients representing the reconstruction loss; Represents the reconstruction loss (a measure of the information retention capability of the encoding-decoding process); The weighting coefficients represent the KL divergence loss. This represents the KL divergence loss (making the posterior distribution closer to the prior distribution); Weighting coefficients representing the loss of physical compliance; This indicates losses related to physical compliance oversight; The weight coefficients representing the semantic consistency loss; This represents the semantic consistency supervision loss.
[0140] Weighting coefficient , , The validation set was optimized to balance the importance of different learning objectives. In the early stages of training, a larger validation set was used. and smaller Gradually increase The value of is used to implement a KL divergence training strategy similar to annealing, which avoids the posterior distribution from collapsing to the prior distribution too early.
[0141] Step 1.4, Inter-layer interaction connections;
[0142] A bidirectional attention connection mechanism is constructed to achieve information interaction between the physical and semantic layers, forming a "double helix" structure. Specifically, two sets of attention modules are designed: one mapping from physical features to semantic features, and the other mapping from semantic features to physical features. Each mapping is implemented through a multi-head attention mechanism, allowing fine-grained information exchange between different features. This bidirectional connection is not a simple feature concatenation, but a dynamic, content-based selective information transfer, ensuring that physical analysis focuses on semantically relevant regions, while semantic understanding is guided by physical constraints. It is important to emphasize that the "double helix" structure in this application is an innovative information flow design, distinct from traditional unidirectional or parallel processing architectures.
[0143] The implementation of inter-layer interaction connections can be divided into the following core parts:
[0144] Bidirectional cross-attention mechanism: This mechanism is the core of the "double helix" structure, enabling deep interaction between physical and semantic features. Physical-to-semantic attention module. and semantic-to-physical attention module Together, they constitute a complete interactive system. Each attention module employs a multi-head attention structure, with the specific calculations as follows:
[0145] Physical-to-semantic attention computation:
[0146] ;
[0147] ;
[0148] ;
[0149] ;
[0150] in, A query matrix representing semantic features; Represents the semantic query weight matrix; Represents semantic feature vectors; A bond matrix representing physical characteristics; Represents the physical bond weight matrix; Represents the physical characteristic vector; A value matrix representing physical characteristics; Represents the physical value weight matrix; This represents the attention output from the physical to the semantic level. This represents the softmax normalization operation; This represents the dimension of the key vector in the physical-to-semantic attention module.
[0151] Semantic-to-physical attention computation:
[0152] ;
[0153] in, A query matrix representing physical characteristics; Represents the physical query weight matrix; Represents the physical characteristic vector; A key matrix representing semantic features; Represents the semantic key weight matrix; Represents semantic feature vectors; A value matrix representing semantic features; Represents the semantic value weight matrix; This represents the semantic-to-physical attention output; This represents the softmax normalization operation; This represents the dimension of the key vector in the semantic-to-physical attention module.
[0154] In practice, each attention module uses 8 heads, each with a dimension of 32, to capture feature dependencies across different dimensions. Optionally, in some implementations, the number of attention heads can be dynamically adjusted based on the complexity of the input image to balance computational efficiency and expressive power.
[0155] Iterative Refinement Layer: To further enhance the fusion depth of physical and semantic information, an iterative refinement layer is introduced to achieve multi-round information exchange. In the... In the round of iteration, the information update formula is:
[0156] ;
[0157] in, Indicates the first The physical feature vector after round iteration; Indicates the first The physical feature vector after round iteration; This represents the learning rate parameter, which controls the speed at which information is updated. A feedforward neural network representing physical features (two fully connected layers + ReLU); Indicates the first The attention output from semantics to physics; Indicates the first The semantic feature vector after round of iterations; Indicates the first The semantic feature vector after round of iterations; Feedforward neural networks representing semantic features; Indicates the first Attention output from physics to semantics.
[0158] Feature Enhancement Gating Unit: To avoid interference from irrelevant information, a feature enhancement gating unit is used to control the intensity of information interaction. The gating mechanism dynamically adjusts the influence of the cross-attention output on the original features through learnable parameters.
[0159] ;
[0160] ;
[0161] ;
[0162] ;
[0163] in, Represents the physical characteristic gating coefficient; This represents the sigmoid activation function; Represents the physical gate weight matrix; This represents the concatenation of physical features and semantics with the output of physical attention. Indicates physical gating bias; Indicates the enhanced physical characteristics; Indicates element-wise multiplication; This represents the semantic-to-physical attention output; Indicates the original physical characteristics; Represents the semantic feature gating coefficient; This represents the semantically gated weight matrix; This represents the concatenation of semantic features with the physical-to-semantic attention output; Indicates semantic gating bias; This represents the enhanced semantic features; This represents the attention output from the physical to the semantic level. Represents the original semantic features.
[0164] Multi-scale interaction strategy: To capture physical-semantic relationships at different levels of abstraction, interactive connections are established at multiple layers of the network. Specifically, cross-connections are established between low-, mid-, and high-level features of the physical parameter extraction network and the semantic understanding network, forming a multi-layered "double helix" structure.
[0165] Low-level connections: Focus on the relationship between local textures and basic semantic elements;
[0166] Mid-level connections: focus on the semantic relationship between object components and regions;
[0167] High-level connectivity: Focuses on the relationship between global physical parameters and scene semantics.
[0168] Connections at different levels employ similar attention mechanisms, but parameters are not shared to adapt to the characteristics of different levels of abstraction. Multi-scale interaction greatly enriches the ways physical and semantic information are fused, enabling the system to simultaneously capture information correlations at both the micro and macro levels. Therefore, this design enhances the system's ability to handle complex scenes, especially for images that simultaneously contain multiple physical phenomena and rich semantic content.
[0169] Attention Visualization and Interpretation Module: To enhance the system's interpretability, an attention visualization module was designed to display the attention relationships between physical and semantic features in real time. This module maps attention weights back to the original image space, generating heatmaps to show the dependence strength between different features. Visualizing attention weights not only helps in understanding the system's decision-making process but also visually demonstrates patterns of physical-semantic associations, providing important clues for subsequent analysis. In its implementation, the system remaps feature-level attention weights to the pixel level through upsampling and bilinear interpolation, and uses color coding to represent dependence strength for easy human understanding. Optionally, attention visualization can also serve as an intermediate debugging tool to help developers optimize system parameters and identify potential problems.
[0170] In summary, the inter-layer interactive connection mechanism achieves deep integration of physical and semantic information through the aforementioned carefully designed components, forming a true "double helix" structure. This structure not only allows physical analysis and semantic understanding to guide and reinforce each other, but also treats conflicts as important discriminative signals, which is one of the core innovations that distinguishes this application from existing technologies.
[0171] Step 2: Use physical parameters to extract network layer analysis input images and generate a set of physical feature vectors;
[0172] Step 2.1: Apply a multi-scale convolutional filter to the input image to extract image features at different scales;
[0173] In one embodiment according to this application, five different scales of convolutional kernels (3×3, 5×5, 7×7, 9×9, 11×11) are used to process the input image. Each scale of the convolutional layer contains 64 filters to capture features at various levels, from detailed texture to macroscopic structure. Optionally, the filter parameters can be automatically adjusted according to the image resolution to ensure that features of appropriate granularity can be extracted at different resolutions. It should be noted that in other embodiments, dilated convolution technology can also be used to obtain different receptive fields by adjusting the dilation rate without increasing the number of parameters or reducing the resolution.
[0174] Step 2.2: Fuse multi-scale features through a feature pyramid structure to form a comprehensive feature map;
[0175] This process employs a bidirectional feature fusion strategy, combining top-down and bottom-up approaches, to ensure the effective integration of information at different scales. Specifically, a bottom-up feature pyramid is first constructed through downsampling, followed by a top-down feature pyramid constructed through upsampling and lateral connections, ultimately generating a comprehensive feature map containing multi-scale information. It should be understood that this structure can simultaneously preserve detailed information and capture global semantics. Furthermore, in some implementations, an attention mechanism can be introduced to assist the feature pyramid fusion process, adaptively adjusting the importance weights of features at different scales to further improve the fusion effect.
[0176] Step 2.3: Apply a dedicated physical parameter estimation module to infer physical parameters from the comprehensive feature map;
[0177] Physical parameters include lighting model parameters, surface reflection characteristics, shadow consistency, and perspective relationships;
[0178] According to embodiments of this application, the module comprises multiple parallel sub-networks, each specifically responsible for estimating a particular physical parameter. For example, the illumination model sub-network outputs parameters such as light source direction, intensity, and color temperature; the reflection property sub-network estimates the diffuse and specular reflection properties of the material; and the perspective relationship sub-network infers the camera's intrinsic and extrinsic parameters. Additionally, in some embodiments, a physical consistency verification sub-network may be included to check whether different physical parameters satisfy basic physical constraints. Optionally, these sub-networks can employ a multi-task learning framework, sharing a low-level feature extraction layer, which not only reduces the number of parameters but also improves overall estimation accuracy through knowledge transfer between tasks.
[0179] Step 2.4: Calculate the uncertainty estimate of the physical parameters to provide confidence information for subsequent probabilistic inference;
[0180] In a preferred embodiment according to this application, uncertainty estimation is achieved through Monte Carlo Dropout, which maintains the activation of the Dropout layer during the inference phase, performs multiple forward propagations on the same input, and uses the statistical output variability as a measure of uncertainty. This uncertainty estimation method does not require additional training and can be directly integrated into existing networks. Furthermore, in other embodiments, ensemble learning methods or Bayesian neural networks can also be used to obtain uncertainty estimates. Although these methods may require more computational resources, they can provide more accurate uncertainty quantification in certain specific scenarios.
[0181] Step 3: Apply the semantic understanding network layer to process the input image, generating semantic feature vectors and relationship graphs;
[0182] Step 3.1: Extract semantic features of the image using a pre-trained visual Transformer model;
[0183] According to embodiments of this application, the model is pre-trained on a large-scale image dataset and possesses powerful semantic understanding capabilities. In a preferred embodiment, a ViT-Base variant is employed, comprising 12 Transformer blocks, each with 12 attention heads, a hidden layer dimension of 768, and a total of approximately 86 million parameters. It should be noted that for different application scenarios, smaller ViT-Small or larger ViT-Large variants can be selected to balance performance and computational resource requirements. Furthermore, in environments with limited computational resources, lightweight Transformer architectures such as MobileViT can be considered, or knowledge distillation techniques can be used to transfer knowledge from a large model to a smaller model, reducing computational burden while maintaining performance.
[0184] Step 3.2: Apply the object detection and segmentation module to identify key objects and regions in the image;
[0185] According to one embodiment of this application, the module is based on the Mask R-CNN architecture, but the backbone network is replaced with a Transformer to enhance feature extraction capabilities. The processing flow includes four stages: region proposal generation, object classification, bounding box regression, and instance segmentation, ultimately outputting the category, location, and pixel-level mask for each identified object. Optionally, in some implementations, a keypoint detection branch can be added for more fine-grained object structure analysis. Furthermore, for scenarios requiring real-time processing, a single-stage detector architecture such as YOLOX or DETR can be used, reducing computational complexity and accelerating processing speed through end-to-end design.
[0186] Step 3.3, construct the object relationship diagram;
[0187] An object relationship graph represents the spatial and functional relationships between objects in an image. In embodiments according to this application, the relationship graph is represented by a graph structure, where nodes correspond to detected objects and edges represent relationships between objects. Relationship types include spatial relationships (e.g., "above", "inside"), functional relationships (e.g., "use", "contain"), and semantic relationships (e.g., "belong to the same category"). The graph is constructed by calculating the similarity and positional relationships between object feature vectors. Furthermore, to enhance relationship reasoning capabilities, external knowledge graphs can be introduced as a source of prior knowledge to guide the relationship prediction process. It should be understood that in other possible implementations, a global relationship modeling method based on an attention mechanism can also be used, which does not require explicit construction of a graph structure and directly captures the interaction relationships between objects through a self-attention mechanism.
[0188] Step 3.4: Apply the event reasoning module to analyze the possible events in the scenario and their rationality;
[0189] According to a preferred embodiment of this application, the module employs Graph Convolutional Neural Networks (GCNs) to process object relationship graphs, inferring interactions between objects and possible events in the scene through a message passing mechanism. Event rationality assessment is based on a pre-trained event knowledge base, which contains event types that common object combinations may form and their probabilities of occurrence. It is important to emphasize that event rationality assessment is a significant innovation of this application, going beyond simple object recognition and relationship understanding to achieve a higher level of scene semantic understanding. Optionally, in some implementations, temporal information (if applicable) can be combined to construct a dynamic event reasoning model, which not only considers the rationality of events in static scenes but also assesses the naturalness of the event development sequence.
[0190] Step 4: Based on the physical feature vector and semantic feature vector, construct a joint physical-semantic distribution using the variational inference engine layer;
[0191] Step 4.1: Input the physical feature vector and semantic feature vector into the variational autoencoder and map them to the shared latent space;
[0192] According to embodiments of this application, a variational autoencoder consists of an encoder and a decoder. The encoder maps high-dimensional features to the mean and variance parameters of low-dimensional latent variables, while the decoder reconstructs the original features from the latent variables. In this embodiment, the latent space dimension is set to 128, and both the encoder and decoder are implemented using multi-layer fully connected networks. Optionally, a convolutional variational autoencoder structure can also be used to better preserve the spatial information of the features. Furthermore, in another embodiment, a hierarchical variational autoencoder (HVAE) architecture can be introduced, which, by constructing a multi-level latent variable structure, more effectively captures the feature relationships at different levels of abstraction.
[0193] Step 4.2: Apply conditional probability modeling in the latent space to calculate the physical compliance probability and semantic consistency probability;
[0194] According to one embodiment of this application, the conditional probability model employs a Conditional Variational Autoencoder (CVAE) structure, allowing the inference of probability distributions given specific conditions. (Physical compliance probability) Semantic consistency probability is calculated by comparing the consistency between actual and expected physical parameters. The result is obtained by evaluating the degree to which the semantic relationships of the image content conform to common sense knowledge. It should be understood that this dual-verification mechanism can comprehensively evaluate the authenticity of an image from both physical and semantic dimensions. Optionally, in some implementations, adversarial training strategies can be introduced to enhance the expressive power and robustness of the conditional probability model through a generator-discriminator structure.
[0195] Step 4.3: Using the Bayesian inference framework, integrate the conditional probabilities into a joint probability distribution;
[0196] The mathematical expression is:
[0197] ;
[0198] in Representing an image is the posterior probability of the real image; Represents the physical parameter space Integrate points; Indicates physical parameters Below, image The probability of physical compliance; Indicates physical parameters Below image The probability of semantic consistency; Representing physical parameters The prior distribution; Indicates the physical parameters The integral variable.
[0199] This integral expresses how physical compliance and semantic consistency jointly influence image authenticity across all possible physical parameter spaces. Furthermore, in practical applications, the weights of physical compliance and semantic consistency can be adjusted according to different scenarios to accommodate different types of images and forgery methods. For example, for landscape images, more emphasis might be placed on physical compliance verification; while for images of people or events, more emphasis might be placed on semantic consistency evaluation.
[0200] Step 4.4: Apply the Monte Carlo integration method to approximate the integral and obtain the final probability estimate of the truth.
[0201] According to embodiments of this application, the Monte Carlo method is used for approximate calculation. Specifically, from the prior distribution... Medium sampling One parameter point:
[0202] ;
[0203] in, , , They represent the 1st, 2nd, and 3rd respectively. A sampled physical parameter vector, Indicates the total number of sampling points;
[0204] Then calculate the function value at each sample point. Finally, the average value is taken as the approximate result of the integral. In actual implementation, The value is set to 1000 to balance computational efficiency and approximate accuracy. Optionally, in environments with limited computational resources, more efficient Monte Carlo variants such as importance sampling can be used to reduce the number of sampling points required. Furthermore, in some other possible implementations, variational inference techniques can be used to directly optimize the approximate representation of the posterior distribution, avoiding explicit integral calculations and further improving computational efficiency.
[0205] Step 5: Implement the physical-semantic cross-attention mechanism to enhance the bidirectional information flow between physical and semantic features;
[0206] Step 5.1: Construct a physical-to-semantic attention mapping to guide semantic understanding to focus on physically anomalous regions;
[0207] According to a preferred embodiment of this application, the mapping generates an attention weight map by calculating the anomaly score at each location in the physical feature map. This weight map is then applied to semantic features, making the semantic analysis process focus more on physically suspicious regions. In a specific implementation, the anomaly score is calculated by comparing the deviation between the actual physical features and the expected physical features. Furthermore, to enhance the stability of the attention mechanism, residual connections can be introduced, preserving the original feature information while enhancing the representation of key regions. Another optional implementation is to employ an adaptive thresholding strategy, dynamically adjusting the screening criteria for anomaly regions based on the complexity and quality of the image, thereby improving the mechanism's adaptability to different types of images.
[0208] Step 5.2: Construct a semantic-to-physical attention mapping so that physical analysis prioritizes semantically important regions;
[0209] In one embodiment of this application, the mapping generates attention weights for physical features based on the importance scores of semantic features, ensuring that the physical analysis process focuses on semantically critical regions (such as main objects, interaction areas, etc.). Importance scores are calculated using indicators such as object classification confidence and node centrality in the relationship graph. It is important to emphasize that this bidirectional attention mapping is one of the core innovations of this application; it enables mutual guidance between physical and semantic analysis, forming a true "double helix" information flow structure. Optionally, in some implementations, differentiated attention weights can be assigned to different semantic entities based on the functionality of objects and their role importance in events, allowing the physical analysis to focus more on key elements in the scene.
[0210] Step 5.3: Deep fusion of physical and semantic features is achieved by iteratively optimizing the two attention maps alternately.
[0211] According to embodiments of this application, the process employs an iterative refinement strategy. In each iteration, the two attention maps influence and update each other, gradually improving the quality of information interaction. Specifically, the first... The physics-semantic attention mapping of the round iteration is affected by the first The semantic-physical attention mapping of the wheel has an impact, and vice versa. Setting the number of iterations to 3 strikes a balance between accuracy and efficiency. Optionally, in some implementations, the number of iterations can be dynamically adjusted, adaptively determining the depth of information exchange based on image complexity. Furthermore, an early stopping mechanism can be introduced, terminating the iteration prematurely when the change in attention distribution is less than a preset threshold, further improving computational efficiency.
[0212] Step 5.4: Extract the fused feature representation for subsequent anomaly detection;
[0213] In one embodiment of this application, the fused features are obtained by nonlinearly combining attention-weighted physical and semantic features. Specifically, a gated fusion unit is employed, which includes a control gate that dynamically adjusts the contribution weights of different features based on the relevance of the input features, ensuring the flexibility and adaptability of the information flow. Therefore, this fusion mechanism can automatically adjust the importance of physical and semantic features according to the characteristics of different images, improving the system's versatility and robustness. It should be understood that in other embodiments, a multi-head attention mechanism can be used to achieve finer-grained feature fusion, or a feature transformation module can be introduced to enhance the expressive power of features through nonlinear transformation.
[0214] Step 6: Analyze the deviation between the physical-semantic joint distribution and the pre-established standard distribution of natural images;
[0215] Step 6.1: Establish a standard distribution model based on large-scale real image data as a reference benchmark;
[0216] According to a preferred embodiment of this application, the model is trained on a dataset containing 1 million real images, covering diverse samples from different scenes, lighting conditions, and shooting devices. The training process employs unsupervised learning, learning the physical-semantic joint distribution features of real images through a variational autoencoder. Furthermore, to improve the model's representativeness, the dataset includes images of different quality levels and sources, ensuring that the standard distribution can cover the diversity of real images. It should be noted that in some implementations, specific standard distribution models can also be constructed according to specific application domains, such as news image standard models or art work standard models, to improve the identification accuracy for specific types of images.
[0217] Step 6.2: Calculate the Kullback-Leibler divergence between the joint distribution and the standard distribution of the image under test, and quantify the degree of deviation;
[0218] In one embodiment of this application, the Kullback-Leibler divergence is an effective tool for measuring the difference between two probability distributions, and its calculation formula is as follows:
[0219] ;
[0220] in Represents distribution and Kullback-Leibler divergence; This represents the summation operator; Indicates the image under test at the th The probability value of the dimension; Represents the natural logarithm function; Indicates the standard distribution at the th The probability value of the dimension; Indicates the dimension of the distribution.
[0221] In practical implementations, due to the complexity of the distribution, the Monte Carlo method is used to approximate the divergence value. Optionally, in some implementations, other distribution metrics such as Jensen-Shannon divergence or Wasserstein distance can also be used to accommodate the different characteristics of different types of distributions. Furthermore, for environments with limited computational resources, low-rank approximation or feature dimensionality reduction techniques can be used to reduce computational complexity while maintaining the validity of the measurement.
[0222] like Figure 2 The figure shows the relationship between the Kullback-Leibler divergence of an image and the probability of it being classified as a true image. The horizontal axis represents the KL divergence between the image distribution and the standard distribution, and the vertical axis represents the probability of the system classifying it as a true image. It can be seen that as the KL divergence increases, the probability of an image being classified as true decreases significantly, and the two show a strong negative correlation. This relationship verifies the effectiveness of using KL divergence as a criterion for judging the authenticity of images, provides a theoretical basis for threshold setting, and demonstrates that this divergence calculation method can accurately reflect the degree of deviation of an image from its natural distribution.
[0223] Step 6.3: Set an adaptive threshold to determine the authenticity of the image based on the degree of deviation;
[0224] According to embodiments of this application, the threshold setting employs a dynamic strategy, adjusting based on the image's content complexity, quality, and application scenario.
[0225] Specifically, the threshold T is composed of the base threshold. and scene adjustment factor composition:
[0226] ;
[0227] in This represents the adaptive threshold ultimately used for the judgment; This represents the baseline threshold, determined through grid search optimization on a validation set containing 50,000 labeled images, with a default value of 0.65. This represents the scene adjustment factor, which is dynamically allocated by the scene classification module based on the image type. The value range is 0.8-1.2. Higher values are set for complex scenes (such as crowds or multi-object interactions), while lower values are set for simple scenes (such as landscapes or still life).
[0228] In practical applications, adjustable threshold options can be provided based on users' different needs for accuracy and recall. For example, stricter thresholds can be set in news media applications to reduce false positives, while more lenient thresholds can be set in rapid screening applications to improve detection rates.
[0229] It should be understood that in other implementations, a multi-threshold strategy may also be adopted, setting three thresholds. The judgment results are divided into:
[0230] Determining the truth: KL divergence < ;
[0231] Possibly true: ≤KL divergence< ;
[0232] Possible forgery: ≤KL divergence< ;
[0233] Determining forgery: KL divergence ≥ ;
[0234] Provides more granular identification results, among which , , These represent the first, second, and third judgment thresholds, respectively.
[0235] like Figure 3As shown, the system's detection accuracy (bar chart) and false alarm rate (line chart) are displayed under different threshold settings. As the threshold increases from 0.1 to 1.0, the detection accuracy gradually decreases, while the false alarm rate increases. This chart helps users find the optimal balance between accuracy and false alarm rate based on the specific application scenario. For example, in news media applications where the false alarm rate requirement is lower, a higher threshold (e.g., 0.8-0.9) can be selected; while in preliminary screening applications requiring a high detection rate, a lower threshold (e.g., 0.3-0.5) can be selected, providing important parameter selection guidance for adaptive threshold settings.
[0236] Step 6.4: Locate the key areas causing the deviation and provide a visual heatmap of the anomalies;
[0237] In embodiments according to this application, heatmap generation is based on feature importance analysis and employs Gradient Weighted Class Activation Mapping (Grad-CAM++) technology to calculate the contribution of each pixel to the final decision. This technology calculates the gradient of the loss function relative to the feature map through backpropagation, and then uses these gradients as weights to perform a weighted summation of the feature maps, ultimately obtaining a heatmap highlighting anomalous regions. It should be noted that visualizing anomaly heatmaps not only helps users understand the identification results but also provides important clues for subsequent evidence collection and analysis. Optionally, in some implementations, saliency detection technology can be combined to further enhance the visual prominence of anomalous regions, or a multi-level heatmap fusion strategy can be adopted to simultaneously display anomalous features at different semantic levels.
[0238] Step 7: Based on the degree of deviation, generate the identification results and interpretability analysis report of the input image;
[0239] Step 7.1: Generate a probability score for image authenticity and quantify the confidence level of the identification;
[0240] According to a preferred embodiment of this application, the probability score is based on the posterior probability calculated in step 4. The scoring is calibrated to ensure consistency and interpretability. Calibration employs a temperature scaling method to adjust the kurtosis of the probability distribution, making the scoring more reliable and intuitive. Furthermore, to facilitate user understanding, the system converts the probability scores into a percentage format from 0-100 and provides corresponding confidence level descriptions. It should be noted that in some implementations, an integrated scoring strategy can be used, combining the prediction results of multiple models to enhance the stability and reliability of the scoring, particularly making it more robust in handling boundary cases.
[0241] Step 7.2: Identify potential abnormal regions and annotate them on the original image;
[0242] In one embodiment of this application, the annotation process, based on the anomaly heatmap generated in step 6.4, converts the heatmap into anomaly region masks through thresholding, then overlays the masks onto the original image, using different colors to mark the anomaly intensity. Optionally, the system can also automatically extract the boundary contours of the anomaly regions and sort the regions according to the degree of anomaly, highlighting the most suspicious parts. Furthermore, in some implementations, different color codes can be used for different types of anomalies (such as physical anomalies and semantic anomalies) to make the annotation results more informative and help users quickly understand the nature of the anomalies.
[0243] Step 7.3 provides an analysis of the cause of the anomaly;
[0244] According to embodiments of this application, the anomaly cause analysis module interprets anomaly regions from both physical and semantic dimensions. Physical dimension analysis includes inconsistent lighting, incorrect shadows, and perspective distortion; semantic dimension analysis includes contradictory object relationships, illogical events, and violations of common sense. The analysis results are output in structured text format, providing intuitive understanding for both professionals and ordinary users. Furthermore, the system automatically sorts the anomaly causes by severity, prioritizing the display of issues that have the greatest impact on image realism. It should be understood that in other implementations, an interactive anomaly analysis interface can be introduced, allowing users to obtain more detailed analysis explanations by clicking on specific areas, or to explore results under different anomaly detection thresholds by adjusting parameters.
[0245] Step 7.4: Generate a detailed identification report;
[0246] In a preferred embodiment according to this application, the report includes the following: basic image information, authenticity probability score, visualization of abnormal regions, physical analysis results, semantic analysis results, comprehensive evaluation conclusions, and recommendations. The report adopts a hierarchical structure, allowing different users to view information of varying levels of detail as needed. Furthermore, in practical applications, customized report templates can be provided based on user identity and application scenarios to meet the needs of professionals in different fields. Optionally, the report generation system can also support multiple output formats, such as PDF, HTML, or JSON, to adapt to different usage scenarios; simultaneously, for applications requiring batch processing, summary reports can be generated, displaying statistical analysis of identification results and anomaly pattern analysis for multiple images.
[0247] like Figure 4As shown, the accuracy of the physical-semantic double helix theory compared with traditional physical verification and traditional semantic inspection methods in identifying ordinary GAN images and high-quality GAN images is illustrated. The results show that the physical-semantic double helix theory maintains a high accuracy of over 90% on both types of images, while the accuracy of traditional methods decreases when dealing with high-quality GAN images. This comparison clearly demonstrates the advantages of the proposed method in handling complex AI-generated images, especially its stable performance when dealing with high-quality forged images.
[0248] like Figure 5 As shown, the identification accuracy of the physical-semantic double helix theory and traditional methods on different types of images (news images, portrait images, landscape images, indoor scenes, and complex events) is demonstrated. The physical-semantic double helix theory maintains high accuracy across all image types, exhibiting good versatility, while traditional methods perform worst on complex event images. This result proves that the proposed method has excellent cross-domain adaptability and stability, and can provide reliable image identification services for different application scenarios.
[0249] like Figure 6 As shown, the overall performance of the physical-semantic double helix theory and traditional methods is compared across six dimensions: accuracy, robustness, interpretability, adaptability, computational efficiency, and generality. The physical-semantic double helix theory outperforms traditional methods in all dimensions except computational efficiency, with its advantages being most pronounced in interpretability and adaptability. This comprehensive performance comparison clearly demonstrates the full technical advantages of the proposed method, providing strong performance support for its deployment and promotion in practical applications.
[0250] An artificial intelligence-based digital image authentication system, used to execute the aforementioned artificial intelligence-based digital image authentication method, includes:
[0251] The three-layer fusion neural network architecture module is used to build a neural network architecture that includes a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer.
[0252] The physics analysis module is used to extract the input image from the network layer analysis using physical parameters and generate a set of physical feature vectors.
[0253] The semantic understanding module is used to process the input image using the semantic understanding network layer, generating semantic feature vectors and relationship graphs.
[0254] The cross-attention module is used to implement the physical-semantic cross-attention mechanism to enhance the bidirectional information flow between physical features and semantic features;
[0255] The joint distribution construction module is used to construct a physical-semantic joint distribution based on physical feature vectors and semantic feature vectors using the variational inference engine layer.
[0256] The distribution analysis module is used to analyze the degree of deviation between the physical and semantic joint distribution of the input image and the pre-established standard distribution of natural images;
[0257] The results output module is used to generate identification results and interpretability analysis reports for the input image based on the degree of deviation.
[0258] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. An artificial intelligence-based digital image authentication method, characterized by, The method comprises the following steps: A three-layer fusion neural network architecture is constructed, including a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer; The physical parameter extraction network layer is used to analyze the input image and generate a set of physical feature vectors; The semantic understanding network layer is applied to process the input image to produce semantic feature vectors and a relationship graph; Based on the physical feature vectors and the semantic feature vectors, the variational inference engine layer is used to construct a physical-semantic joint distribution; A physical-semantic cross-attention mechanism is implemented to enhance the bidirectional information flow between physical features and semantic features; The deviation degree of the physical-semantic joint distribution from a pre-established standard distribution of natural images is analyzed; Based on the deviation degree, an identification result and an explainability analysis report of the input image are generated; The step of implementing the physical-semantic cross-attention mechanism comprises the following steps: A physical-to-semantic attention map is constructed to guide the semantic understanding to focus on the abnormal areas; A semantic-to-physical attention map is constructed to make the physical analysis preferentially process the semantically important areas; Through alternating iterative optimization of the two attention maps, deep fusion of physical and semantic features is achieved; The fused feature representation is extracted for subsequent anomaly detection.
2. The method according to claim 1, wherein the construction method of the physical parameter extraction network layer comprises the following steps: A deep convolutional neural network structure is adopted, including a feature extraction module and a physical parameter estimation module; Multi-scale convolution filters are applied to the input image to extract image features of different scales; Multi-scale features are fused through a feature pyramid structure to form a comprehensive feature map; A dedicated physical parameter estimation module is applied to infer the illumination model parameters, surface reflection characteristics, shadow consistency, and perspective relationship from the comprehensive feature map; The uncertainty estimation of the physical parameters is calculated to provide confidence information for subsequent probabilistic reasoning. The construction method of the semantic understanding network layer comprises the following steps:
3. The method of claim 1, wherein the method further comprises: A Transformer-based visual model is used to extract the semantic features of the input image; An object detection and segmentation module is applied to identify key objects and regions in the input image; An object relationship graph is constructed to represent the spatial and functional relationships between objects in the input image; An event reasoning module is applied to analyze the events in the scene and their rationality. The construction method of the variational inference engine layer comprises the following steps:
4. The method of claim 1, wherein the method is based on artificial intelligence. The physical feature vectors and the semantic feature vectors are input into a variational autoencoder to be mapped to a shared latent space; Conditional probability modeling is applied in the latent space to calculate the physical compliance probability and the semantic consistency probability; The conditional probabilities are integrated into a joint probability distribution through a Bayesian inference framework; The Monte Carlo integration method is applied to calculate the integral to obtain the authenticity probability estimate. The step of analyzing the deviation degree of the physical-semantic joint distribution from the pre-established standard distribution of natural images comprises the following steps:
5. The method of claim 1, wherein the method further comprises: A standard distribution model based on large-scale real image data is established as a reference benchmark; The Kullbac Leibler divergence between the joint distribution of the input image and the standard distribution is calculated to quantify the deviation degree; An adaptive threshold is set to judge the authenticity of the input image based on the deviation degree; The key areas causing the deviation are located, and a visualized abnormal heat map is provided. 6. The method of claim 1, wherein the method further comprises: The step of generating the identification result and the explainability analysis report of the input image comprises: generating a probability score of the authenticity of the input image to quantify the identification confidence; identifying potential abnormal areas and marking them on the input image; providing an abnormal reason analysis to explain the violation of physical laws and semantic logic; generating a detailed identification report including various indicators and comprehensive evaluation results.
7. The method of claim 3, wherein the method further comprises: determining a number of the plurality of images; and determining a number of the plurality of images to be displayed on the display device based on the number of the plurality of images. The step of integrating conditional probabilities into joint probability distribution through the Bayesian inference framework comprises: temperature scaling calibration of the physical semantic joint distribution to improve the reliability of probability estimation by adjusting the temperature parameter; establishing a multi-scale interaction strategy using a variational autoencoder to realize the interaction of physical feature vectors and semantic feature vectors at different abstraction levels through cross-layer connection; applying a feature enhancement gate unit to control the strength of information interaction between physical feature vectors and semantic feature vectors.
8. The method of claim 1, wherein the method further comprises: The collaborative working mechanism of the physical parameter extraction network layer and the semantic understanding network layer comprises: pre-classifying the input image through an image type classifier to automatically adjust the parameter configuration of the multi-scale convolution filter and the attention weight of the Transformer-based visual model according to the classification result; building incremental learning modules in the physical parameter extraction network layer and the semantic understanding network layer to enable the deep convolutional neural network structure and the Transformer-based visual model to continuously learn from new types of fake technology samples and dynamically update the network weights; establishing a multi-modal fusion mechanism to fuse the illumination model parameters and surface reflection characteristics output by the physical parameter estimation module with the key object information identified by the object detection and segmentation module, and simultaneously integrating image metadata for comprehensive analysis.
9. An artificial intelligence-based digital image authentication system, characterized by, An artificial intelligence-based digital image identification method according to any one of claims 1-8, comprising: a three-layer fusion neural network architecture module for building a neural network architecture comprising a physical parameter extraction network layer, a semantic understanding network layer, and a variational inference engine layer; a physical analysis module for analyzing the input image using the physical parameter extraction network layer to generate a set of physical feature vectors; a semantic understanding module for processing the input image using the semantic understanding network layer to produce semantic feature vectors and relationship graphs; a cross-attention module for implementing a physical-semantic cross-attention mechanism to enhance bidirectional information flow between physical features and semantic features; a joint distribution construction module for constructing a physical-semantic joint distribution using the variational inference engine layer based on the physical feature vectors and the semantic feature vectors; a distribution analysis module for analyzing the deviation of the physical-semantic joint distribution of the input image from the pre-established standard distribution of natural images; a result output module for generating the identification result and the explainability analysis report of the input image based on the deviation degree.
Citation Information
Patent Citations
Forgery image detection method, electronic equipment and computer readable storage medium
CN115984178A
Intelligent fruit quality grading system and method based on multi-mode variational auto-encoder
CN120375358A