Bionic machine vision method for large model generated image discrimination

CN122821322APending Publication Date: 2026-09-25ANHUI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202611294505.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0011]以上几种现有技术,包括以下几个方面的共性缺陷:(1)均为纯数据驱动的一次性全局分类器,未利用人类视觉认知的渐进式搜索策略,缺乏可解释性;(2)对所有图像区域同等处理,未聚焦伪造痕迹显著区域,计算资源分配不合理;(3)对简单样本和困难样本使用相同计算量,无法根据样本难度自适应调整推理深度;(4)跨生成器泛化能力不足,面对训练集之外的新型生成器时性能显著下降;(5)即便涉及注视,也仅将其作为辅助正则化信号或静态热图,未把注视行为作为检测核心驱动机制

Benefits of technology

[0044]本发明公开了一种面向大模型生成图像鉴别的仿生机器视觉方法,属于计算机视觉与多媒体安全技术领域。该算法首先构建面向大模型生成图像鉴别任务的眼动轨迹数据集,采集人类观察者在图像真伪判断过程中的注视位置、注视顺序和注视持续时间;然后构建根据待鉴别图像生成眼动轨迹的轨迹生成模型,使其学习人类审查图像时的注视轨迹分布,并将生成轨迹作为注视策略模块的训练监督;在推理阶段,根据预测注视轨迹确定当前注视位置,并围绕该位置提取中心凹、旁中心凹和近外周多尺度窗口特征;将多尺度窗口特征与全局语义特征融合得到当前注视位置的证据向量,并通过序列证据累积和自适应停止机制输出图像鉴别结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821322A_ABST
    Figure CN122821322A_ABST
Patent Text Reader

Abstract

The application discloses a biomimetic machine vision method for large model generated image discrimination, and belongs to the technical field of computer vision and multimedia security. First, an eye movement trajectory dataset for large model generated image discrimination tasks is constructed, and the gaze position, gaze sequence and gaze duration of human observers in the process of judging image authenticity are collected; then, a trajectory generation model for generating eye movement trajectories according to the image to be discriminated is constructed, and the generated trajectory is used as the training supervision of the gaze strategy module; in the inference stage, the current gaze position is determined according to the predicted gaze trajectory, and multi-scale window features are extracted around the position; the multi-scale window features and global semantic features are fused to obtain the evidence vector of the current gaze position, and the image discrimination result is output through the sequence evidence accumulation and adaptive stopping mechanism. Through the eye movement trajectory data, the trajectory generation model and the window evidence extraction mechanism guided by the trajectory, efficient and interpretable discrimination of large model generated images is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and multimedia security technology, specifically involving the authenticity identification of large model-generated images (large model-generated images), and in particular a bionic machine vision method for the identification of large model-generated images. Background Technology

[0002] Large-scale image generation is a technology that uses text descriptions to enable AI to automatically draw images. Mainstream tools include Midjourney, Dall-E 3, and Stable Diffusion. AI software learns patterns from massive amounts of images and then, based on input text descriptions or original images, gradually generates relevant images from noise. The mainstream technique is called diffusion modeling. Core functions include text-to-image generation, image-to-image generation, and style transfer.

[0003] With the rapid development of large-model image generation technologies such as diffusion models and generative adversarial networks (GANs), highly realistic large-model generated images are widely used in content creation, advertising, film and television, and other fields. However, the abuse of large-model generated images has also brought serious multimedia security problems, such as fake news images, deepfake faces, and synthetic ID photos.

[0004] Existing large model-generated image discrimination methods mainly include the following:

[0005] (1) Frequency domain feature-based detection method: High-frequency noise features of images are extracted using Discrete Cosine Transform (DCT) and Spatial Rich Model (SRM), and the unique artifacts left in the frequency domain by the large model-generated image are analyzed. Related patents include: Chinese invention patent CN115880749A, "Multi-domain Feature Fusion Deep Forgery Detection," and Chinese invention patent CN120032234A, "Multi-domain Feature Fusion Based on DINOv2 and DCT Transform." The shortcomings of these existing technologies are: sensitivity to post-processing operations such as JPEG compression, scaling, and filtering, and limited generalization ability.

[0006] (2) End-to-end detection methods based on deep learning: These methods directly learn features for distinguishing between real and fake data using CNN (Convolutional Neural Network) or ViT (Vision Transformer) as the backbone network. Examples include classifiers based on ViT-B / 16 fine-tuning, the AIDE method, and the Chinese invention patent CN116704580A entitled "Deepfake Detection Based on Transformer". The shortcomings of these existing technologies are that they are essentially one-time global classifiers that process the entire image uniformly, which cannot simulate the gradual cognitive process of human experts focusing on suspicious areas and lacks interpretability.

[0007] (3) Reconstruction-based detection methods: DIRE (Diffusion Reconstruction Error) distinguishes the differences before and after reconstruction using a diffusion model, while AEROBLADE (a training-free latent diffusion image detection method based on autoencoder reconstruction error) utilizes the latent diffusion model to reconstruct the error. The drawbacks of these methods are: extremely high computational cost (measured single-image latency of approximately 161.4 ms for DIRE and approximately 148.7 ms for AEROBLADE), significantly higher than the single-image inference latency of approximately 49.6 ms for the method of this invention, making it difficult to meet real-time requirements and having limited effectiveness on images generated by non-diffusion models.

[0008] (4) Reconstruction-based detection methods: DIRE uses a diffusion model to distinguish the differences before and after reconstruction, while AEROBLADE utilizes a latent diffusion model to reconstruct errors. These methods have high computational overhead. The measured single-image delay for DIRE is approximately 161.4 ms, and for AEROBLADE it is approximately 148.7 ms, which is significantly higher than the single-image inference delay of approximately 49.6 ms for the method of this invention, making it difficult to meet the requirements of real-time or near-real-time identification. At the same time, these methods have limited detection performance on images generated by non-diffusion models.

[0009] (5) Detection method based on visual base models: This method uses large-scale pre-trained visual base models such as SigLIP2 and DINOv2 to extract general features, which are then connected to a lightweight classification head. The drawback of this method is that it typically concatenates the features of the two encoders directly and makes a one-time global judgment, lacking a selective attention mechanism for local regions and failing to fully utilize local artifact information.

[0010] (6) Forgery detection methods involving gaze features: GazeForensics (Neural Networks, 2024) uses a pre-trained 3D gaze estimation model to obtain gaze representations and uses MSE (Mean Squared Error) loss to regularize the detection backend; Chinese invention patent CN113627256B, "Forgery Video Detection Method and System Based on Blink Synchronization and Binocular Movement Detection", discloses a biological signal based on blink synchronization and eye movement. Gaze modeling directions include ScanDiff (ICCV 2025), DiffGaze, and Chinese invention patent CN112507799B, which discloses an "Image Recognition Method, MR Glasses and Medium Based on Eye-Motion Gaze Point Guidance", mentioning technologies such as eye-motion gaze point guidance for image recognition. The shortcomings of these existing technologies are as follows: the aforementioned gaze-based methods analyze whether the gaze direction of the "subject being photographed" is consistent, which belongs to passive biometric analysis; or they only use gaze as an auxiliary regularization signal / static heatmap, none of which take the "gazing behavior of the observer (human reviewer)" as the core driving mechanism for detection, nor do they involve sequential evidence accumulation and adaptive stopping. The scan path generation methods are all geared towards general scenarios and are not combined with evidence accumulation decision-making.

[0011] The above-mentioned existing technologies share the following common defects: (1) They are all one-time global classifiers driven by pure data, which do not utilize the progressive search strategy of human visual cognition and lack interpretability; (2) They treat all image regions equally, do not focus on areas with significant forgery traces, and have unreasonable allocation of computational resources; (3) They use the same amount of computation for simple and difficult samples and cannot adaptively adjust the inference depth according to the sample difficulty; (4) They have insufficient generalization ability across generators and their performance drops significantly when facing new generators outside the training set; (5) Even if gaze is involved, it is only used as an auxiliary regularization signal or static heatmap and gaze behavior is not used as the core driving mechanism for detection.

[0012] Therefore, there is an urgent need to study a technical solution for accurately and efficiently detecting images generated by large models, so as to comprehensively solve the common defects mentioned above. Summary of the Invention

[0013] To avoid the shortcomings of the existing technologies, this invention provides a biomimetic machine vision method for identifying images generated by large models. This method effectively integrates local artifact features with global semantic features, uses eye-tracking trajectories to guide the model to focus on key evidence regions, and improves the accuracy, interpretability, and reasoning efficiency of identifying images generated by large models.

[0014] The present invention adopts the following technical solution to solve the technical problem.

[0015] The biomimetic machine vision method for image recognition based on large model generation of the present invention includes the following steps:

[0016] Step 1: Scan Path Sequence Acquisition Step; Process the image to be identified to obtain the scan path sequence G = (g1, g2, …, g max );

[0017] Step 2: Image feature extraction step; extract local artifact features F from the image to be identified. l and global semantic features F g ;

[0018] Step 3: Fixation Path Prediction Step; The fixation strategy module obtains the predicted fixation location through a 4-layer causal Transformer decoder. and path termination probability s t ;

[0019] Step 4: Evidence reading step; The local artifact features F obtained in Step 2... l and global semantic features F g By fusing the data, we obtain the evidence vector e. t ;

[0020] Step 5: Evidence accumulation step; through the obtained evidence vector e t The hidden state h of the evidence accumulator in the evidence accumulation module. t The system is updated, and then the updated evidence accumulator is used to calculate the output discrimination confidence level p. t And identification stop confidence level s c ;

[0021] Step 6: Adaptive stopping and result output step; based on the discrimination confidence level p obtained in Step 5. t And identification stop confidence level s c Determine whether the decision should be stopped; output the identification result based on whether the mandatory stop conditions are met.

[0022] The biomimetic machine vision method for image recognition based on large models in this invention is also characterized by:

[0023] Furthermore, the process of obtaining the scan path sequence G in step 1 includes the following steps;

[0024] Step 11: Image meshing step; After unifying the resolution of the image to be identified, divide it into uniformly sized grid cells;

[0025] Step 12: Grid cell indexing confirmation step; A one-dimensional grid index k is used to identify the position of each grid cell in the image to be identified;

[0026] Step 13: Residual offset calculation steps; calculate the residual offsets Δx and Δy of the gaze point within the grid cell;

[0027] Step 14: Fixation Duration The calculation steps;

[0028] Step 15: Representation of the scan path sequence G; Represent the complete scan path of the image to be identified as the sequence G = (g1, g2, …, g max ).

[0029] Furthermore, in the image meshing process of step 11, bicubic interpolation is used to upsample and enlarge the image, and region average pooling is used to downsample and shrink the image.

[0030] Furthermore, in step 14, the natural logarithm is used to calculate the fixation duration. .

[0031] Furthermore, in step 2, a DINOv2 visual encoder is used to extract local artifact features F. l The SigLIP2-NaFlex visual encoder is used to extract global semantic features F. g .

[0032] Furthermore, in step 3, the causal Transformer decoder includes a causal self-attention layer, a cross-attention layer, a feedforward network, and a pre-layer normalization layer (Pre-LayerNorm).

[0033] Furthermore, in step 4, during evidence reading, a local window of three scales is used to obtain the evidence vector e. t .

[0034] Furthermore, in step 5, a two-layer selective state-space module is used to calculate the discrimination confidence level p. t And identification stop confidence level s c .

[0035] The present invention also discloses a biomimetic machine vision system for image recognition of large models, including an input representation module, a visual feature extraction module, a gaze strategy module, an evidence reading module, an evidence accumulation module, and an adaptive stopping decision module;

[0036] The input representation module is used to convert the image to be identified into a scan path sequence G;

[0037] The visual feature extraction module is used to extract local artifact features F from the image to be identified. l and global semantic features F g ;

[0038] The gaze strategy module is used to obtain the predicted gaze position through a 4-layer causal Transformer decoder. and path termination probability s t ;

[0039] The evidence reading module is used to fuse the local artifact features F l Global semantic features F g and predicting gaze location Obtain the evidence vector e t ;

[0040] The evidence accumulation module is used to accumulate evidence based on evidence vector e. t Predicting gaze position And the hidden state h from the previous moment t-1 Obtain the identification confidence level p t And identification stop confidence level s c ;

[0041] The adaptive stopping decision module is used to determine the stopping confidence level p based on the identification confidence level p. t And identification stop confidence level s c To determine whether the decision should be stopped.

[0042] The biomimetic machine vision system for generating and identifying images from large models also includes a diffusion scanning path teacher module.

[0043] Compared with existing technologies, the beneficial effects of this invention are reflected in:

[0044] This invention discloses a biomimetic machine vision method for large-model generated image identification, belonging to the fields of computer vision and multimedia security technology. The algorithm first constructs an eye-tracking trajectory dataset for large-model generated image identification tasks, collecting the gaze position, gaze order, and gaze duration of human observers during image authenticity judgment. Then, it constructs a trajectory generation model that generates eye-tracking trajectories based on the image to be identified, allowing it to learn the gaze trajectory distribution when humans review images, and uses the generated trajectory as training supervision for the gaze strategy module. In the inference phase, the current gaze position is determined based on the predicted gaze trajectory, and multi-scale window features (fovea, parafovea, and near-periphery) are extracted around this position. The multi-scale window features are fused with global semantic features to obtain the evidence vector for the current gaze position, and the image identification result is output through sequential evidence accumulation and an adaptive stopping mechanism.

[0045] This invention achieves efficient and interpretable identification of images generated by large models by using eye-tracking trajectory data, trajectory generation models, and trajectory-guided window evidence extraction mechanisms.

[0046] During inference, the gaze strategy module bases its decisions on the global semantic features F. gand historical gaze sequence G <t Progressively predict the current gaze position g t The evidence reading module predicts the gaze location g. t Centered on local artifact features F l Local evidence at three scales—central fovea, paracentral fovea, and near periphery—is extracted and compared with global semantic features F. g The evidence vector e at the current gaze position is obtained by fusion. t The evidence accumulation module combines the evidence vector e t Predicting gaze location embedding Emb(g) t ) and the hidden state h from the previous moment t-1 Update the current hidden state h t And output the identification confidence level p. t And identification stop confidence level s c When the accumulated evidence is insufficient to output a discrimination result, the current predicted gaze position g is set. t Add to historical gaze sequence G <t+1 The gaze strategy module continues to predict the next gaze location; when the accumulated evidence meets the stopping condition, the gaze search terminates and the identification result is output. This forms a closed-loop identification process of "gaze location prediction - local evidence reading - sequence evidence accumulation - stopping judgment - next gaze location prediction", which enables the model to prioritize searching image regions that contribute more to the authenticity identification and adaptively adjust the number of gaze locations and inference depth according to the sample difficulty. This solves the problems of the separation between scan path generation and identification decision, insufficient utilization of local artifacts, unreasonable allocation of computing resources, and lack of interpretability of the identification process in the existing technology.

[0047] This invention comprehensively utilizes techniques such as eye-tracking and gaze modeling, deep learning sequence decision-making, diffusion generation models, state space models (SSM), and teacher-student distillation to redefine the large-model generated image identification problem as a "sequence visual search problem simulating the visual scrutiny behavior of human censors." Typical application scenarios include: image authenticity verification, content security review, news and social media forensics, platform risk control, and recognition of synthetic documents / counterfeit goods images. It has broad application prospects in education and entertainment, cultural media, copyright protection, artistic creation, and forensic identification.

[0048] The biomimetic machine vision method for image identification generated by large models of the present invention has the advantages of being able to integrate local artifact features and global semantic features, simulating human review behavior to achieve interpretable detection, adaptively allocating inference computing power, being robust to post-processing perturbations, and having high identification accuracy. Attached Figure Description

[0049] Figure 1This is a flowchart of the biomimetic machine vision method for image recognition based on large models according to the present invention. Figure 2 This is a module logic diagram of the bionic machine vision method for image recognition based on large models according to the present invention.

[0050] Figure 3 This is an adaptive stopping decision flowchart of the biomimetic machine vision method for image recognition based on large models according to the present invention.

[0051] Figure 4 This is a schematic diagram of feature extraction in the visual feature extraction module of the system of the present invention;

[0052] Figure 5 This is a schematic diagram of the gaze strategy module of the system of the present invention;

[0053] Figure 6 This is a schematic diagram of the internal structure of a single-layer causal Transformer decoder in the gaze strategy module of the system of the present invention;

[0054] Figure 7 This is a schematic diagram of the evidence reading module of the system of the present invention;

[0055] Figure 8 This is a schematic diagram of the evidence accumulation module of the system of the present invention;

[0056] Figure 9 This is a schematic diagram of the selective state space module in the evidence accumulation module of the system of the present invention;

[0057] Figure 10 This is a diagram showing the component relationships between the training and inference phases of this invention;

[0058] Figure 11 This is a schematic diagram of the multi-scale gaze window of the present invention.

[0059] Figure 12 This is a flowchart of the human gaze data acquisition experiment of the present invention.

[0060] Figure 13 This is the confidence curve of the model on the image generated by the large model during the process of increasing the number of gaze positions in this invention.

[0061] The present invention will be further described below through specific embodiments and in conjunction with the accompanying drawings. Detailed Implementation

[0062] See Figures 1 to 13 The biomimetic machine vision method for image recognition based on large model generation of the present invention mainly includes the following 5 steps:

[0063] Step 1: Scan Path Sequence Acquisition Step; Process the image to be identified to obtain the scan path sequence G = (g1, g2, …, g max );

[0064] The input representation module processes the image to be identified step by step and finally obtains the scan path sequence G;

[0065] Step 2: Image feature extraction step; extract local artifact features F from the image to be identified. l and global semantic features F g ;

[0066] Step 3: Fixation Path Prediction Step; The fixation strategy module obtains the predicted fixation location through a 4-layer causal Transformer decoder. and path termination probability s t ;

[0067] Step 4: Evidence reading step; The local artifact features F obtained in Step 2... l and global semantic features F g By fusing the data, we obtain the evidence vector e. t ;

[0068] Step 5: Evidence accumulation step; through the obtained evidence vector e t The hidden state h of the evidence accumulator in the evidence accumulation module. t The system is updated, and then the updated evidence accumulator is used to calculate the output discrimination confidence level p. t And identification stop confidence level s c ;

[0069] Step 6: Adaptive stopping and result output step; based on the discrimination confidence level p obtained in Step 5. t And identification stop confidence level s c Determine whether the decision should be stopped; output the identification result based on whether the mandatory stop conditions are met.

[0070] like Figure 1 The flowchart of the biomimetic machine vision method for image recognition of large models in this invention is as follows: instead of relying on one-time black box classification, the detection is redefined as a sequential visual search problem, simulating the visual examination behavior of human examiners who "see-read-accumulate-stop" while thinking and accumulating evidence. It has the advantages of being interpretable, easy to deploy, and able to adaptively allocate computing resources. Figure 1This paper illustrates the overall flow of the biomimetic machine vision method for image identification based on large model generation, as described in this invention. The image to be identified is first processed into a grid, forming a discrete representation composed of image patches. Then, visual feature extraction is used to obtain image features for subsequent analysis. Based on this, a gaze path prediction module simulates the visual search process of a human reviewer, progressively outputting the next gaze position. An evidence reading module reads local evidence from the corresponding region based on the current gaze position and generates an evidence vector. Finally, an adaptive stopping module determines whether to output the identification result based on the accumulated evidence. This flow embodies a sequential identification mechanism of "see-read-accumulate-stop".

[0071] like Figure 1 and Figure 2 In the biomimetic machine vision method for image recognition based on large models of the present invention, the overall data flow is as follows: image to be recognized → input representation and scan path marker-based discretized representation → visual feature extraction module extracts global semantic features F in parallel. g and local artifact features F l →Causal gaze strategy module is based on F g Autoregressive prediction of next gaze location → Evidence reading module retrieves data from F based on the current gaze location. l Reading multi-scale local evidence and comparing it with F g The evidence vector e is obtained by performing 4-head cross-attention fusion on the corresponding global token. t →The evidence accumulation module uses SSM to collect e t Accumulate into hidden state h t And output the current identification confidence level p. t With identification stop confidence level s c →The adaptive stopping decision module determines whether the stopping condition is met. If it is, the final result is output; otherwise, it returns to the causal gaze strategy module to predict the next gaze position and enter the next loop. In practice, the above process requires an average of only 31±4 gaze positions to complete the detection, with a single-image inference latency of approximately 49.6ms.

[0072] Figure 2 The module logic relationship of the algorithm and system of the present invention is illustrated. The system consists of an input representation module, a visual feature extraction module, a gaze strategy module, an evidence reading module, an evidence accumulation module, and an adaptive stopping module in sequence. The input representation module is responsible for converting the image to be identified into a gridded image representation; the visual feature extraction module further extracts local artifact features and global semantic features; the gaze strategy module predicts the next gaze position and path termination probability s based on the historical gaze sequence. t The evidence reading module reads local evidence near the predicted location and fuses it with global semantic information; the evidence accumulation module writes each piece of evidence into the hidden state; and the adaptive stopping module outputs the final identification result when the threshold condition is met.

[0073] In specific implementation, the process of obtaining the scanning path sequence G in step 1 includes the following steps;

[0074] Step 11: Image meshing step; After unifying the resolution of the image to be identified, divide it into uniformly sized grid cells;

[0075] Step 12: Grid cell indexing confirmation step; A one-dimensional grid index k is used to identify the position of each grid cell in the image to be identified;

[0076] Human gaze data is only needed during the training phase, not during the inference phase. The gaze coordinate mapping process is as follows:

[0077] Human-normalized gaze coordinates are denoted as (u, v)∈[0,1]², where u is the horizontal relative position and v is the vertical relative position. The superscript "²" in [0,1]² denotes the Cartesian product; [0,1]² represents the set [0,1] multiplied by itself, resulting in a region in two-dimensional space, represented as a unit square with side length 1 on a two-dimensional plane. The mapping formula for the one-dimensional grid index k is: ,in, This indicates rounding down, k∈[0, 3599]. The one-dimensional grid index k identifies the position of the grid cell in the image to be identified. In actual calculation, The maximum value is 44; The maximum value is 79; then k = 0, 1, 2, ... 3599; thus, the 3600 grid cells are mapped one-to-one with the 3600 numbers, forming a one-to-one mapping relationship between the integer values ​​in the grid cell → [0, 3599].

[0078] Step 13: Residual offset calculation steps; calculate the residual offsets Δx and Δy of the gaze point within the grid cell;

[0079] To preserve the precise position of the gaze point within the grid cell, the calculation process for the residual offsets Δx and Δy of the gaze coordinates (u, v) relative to the center of the grid cell is shown in the following formula (1); where Δx is the residual offset of u and Δy is the residual offset of v;

[0080] (1)

[0081] In formula (1), (Δx, Δy)∈[-0.5, 0.5)², and both Δx and Δy taking the value 0 indicates that the gaze point is exactly located at the center of the grid cell; where, This indicates rounding down. u represents the horizontal relative position of the human normalized gaze coordinates, and v represents the vertical relative position of the human normalized gaze coordinates.

[0082] Step 14: Fixation Duration The calculation steps;

[0083] The original fixation duration *d* represents the time a subject's gaze remains at a fixation point, measured in milliseconds (ms), and is derived from eye-tracking data collected and output by the eye tracker. The original fixation duration *d* naturally follows a log-normal distribution, exhibiting a significant right skewness: most fixation times are concentrated in the 200–400 ms range, while a small number of extremely long fixation tails (hundreds to thousands of milliseconds) exist. These extremely high values ​​severely interfere with parametric tests such as t-tests, ANOVA, and linear mixed models. To compress the right-hand tail, improve distribution symmetry, and satisfy the statistical normality assumption, a logarithmic transformation is applied to the original fixation duration *d* to compress the tail. = ln(d + 1). Adding an offset of 1 is used to accommodate zero-value pseudo-sighting that may occur during the acquisition process. After the transformation, the variable distribution is closer to a normal distribution, reducing the statistical bias caused by extreme long-term sightings.

[0084] Step 15: Representation of the scan path sequence G; Represent the complete scan path of the image to be identified as the sequence G = (g1, g2, …, g max ).

[0085] Each fixation point is represented as a four-dimensional vector g = (k, Δx, Δy, ... ); where k gives the position of the gaze point in the image to be identified, i.e., which grid cell the gaze point is located in, and Δx and Δy give the offset between the position of the gaze point within the grid cell and the center point of that grid cell. The time the subject lingered at this fixation point is given. The complete scan path of the image to be identified is represented as the sequence G = (g1, g2, …, g max ); max is the maximum scan path length. In this invention, the maximum scan path length is set to max = 400 steps. When the actual obtained scan path sequence G is less than 400 steps, the scan path sequence G is expanded and padded to 400, and the scan path sequence G is a 400×4 sequence matrix.

[0086] Additionally, an EOS (sequence end marker) token is introduced into the one-dimensional grid index k, with the corresponding index 3600. Therefore, the output dimension of the patch logits (grid cell logical values) of the causal gaze policy network is 3601 (3600 grid cell indices + 1 EOS), i.e., [0, 3600]. The number 400 steps is mathematically derived: the 80×45 discretized grid has a total of 3600 grid cells, and the core concave window is a 3×3 grid cell (concave window); 3600 / 9=400 precisely allows the model to exhaustively scan the entire image with the highest resolution concave gaze without overlap, serving as a strict upper bound for the "cognitive budget". In practice, due to the adaptive stopping mechanism, most images are completed in far fewer than 400 steps (in the specific calculation process, the average number of steps is 31).

[0087] In specific implementation, during the image meshing process in step 11, bicubic interpolation upsampling is used to enlarge the image, and regional average pooling downsampling is used to shrink the image.

[0088] The image meshing process includes the following two steps:

[0089] 1. Obtain the image to be identified and adjust the resolution of the image to be identified to a standard image with a preset standard resolution, for example, the standard resolution is 1920×1080; the resolution adjustment includes two operations: (1) when the resolution of the image to be identified is less than 1920×1080, use bicubic interpolation upsampling to enlarge the image; (2) when the resolution of the image to be identified exceeds 1920×1080, use regional average pooling downsampling to shrink the image in order to retain high-frequency details;

[0090] 2. Subsequently, the standard image was uniformly divided into a discretized grid of 80×45 pixels, with each grid patch corresponding to 24×24 pixels, for a total of 3600 grid patches. The 24×24 pixel size matches the resolution of the standard eye-tracking experimental display, ensuring that the gaze coordinates correspond precisely to the grid patches;

[0091] In practice, step 14 uses the natural logarithm to calculate the fixation duration. .

[0092] In specific implementation, step 2 uses a DINOv2 visual encoder to extract local artifact features F. l The SigLIP2-NaFlex visual encoder is used to extract global semantic features F. g .

[0093] like Figure 4In this invention, the DINOv2 visual encoder and the SigLIP2-NaFlex visual encoder in the visual feature extraction module are used. The image to be identified is simultaneously input into both encoders, and the two encoders extract two complementary features of the image to be identified in parallel: a first feature and a second feature. The first feature is the local artifact feature F. l The first feature is used to read local evidence from the image to be identified; the second feature is the global semantic feature F. g This is used to characterize the overall semantic content, object relationships, and scene consistency information of the image to be identified. The visual feature extraction module employs two-stream visual encoding; two-stream refers to: [the encoding method is missing here, likely due to an error in the original text]. l Related local artifact flow and global semantic features F g Related global semantic flow.

[0094] Figure 4 The dual-stream structure of the visual feature extraction module is illustrated. The image to be identified is simultaneously input into a global semantic stream and a local artifact stream: the global semantic stream uses a SigLIP2-NaFlex visual encoder to extract the overall semantic consistency, scene logic, and high-level visual context of the image, and is then linearly projected to obtain the global semantic features F. g The local artifact stream employs a DINOv2 visual encoder to capture fine-grained artifacts such as local texture repetition, edge anomalies, unnatural colors, and generation noise. After linear projection, the local artifact features F are obtained. l Both types of features are unified to 256 dimensions for subsequent gaze prediction and evidence reading.

[0095] The local artifact stream employs the DINOv2 visual encoder architecture. DINOv2 is a self-supervised visual foundation model proposed by Meta AI, highly sensitive to local textures and fine-grained visual features (unnatural texture repetition, blurred edges, color anomalies, etc.), making it particularly suitable as a feature extractor for local artifact perception tasks. The DINOv2 visual encoder outputs dense patch tokens; these dense patch tokens are linearly projected to a uniform dimension D=256 to obtain the local artifact features F. l :F l ∈R {N×D} R {N×D} This represents the real space consisting of N D-dimensional vectors. N is the number of grid cells in the image to be identified, N=3600.

[0096] The global semantic stream employs a SigLIP2-NaFlex visual encoder architecture. SigLIP2 is a large-scale vision-language pre-trained model proposed by Google Research, and its NaFlex variant supports native resolution input, flexibly handling images with different aspect ratios without forced cropping / padding. The SigLIP2-NaFlex visual encoder consists of several visual Transformer blocks (multi-head self-attention + feedforward network + layer normalization), outputting the positional features of each grid cell to form a global semantic feature map. This global semantic feature map is then projected onto a uniform dimension D=256 through a linear projection layer to obtain the global semantic feature F. g :F g ∈R^{N×D}.

[0097] The global semantic stream captures overall content consistency and high-level semantics (object shape anomalies, scene logic errors, etc.); the local artifact stream captures local textures and low-level artifacts (diffusion model characteristic noise, GAN checkerboard artifacts, etc.). The global semantic stream and the local artifact stream constitute a complementary two-stream pattern, enabling the method and system of this invention to analyze the authenticity of the image to be identified from multiple dimensions.

[0098] In specific implementation, step 3 includes a causal self-attention layer, a cross-attention layer, a feedforward network, and a pre-layer normalization layer (Pre-LayerNorm).

[0099] Figure 5 This is a schematic diagram of the gaze strategy module of the system of the present invention; in the process of gaze path prediction, the causal gaze strategy is the core component of the inference stage. Given the current visual features and the historical gaze sequence, it autoregressively predicts the next gaze position. and path termination probability s t The path termination probability s t Used to determine whether the model should continue generating the next fixation point at a certain time point.

[0100] like Figure 5 During the operation of the gaze strategy module, the historical gaze sequence G before the current time t is used. <t As input, G <t After being processed layer by layer by a four-layer single-layer causal Transformer decoder structure, the hidden representation at the current time t is obtained. This hidden representation is further used to generate relevant parameters for the next gaze position through multiple output heads, including patch logits (grid cell logic values), residual offsets Δx and Δy, and gaze duration. and path termination probability s tThe patch logits, after being processed by Softmax, represent the probability distribution of the next gaze grid position. The residual offset is used to preserve the precise gaze position within the grid, and the path termination probability s... t Used to determine whether to continue generating subsequent fixations.

[0101] like Figure 5 The multi-head output of the 4-layer causal Transformer decoder (predicting four classes simultaneously at each time step) includes: ① patch logits (grid cell logical values, the original scores given by the model to all candidate grid patches at step t, a 3601-dimensional vector; after Softmax, it represents the probability distribution of the next gaze position; temperature sampling is used during the inference phase to balance exploratory and deterministic approaches); please explain patch logits; ② residual offset Δx and Δy predictions (output mean, limited to the range [-0.5, 0.5) after processing by the tanh function); ③ gaze duration (Mean and standard deviation of log-normal distribution); ④ Path termination probability s t (Scalar, normalized using Sigmoid); probability of path termination s t This is used to determine whether the current gaze path generation process should end.

[0102] The gaze strategy module employs a 4-layer causal Transformer decoder, which is lighter and faster inference compared to the 6-layer structure of the diffusion scan path teacher module.

[0103] Figure 6 This is the invention Figure 5 A schematic diagram of the internal structure of a single-layer causal Transformer decoder in the gaze strategy module. (See diagram below.) Figure 6 Each layer of the 4-layer causal Transformer decoder includes the following four layers: ① Causal self-attention layer (using a lower triangular mask, predicting the t-th step can only utilize the history of the previous t-1 steps, with 8 attention heads, 256 hidden dimensions, and 32 dimensions per head); ② Cross-attention layer with visual features (hidden state is Query, global semantic feature F). g Key / Value; Local artifact feature F l (Not used for policy prediction); ③ Feedforward network (two linear + GELU layers, 1024 hidden layers, 256 outputs); ④ Pre-LayerNorm.

[0104] like Figure 6After the historical gaze sequence is input, it first passes through a Pre-LayerNorm and a causal self-attention layer. The causal self-attention layer uses a lower triangular mask so that the prediction at step t at the current time t can only use the historical gaze sequence G before step t. <t Subsequently, it enters the second Pre-LayerNorm through a residual connection and performs cross-attention calculation with the visual features; where the hidden state serves as the Query, and the global semantic feature F obtained in step 2 is used as the Query. g The layers are used as keys and values; then they pass through residual connections, a pre-layer normalization network (FRN), and a feedforward network (FFN). The feedforward network increases the dimension from 256 to 1024 and then back to 256, using the GELU activation function. The final output is the representation of this layer, which is then processed by the next causal Transformer decoder.

[0105] like Figure 10 The diffusion scan path teacher module provides high-quality training supervision signals for causal gaze strategies by learning the gaze scan path distribution when human observers review images generated by large models. During gaze path prediction, the high-quality scan paths generated by the teacher serve as soft labels to guide the student. Distillation loss is used to calculate the KL divergence between the student's predicted patch distribution and the teacher's patch distribution, and a temperature parameter is used to soften the teacher's distribution, enabling the student to learn richer distribution information rather than simply mimicking the highest probability location.

[0106] The training strategy (progressive learning) for the 4-layer causal Transformer decoder in the gaze policy module consists of two phases:

[0107] (1) In the first stage (the first 5 epochs), teacher forcing is used – when predicting step t, the previous t-1 steps generated by the teacher are used as historical input to ensure stability in the early stage of training;

[0108] (2) In the second stage (after the 5th epoch), scheduled sampling is adopted - the student’s own prediction is used with probability ε and the teacher’s sequence is used with probability 1-ε. ε increases linearly from 0 to 0.3 (approximately 0.06 per epoch), so that the student gradually adapts to autoregressive reasoning and reduces the exposure bias between training and reasoning.

[0109] Figure 10The component relationships between the training and inference phases are illustrated. During the training phase, human gaze sequence training data is input into the diffusion scan path teacher module. This module learns the scan path distribution when humans review images generated by a large model and passes this distribution to the causal gaze policy student module through distillation supervision. During the inference phase, the diffusion scan path teacher module no longer participates in computation. The image to be identified is input into the causal gaze policy student module, which predicts the gaze position and then outputs the identification confidence p through the evidence reading and evidence accumulation modules. t With identification stop confidence level s c Finally, the adaptive stopping module provides the identification result.

[0110] In specific implementation, during step 4, when reading evidence, a local window of three scales is used to obtain the evidence vector e. t .

[0111] Figure 7 This is a schematic diagram of the evidence reading module of the system of the present invention; in the present invention, the evidence reading module is used to realize the local artifact feature F. l and global semantic features F g Feature fusion. The evidence reading module simultaneously receives local artifact features F. l Global semantic features F g and the predicted gaze position at the current time t For F l The evidence reading module reads local windows of three scales—13×13, 7×7, and 3×3—centered on the current gaze position, and obtains local evidence vectors of three scales (13×13 scale local evidence vector, 7×7 scale local evidence vector, and 3×3 scale local evidence vector) through bilinear sampling and average pooling; for F g The evidence reading module retrieves the global semantic token corresponding to the current gaze position; simultaneously, it predicts the gaze position. The embedding is used as the query. Subsequently, the local evidence vectors at three scales and the global semantic token are used together as the key / value pair in the four-head cross-attention calculation, outputting the evidence vector e at the current gaze position. t .

[0112] The human visual system has non-uniform spatial resolution, which is the fundamental biological basis for the existence of active vision and foveated vision. The spatial resolution of the human visual system is generally divided into three regions: (1) Fovea, with a visual field of about 1° to 2°, has the highest resolution and the smallest coverage; (2) Parafovea, with a visual field of about 2° to 5°; and (3) Near Periphery, with a visual field of about 5° to 30°, has the lowest resolution but the widest coverage.

[0113] The evidence reading module employs a multi-scale gaze reader, which simulates this multi-resolution characteristic using three concentric rectangular windows. Assuming a standard monitor viewing distance of 60cm, 1° of viewing angle corresponds to approximately 38 pixels. Based on this, three square windows are designed (centered on the currently predicted gaze grid cell), such as... Figure 11 As shown.

[0114] 1. Near-peripheral window: 13×13 grid units (312×312 pixels), corresponding to an approximately 8° viewing angle, providing a large range of context, which helps detect overall structural anomalies and large-scale semantic inconsistencies. The near-peripheral window is sensitive to contrast and motion changes. Once an anomaly is detected (such as local texture repetition or blurred edges), the abnormal area will be scanned as the target.

[0115] 2. Side-central concave window: The size is 7×7 grid units (168×168 pixels), corresponding to an approximately 5° viewing angle, providing medium-range local features, which helps to detect local texture anomalies and medium-scale artifacts; before the scan is performed, if multiple anomalies (targets) are detected, the side-central concave window will extract coarse-grained features in advance to determine the priority of the targets;

[0116] 3. Central concave window: The size is 3×3 grid cells (i.e. 72×72 pixels), corresponding to an approximately 2° viewing angle. The central concave window is finally aligned with the target, providing the finest local features, which helps to detect fine-grained artifacts (diffusion model characteristic noise, GAN checkerboard artifacts, etc.).

[0117] Figure 11 The spatial structure of the multi-scale gaze windows is shown. Each of the three windows is centered on the current gaze position and corresponds to a visual region: the near periphery, the parafoveal, and the foveal. The near periphery window is a 13×13 grid cell, corresponding to approximately 312×312 pixels and an approximately 8° field of view, used to acquire a large-scale context; the parafoveal window is a 7×7 grid cell, corresponding to approximately 168×168 pixels and an approximately 5° field of view, used to extract mesoscale local anomalies; and the foveal window is a 3×3 grid cell, corresponding to approximately 72×72 pixels and an approximately 2° field of view, used to read the highest resolution fine-grained artifact evidence.

[0118] The evidence reading process in this invention is based on three concentric rectangular gaze windows, and an adaptive weighted fusion of local evidence extracted from the three scale windows is performed through a 4-head cross-attention mechanism, enabling the model to dynamically adjust the importance of evidence at different scales according to the current image content, gaze position and global semantic information.

[0119] 1. For each scale window of the three concentric rectangular windows, from the local artifact feature F l The local window features are cropped with the current gaze position as the center (using bilinear grid sampling); then compressed into a 256-dimensional vector by average pooling; local artifact feature vectors at three scales are obtained.

[0120] 2. Simultaneously, from the global semantic feature F g Retrieve the global token corresponding to the current gaze position.

[0121] 3. Concatenate and stack the local artifact feature vectors of the three scales with the global token to form a matrix, which serves as the key / value pair for the 4-head cross-attention. The predicted gaze position is then obtained by inputting the current gaze position into the gaze strategy module. The embedding vector (including the grid cell index k and residual offsets Δx and Δy) is used as the Query; the outputs of the 4-head cross-attention are concatenated and then linearly projected, then connected with the Query residuals, and layer normalized to finally obtain the evidence vector e of the current gaze position. t ∈R 256 This step enables the system to adaptively weigh the importance of the three scales based on context. The current gaze position at time t is the t-th element g of the scan path sequence G. t .

[0122] In specific implementation, step 5 uses a two-layer selective state-space module to calculate the discrimination confidence level p. t And identification stop confidence level s c .

[0123] Figure 8 This is a schematic diagram of the evidence accumulation module of the present invention; the evidence accumulation module of the present invention uses a two-layer selective state space module (Selective SSM, based on Mamba architecture) as the evidence accumulator. The input of the evidence accumulation module is the current evidence vector e. t Predicting gaze position And the hidden state h from the previous moment t-1 Among them, predicting gaze location Includes one-dimensional grid index k, residual offsets Δx and Δy, and gaze duration. After embedding, the position representation Emb( The evidence accumulation module will accumulate the evidence vector e. t With location embedding Emb ( The hidden state h is formed by splicing the elements together to create a 512-dimensional representation, which is then reduced to 256 dimensions through linear projection and fed into a two-layer selective state space module (SSM) to update the hidden state h. t The updated hidden state h t Input two separate MLP headers: the Confidence Header for Authentication and the MLP for Discrimination. p And identification of stop confidence head MLP s Then, the confidence head MLP is used to determine the confidence level. p And identification of stop confidence head MLP s Output the discrimination confidence level p respectively t And identification stop confidence level s c .

[0124] Figure 9 This is a schematic diagram of the selective state space module in the evidence accumulation module of this invention; the selective state space module has linear time complexity and can efficiently process long sequences. Hidden state h t The update equation is shown in the following formula (2);

[0125] h t = SSM([e t Emb( )], h t-1 (2)

[0126] In formula (2), e t Emb( is the gaze position evidence vector at the current time t); () represents the predicted gaze position at the current time t. Embedded representation; h t Let h be the hidden state at time t. t-1 Let be the hidden state of the previous time step t-1. In the selective state space module, the state matrix B, the state matrix C, and the step size Δ are input-dependent (dynamically calculated based on the current input), thereby achieving selective memorization of the input sequence, which can retain evidence that contributes significantly to the discrimination and suppress redundant information in irrelevant regions.

[0127] like Figure 9 This is a single-layer structure for the Selective State-Space Module (SSM). The Selective State-Space Module (SSM) uses the current input x... t and the hidden state h from the previous moment t-1 Based on, where x t Let e ​​be the evidence vector t With predicted gaze location Embedded ( The 256-dimensional vector obtained by concatenating and then linearly projecting the vectors is x.t =Linear([e t Emb( ]), x t ∈R 256 Then according to x t Parameter B required for dynamic generation of state update t C t and Δ t Then perform a selective state update to obtain a new hidden state h. t Unlike sequence models with fixed parameters, selective SSM can determine which evidence should be remembered and which irrelevant information should be suppressed based on the current input. The updated state is then processed through linear projection and residual connections to form the output representation of this layer, which is then fed into the next layer of SSM or subsequent classification and stopping decision modules.

[0128] Each fixation point is represented as a four-dimensional vector g = (k, Δx, Δy, ... A four-dimensional vector (k, Δx, Δy, The mapping is a 256-dimensional vector, where the 256 dimensions include: the grid cell index k is embedded into 128 dimensions through a lookup table, the residual offset (Δx, Δy) is mapped into 64 dimensions through a linear layer, and the gaze duration. The linear layer mapping yields a 64-dimensional array, and the concatenation of the three dimensions followed by linear projection results in a 256-dimensional Emb. ).

[0129] Evidence vector e t (256-dimensional) and Emb( The 256-dimensional array is stitched together to a 512-dimensional array, then linearly projected down to 256 dimensions before being input into a two-layer selective state-space module (SSM) to update the hidden state h. t .

[0130] Based on updating the hidden state h t The post-evidence accumulator outputs two values: ① Confidence level p t = P(This image is generated from a large model | h) t The probability of determining an image as generated by a large model is calculated using two layers of MLP (hidden layer 128, ReLU) and Sigmoid (where the hidden state h is used to calculate the probability). t Input a two-layer MLP and calculate the discrimination confidence level p. t ); ② Identify the stopping confidence level s c The hidden state h is calculated using a separate two-layer MLP (hidden layer 128, ReLU) and Sigmoid algorithm to indicate whether the current accumulated evidence is sufficient to terminate the identification process and output the identification result. t Input the two-layer MLP and calculate the discrimination stopping confidence s.c .

[0131] The above-mentioned confidence head MLP p And identification of stop confidence head MLP s They have the same structure, but their parameters are not shared; they are independent of each other. Among them, the discrimination confidence level p... t The calculation process is shown in the following formula (3);

[0132] (3)

[0133] Identify stopping confidence level s c The calculation process is shown in the following formula (4);

[0134] (4)

[0135] In formulas (3) and (4), h t p represents the hidden state after updating at the t-th gaze position. t The confidence score represents the probability that the image to be identified is a generated image of the large model; s c To determine the stopping confidence level, we indicate whether the currently accumulated evidence is sufficient to terminate the identification process and output the identification result; The Sigmoid activation function is expressed as follows: ReLU is the linear rectified activation function, and its expression is: ; and These are the identification confidence head MLPs. p The weight matrix and bias terms of the first linear layer; and These are the identification confidence head MLPs. p The weight matrix and bias terms of the second linear layer; and These are the identification stop confidence head MLPs. s The weight matrix and bias terms of the first linear layer; and These are the identification stop confidence head MLPs. s The weight matrix and bias terms of the second linear layer.

[0136] In step 6, the adaptive stopping decision process is as follows: Figure 3 As shown. This invention employs a dual-threshold joint judgment mechanism. The adaptive stopping condition for the decision-making process is: the discrimination confidence level p. t >Probability threshold τ and identification stopping confidence s cThe stopping threshold τ_stop is triggered when both conditions are met simultaneously. The probability threshold τ = 0.9 (validation set ablation experiments show that this value achieves the best balance between discrimination accuracy and average number of fixation locations), and the stopping threshold τ_stop = 0.5 (standard binary classification decision threshold).

[0137] like Figure 3 At each gaze location, obtain the identification confidence level p from the evidence accumulation module. t And identification stop confidence level s c First, based on the aforementioned adaptive stopping conditions, it is determined whether both exceed a preset threshold simultaneously, such as a discrimination confidence greater than 0.9 and a discrimination stopping confidence greater than 0.5. If the conditions are met, the discrimination result is directly output. If the stopping conditions are not met, it is further determined whether the current number of gaze positions has reached the maximum scan path length (max). If the maximum scan path length (max) is reached, it is forcibly stopped and the result is output. If the maximum scan path length (max) is not reached, it returns to the gaze strategy module to predict the next gaze position and enters the next round of evidence reading and accumulation.

[0138] The decision-making process for stopping (each gaze location) includes the following four steps:

[0139] Step D1: Obtain the discriminative confidence level p from the evidence accumulator. t And identification stop confidence level s c ;

[0140] Step D2: Based on the identification confidence level p t And identification stop confidence level s c Determine whether the stopping condition is met. If it is met, proceed to step D3; otherwise, proceed to step D4.

[0141] Step D3: Trigger stop, based on the current identification confidence level p t Determine whether the image to be identified is a generated image of a large model, and output the determination result as the final result: p t If the value is greater than 0.5, it is determined to be an image generated by a large model; otherwise, it is determined to be a real photo.

[0142] Step D4: Determine whether the current number of gaze positions has reached the maximum scan path length (max). If it has, the output will be forcibly stopped. Otherwise, the gaze path prediction step in step 3 will be triggered to predict the next gaze position and enter the next loop.

[0143] In the process of identifying large model generated images, the method of the present invention typically reaches the stopping condition within about 10 to 15 steps for large model generated images with obvious features; for difficult samples after post-processing, it will take about 50 to 80 steps to collect sufficient evidence.

[0144] Extensive testing has demonstrated that the average number of fixation locations on the test set is 31 ± 4 steps, far fewer than the maximum number of steps (max = 400). The adaptive stopping mechanism effectively avoids overcomputation on simple samples.

[0145] In the process of testing the method and system of the present invention, a dedicated gaze-aware benchmark dataset was constructed to support training and evaluation. The construction process includes the following steps.

[0146] 1. Image Sources and Scale: A total of 3000 images are included, sourced from three sources: images generated by Kimi, images generated by Doubao, and real photos, with 1000 images from each source. Among them, the images generated by Kimi and Doubao serve as samples of images generated by the large model, while the real photos serve as comparison samples of real images, used to construct the positive and negative sample set for the large model-generated image discrimination task.

[0147] 2. Eye-tracking experiment: An EyeLink 1000 Plus eye tracker was used, with a sampling frequency of 1000Hz and an observation distance of approximately 60cm (corresponding to 13×13, 7×7, and 3×3 patch windows to the near peripheral, parafoveal, and foveal visual fields). Forty participants were recruited (20 women and 20 men, mean age 21.3±2.1 years). Before the experiment, participants were informed of the task content, experimental procedure, data usage, and privacy protection measures. Participants voluntarily participated and signed informed consent forms. Eye-tracking data was anonymized and used only for research. The experimental procedure included calibration / validation, a practice phase, and the formal experimental phase. Each participant completed 150 trials (divided into 3 trial cycles, spanning 5 categories: natural scenery / people / animals / food still life / daily life, presented in random order). The trial structure consisted of: drift check, viewing of stimulus images, authenticity judgment, and a 1-second blank screen interval.

[0148] 3. Data Labeling and Splitting: After quality control, approximately 6000 valid annotation trials with gaze were obtained. On average, each image was reviewed by two different participants to improve the reliability of the behavioral supervision signal. The dataset was divided into a training set (2100 images, 4200 trials), a validation set (450 images, 900 trials), and a test set (450 images, 900 trials) in a 70 / 15 / 15 ratio. During the partitioning, it was ensured that all trials for the same image belonged to the same subset to avoid data leakage.

[0149] Figure 12The experimental procedure for collecting human gaze data is illustrated. The experiment begins with eye-tracking calibration and verification, using a 9-point eye-tracking calibration. The calibration accuracy is assessed to ensure it is less than 0.5° of visual angle. If it fails, recalibration is performed. Once successful, the experiment proceeds to the practice phase. The practice phase includes five practice trials where participants view stimulus images, determine whether the images are generated by a large model or real human images, and record their results to familiarize themselves with the formal experimental procedure. The formal experimental phase presents 150 images categorized as natural landscapes, people, animals and plants, still life food, and everyday scenes. Each round of trials includes drift verification, presentation of stimulus images, judgment and recording of results, and presentation of a one-second blank image to avoid visual persistence. All trials are repeated.

[0150] The present invention also discloses a biomimetic machine vision system adapted to the above method for image recognition of large models, characterized in that it includes an input representation module, a visual feature extraction module, a gaze strategy module, an evidence reading module, an evidence accumulation module, and an adaptive stopping decision module.

[0151] The input representation module is used to convert the image to be identified into a scan path sequence G;

[0152] The visual feature extraction module is used to extract local artifact features F from the image to be identified. l and global semantic features F g ;

[0153] The gaze strategy module is used to obtain the predicted gaze position through a 4-layer causal Transformer decoder. and path termination probability s t ;

[0154] The evidence reading module is used to fuse the local artifact features F l Global semantic features F g and predicting gaze location Obtain the evidence vector e t ;

[0155] The evidence accumulation module is used to accumulate evidence based on evidence vector e. t Predicting gaze position And the hidden state h from the previous moment t-1 Obtain the identification confidence level p t And identification stop confidence level s c ;

[0156] The adaptive stopping decision module is used to determine the stopping confidence level p based on the identification confidence level p. t And identification stop confidence level s c To determine whether the decision should be stopped.

[0157] In practice, the bionic machine vision system for generating and identifying images from large models also includes a diffusion scanning path teacher module.

[0158] The diffusion scan path teacher module is used during the training phase and is only used in the training process; it does not participate in inference. The diffusion scan path teacher module is a conditional latent diffusion model used to learn the gaze scan path distribution when a human observer reviews images generated by a large model, providing high-quality training supervision signals for the causal gaze strategy student; these high-quality training supervision signals are input to the causal gaze strategy module.

[0159] Figure 10 The training supervision relationship and inference call relationship between the diffusion scan path teacher, causal gaze strategy student, multi-scale gaze reader, evidence accumulator and adaptive stopping decision module are shown. The dashed line represents supervision only during the training phase.

[0160] In the diffusion scan path teacher module, each scan path token is embedded and combined with the learned EOS token to form a latent sequence z_{1:max}; where max = 400, which is the maximum scan path length; the diffusion process is performed on this latent sequence z to model the multimodal distribution of human scan paths. The conditional input of the diffusion scan path teacher module consists of two parts: (1) image conditions, i.e., global semantic features F g (1) Inject denoising network through cross attention; (2) Task label conditions, i.e., real image labels (large model generated images / real photos), are converted into condition vectors by binary classification task tokens.

[0161] The teacher module of the diffusion scan path employs cosine noise scheduling, with 50 diffusion steps. The forward diffusion process progressively adds noise to the latent representation of the original scan path until it becomes pure Gaussian noise; the reverse denoising process starts from pure Gaussian noise and gradually recovers it through a denoising network. Compared to linear scheduling, cosine scheduling provides smoother noise changes in the early and late stages, contributing to the generation of higher-quality samples.

[0162] The denoising network of the teacher module in the diffusion scanning path is a 6-layer causal denoising Transformer. Each layer of the 6-layer causal denoising Transformer includes the following four layers: ① a causal self-attention layer, where temporal causality is ensured through a lower triangular mask; ② a cross-attention layer (attention head) with image conditions, where the current hidden state is Query and the image features are Key / Value; ③ a feedforward network (GELU activation); and ④ a pre-layer normalization layer (Pre-LayerNorm). In one embodiment, the 6-layer causal denoising Transformer has a hidden dimension of 256, 8 attention heads, each with a dimension of 32, and uses learnable absolute positional encoding for positional coding, with a maximum sequence length of max.

[0163] The following parameters are predicted after the output of the denoising network is decoded:

[0164] ①Grid cell distribution (3601-dimensional discrete logits, Softmax normalization);

[0165] ② Gaussian parameters (mean and logarithmic variance) of residual offsets Δx and Δy;

[0166] ③ Duration of fixation The parameters (mean and standard deviation) are log-normally distributed. The training loss is the standard diffusion denoising loss. After training, the teacher is frozen and used only to generate scan path samples, not for inference.

[0167] In the biomimetic machine vision method for image recognition based on large models of this invention, the training phase includes training the trainable parameters in the diffusion scanning path teacher module, the causal gaze strategy student module, the visual feature extraction module, the evidence reading module, the evidence accumulation module, and the adaptive stopping module. All trainable parameters involved in training are updated using the AdamW optimizer with an initial learning rate of 1e-4, employing a cosine decay strategy, a batch size of 32, and early stopping based on the validation set AUROC (Area Under the Receiver Operating Characteristic Curve). After training, the diffusion scanning path teacher module is frozen and used only to generate scanning path supervision signals during the training phase; it does not participate in the inference phase computation. The inference phase only calls the causal gaze strategy student module, the visual feature extraction module, the evidence reading module, the evidence accumulation module, and the adaptive stopping module to complete the recognition.

[0168] In the experiment, the images to be identified were uniformly adjusted to a resolution of 1920×1080, and the maximum scan path length was set to max=400. Evaluation metrics included Accuracy, F1, AUROC, AUPRC (Area Under the Precision-Recall Curve), single-image delay, and average number of fixation locations. Fixation correlation analysis also used fixation prediction error and model-human ScanMatch (scan path similarity index). The main results were averaged over three random seeds, and the mean ± standard deviation was reported. The AUROC confidence interval was calculated using the test image bootstrap method.

[0169] On the main test set, this invention achieved an accuracy of 93.7±0.2%, F1 score of 93.3±0.2%, AUROC score of 98.2±0.1%, AUPRC score of 97.9±0.1%, an average number of fixations of 31±4 steps, and a single-image delay of approximately 49.6 ms. Comparison with four representative baselines (AUROC / single-image delay):

[0170] 1. ViT-B / 16 (Visual Transformer end-to-end classifier): 91.5% / 7.2ms. A one-time global classifier that cannot utilize local artifacts and the accumulation of sequence evidence, performing poorly on difficult samples;

[0171] 2. DIRE (Diffusion Reconstruction Error): 94.3% / 161.4ms. It performs well for diffusion-generated images but has a huge computational cost.

[0172] 3. AEROBLADE (Latent Diffusion Reconstruction Error, No Training Required): 94.8% / 148.7ms. Also relies on reconstruction error, resulting in high latency;

[0173] 4. AIDE (Auxiliary Information Augmentation, State-of-the-art Static Detector): 96.6% / 11.3ms. Still a one-off global classifier, lacking sequence evidence accumulation.

[0174] The AUROC of this invention is 1.6 percentage points higher than the strongest static baseline AIDE; although it is slower than a single-shot classifier due to sequence characteristics, it is significantly faster than the reconstruction classifier (DIRE / AEROBLADE), achieving a good balance between accuracy and efficiency.

[0175] In constructing a post-processing robustness test set, JPEG compression (quality factor Q=50), resolution scaling (0.75×), Gaussian blur (σ=1.0), and screenshot re-encoding (simulating secondary encoding distortion from social media platforms / screen captures) were applied. The AUROC results of this invention are compared with the baseline as follows:

[0176] 1. JPEG compression (Q=50): This invention 94.0%, ViT-B / 16 83.5%, DIRE 82.1%, AEROBLADE 84.4%, AIDE 89.2%;

[0177] 2. Resolution scaling (0.75×): This invention 94.6%, ViT-B / 16 85.0%, DIRE 86.7%, AEROBLADE 87.2%, AIDE 90.1%;

[0178] 3. Gaussian blur (σ=1.0): This invention 93.2%, ViT-B / 16 81.8%, DIRE 80.4%, AEROBLADE 83.1%, AIDE 88.5%;

[0179] 4. Screenshot recoding: This invention 92.8%, ViT-B / 16 84.2%, DIRE 85.5%, AEROBLADE 86.0%, AIDE 87.9%.

[0180] The gaze-guided sequence evidence accumulation mechanism can continuously focus on stable local artifacts and semantically inconsistent regions under post-processing perturbations, thus maintaining good robustness; DIRE / AEROBLADE relies on pixel-level traces, and its performance degrades significantly under compression and blurring.

[0181] The present invention provides a biomimetic machine vision method for image recognition based on large models. This method employs a human gaze-guided sequential evidence accumulation framework, executing a complete "see-read-accumulate-stop" process. At an abstract level, the system comprises four essential modules: a gaze strategy module, an evidence reading module, an evidence accumulation module, and an adaptive stopping module. In preferred implementations, it also includes an input representation module and a visual feature extraction module. A preferred implementation of the visual feature extraction module is a two-stream visual encoder, and a preferred implementation of the evidence reading module is a multi-scale gaze reader.

[0182] The biomimetic machine vision method for image recognition based on large model generation in this invention has the following technical features:

[0183] 1. For the first time, image recognition generated by large models is redefined as a sequence visual search problem guided by human gaze, giving the recognition process interpretability and making the decision-making process transparent and traceable through gaze path visualization;

[0184] 2. The dual-stream visual encoder (SigLIP2-NaFlex global semantic stream + DINOv2 local artifact stream) has complementary functions, achieving an AUROC of 98.2% on the main test set, which is significantly better than existing methods;

[0185] 3. The diffusion scanning path teacher-causal gaze strategy and student teacher-student distillation framework effectively learn the multimodal distribution of human gaze, requiring only an average of 31±4 steps to complete the detection;

[0186] 4. The SSM time-series evidence accumulator + adaptive stopping mechanism can adaptively adjust the inference depth according to the sample difficulty, stopping early for simple samples and conducting in-depth analysis for difficult samples;

[0187] 5. Cross-generator migration AUROC reaches 91.7% / 92.4%, post-processing robustness is comprehensively superior, and generalization ability is strong.

[0188] like Figure 13 The method of this invention and the detection under conditions with and without human gaze supervision are used to illustrate the trend of the discrimination confidence of the two detection methods with the number of gaze positions, which is used to illustrate the convergence effect of the adaptive stopping mechanism—the gaze-supervised version converges faster and the prediction is more stable.

[0189] Figure 13 The results show the trend of the identification stopping confidence of the method of the present invention and the method without human gaze supervision as a function of the number of gaze steps, which is used to illustrate the convergence effect of the adaptive stopping mechanism. Figure 13 This shows the discrimination stopping confidence s of the model output as the number of fixation steps increases. c The trend of change. The horizontal axis represents the number of fixations, and the vertical axis represents the confidence level of discrimination stop. c The blue curve represents the comparison method without human gaze supervision, and the orange curve represents the algorithm of this invention. It can be seen that the algorithm of this invention identifies the stopping confidence s in fewer gaze steps. c The faster ascent and earlier attainment of higher confidence levels indicate that human gaze supervision guides the model to accumulate valid evidence more quickly, enabling the adaptive stopping module to determine within fewer gaze steps that the current evidence is sufficient to output a discrimination result. The light-colored area around the curve represents the range of result fluctuations; the algorithm of this invention exhibits smaller fluctuations, indicating that its stopping judgment process is more stable.

[0190] For those skilled in the art, without departing from the core ideas and basic technical features of this invention, the present invention is not limited to the specific implementation of image identification using a large model generated as an example in the above embodiments. The gaze guidance, local evidence reading, sequential evidence accumulation, and adaptive stopping decision mechanisms proposed in this invention can also be extended to other image or video analysis tasks that require simulating the human visual review process, such as image authenticity identification, video frame-level authenticity detection, image / video quality evaluation, image aesthetic evaluation, content security review, and user subjective satisfaction prediction. Any equivalent substitutions, modifications, or combinations made based on the technical concept of this invention to the input data type, feature extraction network, gaze window scale, evidence accumulation model, or stopping decision strategy should fall within the protection scope defined by the claims of this invention.

[0191] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A biomimetic machine vision method for image recognition based on large models, characterized by: Includes the following steps: Step 1: Scan Path Sequence Acquisition Step; Process the image to be identified to obtain the scan path sequence G = (g1, g2, …, g max ); Step 2: Image feature extraction step; Local artifact features F are extracted from the image to be identified. l and global semantic features F g ; Step 3: Fixation Path Prediction Step; The fixation strategy module obtains the predicted fixation location through a 4-layer causal Transformer decoder. and path termination probability s t ; Step 4: Evidence reading step; The local artifact features F obtained in Step 2... l and global semantic features F g By fusing the data, we obtain the evidence vector e. t ; Step 5: Evidence accumulation step; through the obtained evidence vector e t The hidden state h of the evidence accumulator in the evidence accumulation module. t The system is updated, and then the updated evidence accumulator is used to calculate the output discrimination confidence level p. t And identification stop confidence level s c ; Step 6: Adaptive stopping and result output step; based on the discrimination confidence level p obtained in Step 5. t And identification stop confidence level s c To determine whether the decision should be stopped; Output the identification result based on whether the forced stop condition is met.

2. The biomimetic machine vision method for image recognition based on large model generation according to claim 1, characterized in that, The process of obtaining the scan path sequence G in step 1 includes the following steps; Step 11: Image meshing step; After unifying the resolution of the image to be identified, divide it into uniformly sized grid cells; Step 12: Grid cell indexing confirmation step; A one-dimensional grid index k is used to identify the position of each grid cell in the image to be identified; Step 13: Residual offset calculation steps; calculate the residual offsets Δx and Δy of the gaze point within the grid cell; Step 14: Fixation Duration The calculation steps; Step 15: Representation of the scan path sequence G; Represent the complete scan path of the image to be identified as the sequence G = (g1, g2, …, g max ).

3. The biomimetic machine vision method for image recognition based on large model generation according to claim 2, characterized in that, In step 11, during the image meshing process, bicubic interpolation is used to upsample and enlarge the image, while regional average pooling is used to downsample and shrink the image.

4. The biomimetic machine vision method for image recognition based on large model generation according to claim 2, characterized in that, In step 14, the natural logarithm is used to calculate the fixation duration. .

5. The biomimetic machine vision method for image recognition based on large model generation according to claim 1, characterized in that, In step 2, the DINOv2 visual encoder is used to extract local artifact features F. l The SigLIP2-NaFlex visual encoder is used to extract global semantic features F. g .

6. The biomimetic machine vision method for image recognition based on large model generation according to claim 1, characterized in that, In step 3, the causal Transformer decoder includes a causal self-attention layer, a cross-attention layer, a feedforward network, and a pre-layer normalization layer (Pre-LayerNorm).

7. The biomimetic machine vision method for image recognition based on large model generation according to claim 1, characterized in that, In step 4, during evidence reading, a local window of three scales is used to obtain the evidence vector e. t .

8. The biomimetic machine vision method for image recognition based on large model generation according to claim 1, characterized in that, In step 5, a two-layer selective state-space module is used to calculate the discrimination confidence level p. t And identification stop confidence level s c .

9. A biomimetic machine vision system for image recognition based on large models, characterized in that, It includes an input representation module, a visual feature extraction module, a gaze strategy module, an evidence reading module, an evidence accumulation module, and an adaptive stopping decision module; The input representation module is used to convert the image to be identified into a scan path sequence G; The visual feature extraction module is used to extract local artifact features F from the image to be identified. l and global semantic features F g ; The gaze strategy module is used to obtain the predicted gaze position through a 4-layer causal Transformer decoder. and path termination probability s t ; The evidence reading module is used to fuse the local artifact features F l Global semantic features F g and predicting gaze location Obtain the evidence vector e t ; The evidence accumulation module is used to accumulate evidence based on evidence vector e. t Predicting gaze position And the hidden state h from the previous moment t-1 Obtain the identification confidence level p t And identification stop confidence level s c ; The adaptive stopping decision module is used to determine the stopping confidence level p based on the identification confidence level p. t And identification stop confidence level s c To determine whether the decision should be stopped.

10. The bionic machine vision system for image recognition based on large model generation according to claim 1, characterized in that, It also includes a diffusion scan path teacher module.

Citation Information

Patent Citations

  • Image recognition method based on eye-tracking gaze point guidance, MR glasses and media

    CN112507799B

  • A Method and System for Detecting Forged Videos Based on Blink Synchronization and Binocular Movement Detection

    CN113627256B

  • Face deep false detection method based on multi-modal feature fusion

    CN115880749A

  • Face forgery detection method based on depth information decoupling

    CN116704580A

  • Deep counterfeit multi-label sorting and positioning method based on multi-domain feature fusion

    CN120032234A