AI generation military image detection method based on training in adaptive test

Through the adaptive test-time training method, utilizing the military layered perception module and multi-task loss function, the problems of insufficient targeting and robustness in the existing technology for detecting forged military images generated by AI are solved, and efficient and accurate detection of military images is achieved.

CN120689327APending Publication Date: 2025-09-23SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510825314.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies lack targeted and robust solutions when detecting forged content in AI-generated military images. They have difficulty effectively identifying the layering patterns and manufacturing consistency of military equipment, and lack the ability to adapt to new generative models.

Method used

A method based on adaptive test-time training is adopted to extract multi-level features through the military hierarchical perception module. Combined with cross-layer feature fusion and attention mechanism, the model is trained using a multi-task loss function to achieve fine-grained modeling and dynamic adjustment of military images.

Benefits of technology

It significantly improves the accuracy and robustness of forgery detection of AI-generated military images, can dynamically adapt to the distribution drift and forgery feature changes of different generation models, and improves the generalization ability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0FI3CMBHQF9BIFKYMIVTXFO9DX89J20LNZNJO23H
    Figure 0FI3CMBHQF9BIFKYMIVTXFO9DX89J20LNZNJO23H
  • Figure 2XTV34WJQTXYVRYPD7LW68ON6J6ZQMFIT8M6KFQZ
    Figure 2XTV34WJQTXYVRYPD7LW68ON6J6ZQMFIT8M6KFQZ
  • Figure 3NKCMNLNJ2ZFLQ1TWZ59NIA9UTNFNQ1SCPFUDBJC
    Figure 3NKCMNLNJ2ZFLQ1TWZ59NIA9UTNFNQ1SCPFUDBJC
Patent Text Reader

Abstract

The invention relates to the technical field of image detection, in particular to an AI generation military image detection method based on self-adaptive test training, which comprises the following steps: S1, preprocessing an input image to obtain a preprocessed image; s2, extracting multi-level features of the preprocessed image by a military hierarchical sensing module, and obtaining fusion features according to a cross-layer feature fusion mechanism; s3, weighting fusion features by an attention mechanism to obtain weighted features, and obtaining an image forgery probability according to an MLP classifier; s4, in a training stage, combining the loss of the main classification task and the loss of the self-supervised auxiliary task to analyze the loss so as to update model parameters to obtain a training model, and adopting a training strategy during self-adaptive testing to obtain an unknown forged model; and S5, inputting a military image to be detected into the unknown counterfeit model, and outputting an authenticity judgment result. The method can significantly improve the detection accuracy and generalization ability of AI generated military images, is suitable for multiple practical application scenes such as media information security, and has high practical value and wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image detection technology, and in particular to a method, device, and storage medium for detecting military images generated by AI based on adaptive testing training. Background Art

[0002] With the rapid development of AI-generated content (AIGC) technology, deep learning-based image generation models (such as DALL-E 2 and Stable Diffusion) have become capable of producing highly realistic images that are often visually indistinguishable from real photographs. This advancement has brought numerous benefits to the creative industries and information dissemination, but it has also created serious security and social cognition challenges, particularly in the military. If AI-generated military imagery is maliciously fabricated and disseminated, it can easily cause public panic and even threaten national security and the reliability of military intelligence. Therefore, how to effectively detect and identify forged AI-generated military imagery has become a critical technical challenge that needs to be addressed in the fields of AI security and national defense information security.

[0003] Currently, AIGC image forgery detection technology is mainly divided into active forensics methods and passive forensics methods. Active forensics methods such as digital watermarks and steganographic markers can achieve traceability, but they need to be embedded in advance during the generation stage, which makes it difficult to deal with forged content that has spread on a large scale. Passive forensics methods analyze the characteristics of the image itself to identify authenticity and have become the mainstream research direction. In recent years, many patents have been published in the field of deep forgery detection at home and abroad, but there are still shortcomings: First, the patent publication number CN114757877B, entitled "A deep forgery detection method based on frequency domain filter residuals", proposed a frequency domain filter residual feature, and the patent publication number CN118196865B, entitled "Generalizable deep forgery image detection method and system based on noise perception", proposed a noise perception detection strategy. These methods effectively capture anomalies in high-frequency or noise distributions in forged images, improving the detection of some deepfakes. However, these solutions often rely on statistical features in a single domain (spatial or frequency), making it difficult to account for the complex, multi-layered structural features of military images. They also have limited generalization to novel generative models and lack specialized modeling of military equipment structures and manufacturing specifications. Furthermore, while various feature fusion and deep fusion network structures have been proposed in existing technologies, these approaches are primarily targeted at general natural images and lack specific design considerations for visual structures unique to the military domain (such as equipment hierarchies and standardized components). Their feature fusion strategies fail to fully account for the hierarchical patterns and manufacturing consistency of military equipment, resulting in limited detection effectiveness in military scenarios. Furthermore, while existing technologies have introduced adaptive discrimination mechanisms and generalization enhancement strategies, they primarily focus on conventional scenes such as faces and lack the ability to adapt to the unique distributional shifts and generative model diversity of military images. Existing methods struggle to fine-grainedly model and dynamically adjust to the subtle differences between authentic manufacturing standards for military equipment and AI-generated forgeries.

[0004] Therefore, existing technologies have not yet formed a highly targeted and robust solution to the problem of detecting AI-generated military image forgeries. It is urgent to propose new detection methods and systems that can combine military field knowledge, have layered perception capabilities, and support adaptive training, so as to improve the accuracy and generalization ability of military image forgery detection and ensure national defense information security and social stability. Summary of the Invention

[0005] The purpose of the present invention is to address the shortcomings of the prior art and provide an AI-generated military image detection method based on adaptive test-time training, comprising the following steps: S1: Perform preprocessing including image cropping, Gaussian blurring, and JPEG compression on the input military image to obtain a preprocessed image; S2: extracting multi-level features including global semantic features and multi-scale features from the pre-processed image through a military layered perception module, and fusing the multi-level features using a cross-layer feature fusion mechanism to obtain fused features; S3: Using the attention mechanism to perform feature weighting on the fusion features to obtain weighted features, the multi-layer perceptron (MLP) classifier performs binary classification based on the weighted feature map to obtain the image forgery probability, thereby achieving classification of true and false military images; S4: During the training phase, the main classification task loss and the self-supervised auxiliary task loss, including the local consistency verification task and the multi-scale feature signature analysis task, are combined to analyze the task loss. The model parameters are updated according to the task loss to obtain a training model. An adaptive test-time training strategy is adopted to adjust the training model according to the test sample to obtain an unknown forgery model. S5: Input the military image to be detected into the unknown forgery model and output the authenticity determination result.

[0006] Preferably, in step S2, extracting multi-level features including global semantic features and multi-scale features from the pre-processed image by a military layered perception module further includes: The last layer of the visual Transformer backbone network in the military hierarchical perception module outputs backbone network features; The global semantic branch extracts global semantic features based on the backbone network features , as shown below: in, Represents the feature transformation operation of the last layer of the Transformer backbone network, 、 、 Full The height, width and number of channels of the local feature map; The multi-scale feature branch extracts multi-scale features from the 1 / 4, 1 / 2 and 3 / 4 depths of the visual Transformer backbone network, and the features extracted in the i-th layer are recorded as ,in They correspond to the early, middle and late stages of feature abstraction, as shown below: in, Represents the feature extraction operation of the i-th layer of the visual Transformer backbone network.

[0007] Preferably, the multi-level features are fused using a cross-layer feature fusion mechanism to obtain a fused feature map, further comprising: The cross-layer feature fusion mechanism first passes the i-th layer feature Convolution to obtain lateral connections and residual connections , as shown below: in, and are the i-th layer Convolution operation to achieve lateral and residual connections; The bottom-up fusion path uses adaptive weights to integrate features layer by layer. The bottom-up features of each layer i Iterate and calculate as follows: If i=1, then ;otherwise ; in, is the downsampling operation, is the feature after processing in the previous layer, is the learnable weight of the i-th layer, for i=1, For lateral connection only; The dual-path fusion of each layer obtains the fusion feature in, is the attention module of the i-th layer, 、 is the adaptive fusion weight of this layer.

[0008] Preferably, in step S3, the fusion feature is weighted by using the attention mechanism to obtain a weighted feature, further comprising: The attention mechanism is used to weight the fusion features in channels and spaces, that is, the processed multi-scale features and global semantic features Input feature fusion module Perform weighting to obtain the weighted features: Among them, the fusion module The HiLo Attention mechanism and the ECA-Net mechanism are combined. HiLo Attention prioritizes military-salient areas, while the ECA-Net mechanism refines channel importance to highlight military-related feature channels.

[0009] Preferably, in step S4, during the training phase, the main classification task loss and the self-supervised auxiliary task loss including the local consistency verification task and the multi-scale feature signature analysis task are combined to analyze the task loss, and the model parameters are updated according to the task loss to obtain the training model, further comprising: The binary cross entropy loss is calculated by the main classification task ; Calculate the LCV loss through the local consistency verification task; The signature analysis loss is calculated by the multi-scale feature signature analysis task. , and through adaptive weights and adaptive weights to strike a balance; The target training model is based on the binary cross entropy loss , the LCV loss and the signature analysis loss Update the model parameters to obtain the training model, the target training model, as shown below, The self-supervised auxiliary task uses independent task parameters , the main classification task adopts , and the feature extractor parameters are recorded as .

[0010] Preferably, in step S4, calculating the LCV loss through the local consistency verification task further includes: The local consistency verification task divides the weighted feature map into 6×6 grids through the unfold operator and adaptive average pooling to obtain local feature vectors, and each feature vector is reduced to a representative feature vector Respectively represent the i-th row and j-th column of the grid; Normalize each representative feature vector to obtain the normalized vector .

[0011] Calculate the cosine similarity of adjacent feature vectors based on the representative vector and the normalized vector: in and are the normalized eigenvectors at grids (i, j) and (k, l) respectively; The LCV loss function calculates the LCV loss based on the cosine similarity as follows: in is the temperature parameter, the denominator is summed over all pairs of partial eigenvectors, and for adjacent partial eigenvectors sharing an edge, , for diagonally adjacent eigenvectors, , to reflect actual continuity expectations.

[0012] Preferably, in step S4, the signature analysis loss is calculated by the multi-scale feature signature analysis task , further including: Convert RGB image Projection onto the grayscale luminance plane ; The grayscale brightness plane With uniform smoothing kernel Perform convolution to obtain a smooth representation as follows: The grayscale brightness plane The spatial noise map is obtained by taking the absolute residual of the surface and the smoothed representation : ; The spatial noise map Projected to the frequency domain by two-dimensional fast Fourier transform FFT, , its frequency domain noise diagram The mathematical expression is as follows: in is the noise intensity at the spatial coordinate (x, y), is the amplitude of the frequency domain noise component at frequency (u, v); Through dual analysis in the spatial domain and frequency domain, the multi-scale feature signature analysis task statistically calculates three feature indicators at each scale, including spatial mean, frequency domain mean and spatial standard deviation. For each scale s, a dedicated noise prediction head generates a noise statistical prediction value at that scale based on the feature extractor fusion feature. , the noise statistical prediction value and the true statistic Make a comparison; Statistical loss function According to the noise statistics prediction value and the true statistic Calculate the spatial mean loss, frequency domain mean loss and spatial standard deviation loss respectively. The statistical loss function The calculation formula is as follows: in is the weight factor; According to the statistical loss function Get the overall loss at each scale ; According to the overall loss, the signature analysis loss is obtained , as shown below: in is the number of scales, and the contribution of each scale is determined by the scale weight adjust.

[0013] Preferably, an adaptive test-time training strategy is adopted to adjust the training model according to the test sample to obtain the unknown forged model, further comprising: Freeze classifier parameters during testing ; Minimize the weighted combination of self-supervised losses according to the SGD optimizer as follows: in, is the shared feature extractor parameters; The position forgery model is obtained according to the optimized combination, wherein the optimized combination is as follows: in, are the optimized parameters of the feature extractor during testing.

[0014] Based on the same concept, the present invention also provides a computer device including a memory and one or more processors, wherein the memory stores computer code, and when the computer code is executed by the one or more processors, the one or more processors execute the AI-generated military image detection method based on adaptive test-time training as described in any one of the embodiments.

[0015] Based on the same concept, the present invention also provides a computer-readable storage medium, which stores computer code. When the computer code is executed, the AI-generated military image detection method based on adaptive test-time training as described in any of the embodiments is executed.

[0016] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention extracts multi-level features including global semantic features and multi-scale features from the pre-processed image through a military hierarchical perception module, thereby effectively capturing anomalies of military targets in various hierarchical structures. The present invention then fuses the multi-level features using a cross-layer feature fusion mechanism to obtain fused features, thereby effectively improving detection performance by collaboratively extracting and fusing military-related hierarchical feature representations, thereby improving the accuracy and robustness of forgery detection. (2) The present invention uses the attention mechanism to perform feature weighting on the fused features to obtain weighted features, thereby focusing on military salient areas and refining the importance of channels to highlight military-related features.

[0017] (3) The present invention combines the main classification task loss with the self-supervised auxiliary task loss including the local consistency verification task and the multi-scale feature signature analysis task to analyze the task loss, and updates the model parameters according to the task loss to obtain a training model, thereby effectively capturing the subtle anomalies in the local structural continuity of AI forged images. By adopting an adaptive test-time training strategy, it can achieve dynamic adaptation to the local inconsistency and noise distribution changes brought about by different AI generation models, significantly improving the generalization ability and robustness of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Various other advantages and benefits will become apparent to those skilled in the art by reading the following detailed description of the preferred embodiment.The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the invention.

[0019] Figure 1 A flowchart of the method for generating military image detection based on AI trained during adaptive testing according to the present invention; Figure 2 Another flowchart of the present invention's method for generating military image detection based on AI trained during adaptive testing. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Obviously, the embodiments described are part of the embodiments of this application, rather than all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative work are within the scope of protection of this application.

[0021] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0022] First embodiment In summary, existing patented technologies have made certain progress in the field of AIGC image forgery detection. However, when dealing with AI-generated image forgery detection in military scenarios, there are the following prominent deficiencies: First, existing detection models are mainly trained based on general-domain datasets and lack the learning and understanding of visual features specific to the military field, resulting in limited performance in military image detection. Second, the structural design of general forgery detection models fails to optimize the special vulnerabilities of military images. Real military equipment usually has strict manufacturing standards and structural laws, while AI-forged images often have difficult-to-detect anomalies in these details, and special feature extraction and fusion mechanisms are urgently needed to identify them. Finally, with the rapid iteration of various AI generation models, image forgery techniques are becoming increasingly diverse. Static detection methods are difficult to adapt to the distribution drift and forgery style changes brought about by new generation technologies, and lack adaptive capabilities.

[0023] See also Figure 1 and Figure 2 As shown, the AI-generated military image detection method based on adaptive test-time training provided in this embodiment includes: S1: Perform preprocessing including image cropping, Gaussian blurring, and JPEG compression on the input military image to obtain a preprocessed image. Specifically, in this embodiment, the military image is received. ,in and are the height and width of the image respectively, 3 is the number of color channels, the image can come from a public military image library, network collection or actual scene shooting, the input military image is cropped to 224×224 size, and during training, Gaussian blur and JPEG compression are applied with a probability of 0.1 for data enhancement.

[0024] To address the need for robust detection of AI-generated forged images in sensitive military scenarios, a specially designed detection framework, SentinelFakeNet (SFNet), is proposed. The core of SFNet is the Military Hierarchical Perception Module (MHP Module), which can collaboratively extract and fuse military-related hierarchical feature representations to improve the accuracy and robustness of forgery detection.

[0025] The design of the MHP module fully considers the unique structural and visual characteristics of military assets. Military targets typically exhibit distinct hierarchical structural features, ranging from basic visual primitives (such as tank armor plating textures), mid-level component arrangements (such as missile launcher layouts), to high-level global contours (such as warship shapes). AI-generated models often struggle to fully reproduce these unique military visual signatures, so deep modeling of these hierarchical features can effectively improve detection performance.

[0026] S2: The military hierarchical perception module is used to extract multi-level features including global semantic features and multi-scale features from the preprocessed image, and the cross-layer feature fusion mechanism is used to fuse the multi-level features to obtain fused features.

[0027] Preferably, in step S2, extracting multi-level features including global semantic features and multi-scale features from the pre-processed image by a military layered perception module further includes: The final layer of the visual Transformer backbone network in the military hierarchical perception module outputs backbone network features. Specifically, in this embodiment, the visual Transformer backbone network (CLIP-ViT-L / 14) is used as the feature extraction backbone network. This network has rich visual semantic understanding capabilities and is pre-trained on 400 million sets of images and text. Based on this backbone network, the MHP module designs a cross-level feature fusion (CLFF) mechanism to effectively capture anomalies in military targets at all levels of the structure. The global semantic branch extracts global semantic features based on the backbone network features , as shown below: in, Represents the feature transformation operation of the last layer of the Transformer backbone network, 、 、 are the height, width, and number of channels of the global feature map, respectively. Specifically, in this embodiment, the global semantic branch directly uses the backbone network features output by the last layer of the backbone network to obtain a holistic global representation for detecting forgery anomalies in the overall proportion and contour of AI-generated images; The multi-scale feature branch extracts multi-scale features from the 1 / 4, 1 / 2 and 3 / 4 depths of the visual Transformer backbone network, and the features extracted in the i-th layer are recorded as ,in They correspond to the early, middle and late stages of feature abstraction, as shown below: in, Represents the feature extraction operation of the i-th layer of the visual Transformer backbone network. Specifically, in this embodiment, the multi-scale feature branch adopts an improved feature pyramid network FPN structure, which is specially used to capture the hierarchical features of military equipment.

[0028] Preferably, a cross-layer feature fusion mechanism is used to fuse multi-level features to obtain a fused feature map, further comprising: The cross-layer feature fusion mechanism first passes through the i-th layer feature Convolution to obtain lateral connections and residual connections , as shown below: in, and are the i-th layer Convolution operation realizes lateral and residual connections. Specifically, in this embodiment, the CLFF mechanism designs lateral connections and residual connections to achieve effective integration of multi-level features; The bottom-up fusion path uses adaptive weights to integrate features layer by layer. The bottom-up features of each layer i Iterate and calculate as follows: If i=1, then ;otherwise ; in, is the downsampling operation, is the feature after processing in the previous layer, is the learnable weight of the i-th layer, for i=1, For lateral connection only; The dual paths of each layer are fused to obtain the fusion features in, is the attention module of the i-th layer, 、 is the adaptive fusion weight of this layer.

[0029] S3: The attention mechanism is used to weight the fused features to obtain weighted features. The multi-layer perceptron (MLP) classifier performs binary classification based on the weighted feature map to obtain the image forgery probability and realize the classification of true and false military images.

[0030] Preferably, in step S3, the fusion features are weighted using the attention mechanism to obtain weighted features, further comprising: The attention mechanism is used to weight the fusion features in channels and space, that is, the processed multi-scale features and global semantic features Input feature fusion module Perform weighting to obtain weighted features: Among them, the fusion module The HiLo Attention mechanism and the ECA-Net mechanism are combined. HiLo Attention prioritizes military-salient areas, while the ECA-Net mechanism refines channel importance to highlight military-related feature channels.

[0031] To further improve the robustness of the present invention in detecting forged military images generated by diverse AI generative models, a Military Adaptive Test-Time Training (MATTT) strategy is proposed. This strategy enables the detection system to dynamically adapt to the distribution shift and forgery feature changes brought about by different generative models through multi-task training and test-time adaptation mechanisms.

[0032] S4: During the training phase, the main classification task loss and the self-supervised auxiliary task loss including the local consistency verification task and the multi-scale feature signature analysis task are combined to analyze the task loss. The model parameters are updated according to the task loss to obtain the training model. The adaptive test-time training strategy MATTT is adopted to adjust the training model according to the test sample to obtain the unknown forged model. Specifically, in this embodiment, SFNet adopts a multi-task learning strategy to jointly optimize the main classification task and the two self-supervised auxiliary tasks. The main classification task goal is to distinguish forged military images from real military images, and the binary cross entropy loss is adopted. , the auxiliary task LCV and the multi-scale feature signature analysis task MSSA correspond to the consistency verification loss respectively and signature analysis loss ,Adaptive Test-Time Training (MATTT) is used to achieve adaptive adjustment of unknown forged models.

[0033] Preferably, in step S4, during the training phase, the main classification task loss and the self-supervised auxiliary task loss including the local consistency verification task and the multi-scale feature signature analysis task are combined to analyze the task loss, and the model parameters are updated according to the task loss to obtain a training model, further comprising: Calculate the binary cross entropy loss through the main classification task ; LCV loss is calculated through the local consistency verification task. Specifically, in this embodiment, the feature map is divided into grids, and cosine similarity and position mask are used to evaluate local consistency to identify subtle inconsistencies in AI forgeries. Compute signature analysis loss via multi-scale feature signature analysis task , and through adaptive weights and adaptive weights to strike a balance; The target training model is based on binary cross entropy loss , LCV loss and signature analysis loss Update the model parameters to obtain the training model and the target training model as shown below. Among them, the self-supervised auxiliary task uses independent task parameters , the main classification task adopts , and the feature extractor (i.e., MHP module) parameters are recorded as .

[0034] The Local Consistency Verification (LCV) task detects subtle violations of standardized design specifications in AI-generated forged military imagery. While authentic military assets are typically manufactured to strict engineering standards, AI-generated forgeries often exhibit subtle but noticeable structural inconsistencies.

[0035] Preferably, in step S4, calculating the LCV loss through the local consistency verification task further includes: The local consistency verification task uses the unfold operator and adaptive average pooling to divide the weighted feature map into a 6×6 grid to obtain local feature vectors. Each feature vector is reduced to a representative feature vector. Respectively represent the i-th row and j-th column of the grid. Specifically, in this embodiment, the weighted feature map output by the MHP module is divided into Grid, 36 local patches are obtained; Normalize each representative feature vector to obtain the normalized vector Specifically, in this embodiment, the normalized vector is a unit vector.

[0036] Calculate the cosine similarity of adjacent feature vectors based on the representative vector and the normalized vector: in and are the normalized eigenvectors at grids (i, j) and (k, l), respectively. Specifically, in this embodiment, a contrast consistency framework is used to specifically detect abnormal patterns that violate natural imaging, and the pattern consistency between adjacent patches is quantified by the cosine similarity matrix S; The LCV loss function calculates the LCV loss based on cosine similarity as follows: in is the temperature parameter, the denominator is summed over all pairs of partial eigenvectors, and for adjacent partial eigenvectors sharing an edge, , for diagonally adjacent eigenvectors, , to reflect the actual continuity expectation. Specifically, in this embodiment, to define the positive sample pairs for contrastive learning, a position mask matrix P is introduced, and the LCV loss function is averaged within the batch to strengthen the consistency of features in adjacent regions and ensure the reasonable distinction between non-adjacent regions.

[0037] Through the above design, the LCV task can effectively capture subtle anomalies in the local structural continuity of AI-forged images, providing a key basis for the final judgment.

[0038] Spatial and frequency-domain noise features are extracted at multiple scales to analyze the micro-statistical differences between AI-generated and real military imagery. To further enhance forgery detection capabilities, this embodiment integrates the Multi-Scale Signature Analysis (MSSA) task into the MATTT strategy. The MSSA task is based on the observation that real military equipment, due to its unique material composition, color distribution, and texture, often exhibits noise patterns and characteristic "signatures" in images that are difficult for AI-generated systems to accurately reproduce.

[0039] Preferably, in step S4, the signature analysis loss is calculated by the multi-scale feature signature analysis task , further including: The MSSA task is characterized on three strategic scales: , for each scale s, the RGB image Projection onto the grayscale luminance plane ,Specifically, in this embodiment, the projection process is achieved by weighted summation of ,each color channel; Grayscale brightness plane With uniform smoothing kernel Perform convolution to obtain a smooth representation as follows: Grayscale brightness plane The absolute residual between the surface and the smooth representation is used to obtain the spatial noise map : Specifically, in this embodiment, this operation can reveal inconsistencies in the micro-texture level of AI-generated military assets, as AI-generated models often have difficulty accurately reproducing the complex surface properties (e.g., ballistic armor texture) and color patterns (e.g., camouflage pattern distribution) of military-grade materials. In order to further extract the characteristic signature of real military equipment in the frequency domain, the spatial noise map Projected to the frequency domain by two-dimensional fast Fourier transform FFT, , its frequency domain noise diagram The mathematical expression is as follows: in is the noise intensity at the spatial coordinate (x, y), is the amplitude of the frequency domain noise component at the frequency (u, v). Specifically, in this embodiment, the unique frequency characteristics of real military assets due to their special material composition and specific texture are utilized; Through dual analysis in the spatial and frequency domains, the multi-scale feature signature analysis task calculates three characteristic indicators at each scale, including spatial mean, frequency mean, and spatial standard deviation. For each scale s, a dedicated noise prediction head generates a noise statistical prediction value at that scale based on the fusion features of the feature extractor (MHP module). , the noise statistics prediction value and the true statistic Specifically, in this embodiment, through the spatial-frequency dual-domain and multi-scale joint feature signature analysis, MSSA can effectively capture the subtle differences in high-order statistical properties between real military equipment and AI forgeries, greatly improving the accuracy and robustness of forgery detection; Statistical loss function Predicted values ​​based on noise statistics and the true statistic Calculate the spatial mean loss, frequency domain mean loss and spatial standard deviation loss respectively, and the statistical loss function The calculation formula is as follows: in is a weight factor. Specifically, in this embodiment, the weight factor Take 0.3, and the statistical loss function calculates the spatial mean loss, frequency domain mean loss, and spatial standard deviation loss respectively; According to the statistical loss function Get the overall loss at each scale Specifically, in this embodiment, the overall loss is the weighted sum of all prediction statistic losses at this scale; Based on the overall loss, the signature analysis loss is obtained , as shown below: in is the number of scales, and the contribution of each scale is determined by the scale weight Adjustment, specifically, in this embodiment, the total loss of MSSA Perform weighted aggregation across all scales, is 3.

[0040] Preferably, an adaptive test-time training strategy is adopted to adjust the training model according to the test sample to obtain an unknown forged model, further comprising: Freeze classifier parameters during testing Specifically, in this embodiment, following the principle of Test-Time Training (TTT), the classifier parameters Keep frozen, only the shared feature extractor parameters Optimize Minimize the weighted combination of self-supervised losses according to the SGD optimizer as follows: in, is the shared feature extractor parameters; The position forging model is obtained according to the optimized combination, where the optimized combination is as follows: in, are the optimized parameters of the feature extractor during testing.

[0041] This MATTT strategy enables the model to dynamically adapt to local inconsistencies and noise distribution changes brought about by different AI-generated models, significantly improving the generalization ability and robustness of detection.

[0042] S5: Input the military image to be detected into the unknown forgery model and output the authenticity determination result.

[0043] This embodiment fills a technological gap in the field of AI-generated military image forgery detection, combining layered perceptual feature extraction with adaptive test-time training. The proposed method can effectively improve the detection accuracy and generalization capabilities of AI-generated forgeries in complex military scenarios, and is suitable for a variety of practical applications, including national defense security, military intelligence, and public opinion monitoring. This technology not only enhances the level of military information security, but also provides strong technical support for intelligent content review and data traceability in related industries, and has broad commercial value and application prospects.

[0044] Second embodiment In this embodiment, a computer device is provided, including a memory and one or more processors. Computer code is stored in the memory. When the computer code is executed by the one or more processors, the one or more processors execute the steps of the military image detection method generated by AI trained during adaptive testing in the first embodiment.

[0045] In some embodiments of the present application, a computer-readable storage medium is also provided. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the steps of the AI-generated military image detection method based on adaptive test-time training as described in any one of the first embodiments.

[0046] It is understandable that for the aforementioned AI-generated military image detection method based on adaptive test-time training, if it is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc. Various media that can store program codes.

[0047] Computer-readable storage media may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0048] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An AI-generated military image detection method based on adaptive test-time training, characterized in that: The following steps are involved: S1: Perform preprocessing including image cropping, Gaussian blurring, and JPEG compression on the input military image to obtain a preprocessed image; S2: extracting multi-level features including global semantic features and multi-scale features from the pre-processed image through a military layered perception module, and fusing the multi-level features using a cross-layer feature fusion mechanism to obtain fused features; S3: Using the attention mechanism to perform feature weighting on the fusion features to obtain weighted features, the multi-layer perceptron (MLP) classifier performs binary classification based on the weighted feature map to obtain the image forgery probability, thereby achieving classification of true and false military images; S4: During the training phase, the main classification task loss and the self-supervised auxiliary task loss, including the local consistency verification task and the multi-scale feature signature analysis task, are combined to analyze the task loss. The model parameters are updated according to the task loss to obtain a training model. An adaptive test-time training strategy is adopted to adjust the training model according to the test sample to obtain an unknown forgery model. S5: Input the military image to be detected into the unknown forgery model and output the authenticity determination result.

2. The AI-generated military image detection method based on adaptive test-time training according to claim 1 is characterized in that: In step S2, multi-level features including global semantic features and multi-scale features are extracted from the pre-processed image by a military layered perception module, further comprising: The military layered perception module receives the preprocessed image, and the last layer of the visual Transformer backbone network in the military layered perception module outputs backbone network features; The global semantic branch extracts global semantic features based on the backbone network features , as shown below: in, Represents the feature transformation operation of the last layer of the Transformer backbone network, 、 、 are the height, width and number of channels of the global feature map respectively; The multi-scale feature branch extracts multi-scale features from the 1 / 4, 1 / 2 and 3 / 4 depths of the visual Transformer backbone network, and the features extracted in the i-th layer are recorded as ,in They correspond to the early, middle and late stages of feature abstraction, as shown below: in, Represents the feature extraction operation of the i-th layer of the visual Transformer backbone network.

3. The AI-generated military image detection method based on adaptive test-time training according to claim 2, characterized in that: The multi-level features are fused using a cross-layer feature fusion mechanism to obtain a fused feature map, further comprising: The cross-layer feature fusion mechanism first passes the i-th layer feature Convolution to obtain lateral connections and residual connections , as shown below: in, and are the i-th layer Convolution operation to achieve lateral and residual connections; The bottom-up fusion path uses adaptive weights to integrate features layer by layer. The bottom-up features of each layer i Iterate and calculate as follows: If i=1, then ;otherwise ; in, is the downsampling operation, is the feature after processing in the previous layer, is the learnable weight of the i-th layer, for i=1, For lateral connection only; The dual-path fusion of each layer obtains the fusion feature in, is the attention module of the i-th layer, 、 is the adaptive fusion weight of this layer.

4. The AI-generated military image detection method based on adaptive test-time training according to claim 3 is characterized in that: In step S3, the fusion features are weighted using the attention mechanism to obtain weighted features, further comprising: The attention mechanism is used to weight the fusion features in channels and spaces, that is, the processed multi-scale features and global semantic features Input feature fusion module Perform weighting to obtain the weighted features: Among them, the fusion module The HiLo Attention mechanism and the ECA-Net mechanism are combined. HiLo Attention prioritizes military-salient areas, while the ECA-Net mechanism refines channel importance to highlight military-related feature channels.

5. The AI-generated military image detection method based on adaptive test-time training according to claim 4 is characterized in that: In step S4, during the training phase, the main classification task loss and the self-supervised auxiliary task loss including the local consistency verification task and the multi-scale feature signature analysis task are combined to analyze the task loss, and the model parameters are updated according to the task loss to obtain a training model, further comprising: The binary cross entropy loss is calculated by the main classification task ; Calculate the LCV loss through the local consistency verification task; The signature analysis loss is calculated by the multi-scale feature signature analysis task. , and through adaptive weights and adaptive weights to strike a balance; The target training model is based on the binary cross entropy loss , the LCV loss and the signature analysis loss Update the model parameters to obtain the training model, the target training model, as shown below, The self-supervised auxiliary task uses independent task parameters , the main classification task adopts , and the feature extractor parameters are recorded as .

6. The AI-generated military image detection method based on adaptive test-time training according to claim 5, characterized in that: In step S4, calculating the LCV loss through the local consistency verification task further includes: The local consistency verification task divides the weighted feature map into 6×6 grids through the unfold operator and adaptive average pooling to obtain local feature vectors, and each feature vector is reduced to a representative feature vector Respectively represent the i-th row and j-th column of the grid; Normalize each representative feature vector to obtain the normalized vector ; Calculate the cosine similarity of adjacent feature vectors based on the representative vector and the normalized vector: in and are the normalized eigenvectors at grids (i, j) and (k, l) respectively; The LCV loss function calculates the LCV loss based on the cosine similarity as follows: in is the temperature parameter, the denominator is summed over all pairs of partial eigenvectors, and for adjacent partial eigenvectors sharing an edge, , for diagonally adjacent eigenvectors, , to reflect actual continuity expectations.

7. The AI-generated military image detection method based on adaptive test-time training according to claim 6, characterized in that: In step S4, the signature analysis loss is calculated by the multi-scale feature signature analysis task , further including: Convert RGB image Projection onto the grayscale luminance plane ; The grayscale brightness plane With uniform smoothing kernel Perform convolution to obtain a smooth representation as follows: The grayscale brightness plane The spatial noise map is obtained by taking the absolute residual of the surface and the smoothed representation : ; The spatial noise map Projected to the frequency domain by two-dimensional fast Fourier transform FFT, , its frequency domain noise diagram The mathematical expression is as follows: in is the noise intensity at the spatial coordinate (x, y), is the amplitude of the frequency domain noise component at frequency (u, v); Through dual analysis in the spatial domain and frequency domain, the multi-scale feature signature analysis task statistically calculates three feature indicators at each scale, including spatial mean, frequency domain mean and spatial standard deviation. For each scale s, a dedicated noise prediction head generates a noise statistical prediction value at that scale based on the feature extractor fusion feature. , the noise statistical prediction value and the true statistic Make a comparison; Statistical loss function According to the noise statistics prediction value and the true statistic Calculate the spatial mean loss, frequency domain mean loss and spatial standard deviation loss respectively. The statistical loss function The calculation formula is as follows: in is the weight factor; According to the statistical loss function Get the overall loss at each scale ; According to the overall loss, the signature analysis loss is obtained , as shown below: in is the number of scales, and the contribution of each scale is determined by the scale weight adjust.

8. The AI-generated military image detection method based on adaptive test-time training according to claim 7, characterized in that: Adopting an adaptive test-time training strategy, adjusting the training model according to the test sample to obtain an unknown forged model, further comprising: Freeze classifier parameters during testing ; Minimize the weighted combination of self-supervised losses according to the SGD optimizer as follows: in, is the shared feature extractor parameters; The position forgery model is obtained according to the optimized combination, wherein the optimized combination is as follows: in, are the optimized parameters of the feature extractor during testing.

9. A computer device comprising a memory and one or more processors, wherein the memory stores computer code, and when the computer code is executed by the one or more processors, the one or more processors execute the steps of the AI-generated military image detection method based on adaptive test-time training as described in any one of claims 1-8.

10. A computer-readable storage medium storing computer code, wherein when the computer code is executed, the steps of the method for generating military image detection based on AI trained during adaptive testing according to any one of claims 1 to 8 are performed.

Citation Information

Patent Citations

  • A deepfake detection method based on frequency domain filtering residual

    CN114757877B

  • Generalizable deep fake image detection method and system based on noise perception

    CN118196865B

Cited By

  • AIGC image fine granularity detection method based on adaptive layering mechanism

    CN121236569A

  • Test training method based on consistency regularization

    CN121599032A