Remote sensing image robust tampering positioning method based on segmentation all-model SAM

Through the dual-stream feature encoding and dynamic tampering signal enhancement module DTSEM based on segmented all model SAM, combined with the weighted loss function, the problem of degradation of detection performance in post-processing operations is solved, and robustness and high-precision detection of complex texture backgrounds are achieved.

CN120298404AActive Publication Date: 2025-07-11NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Application Number
CN202510775332.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing remote sensing image tamper detection methods have deteriorated after facing post-processing operations, making it difficult to distinguish between normal texture changes and tamper traces, and lack of explicit tamper trace guidance mechanism, resulting in insufficient detection accuracy and robustness.

Method used

Using a method based on segmentation of all models, the detection capability of multi-scale tampering regions is improved through dual-stream feature encoding, dynamic tampering signal enhancement module DTSEM and parallel branch forgery decoder PFD, combining weighted cross-parallel loss and weighted binary cross-entropy loss.

Benefits of technology

Effectively capture tampering traces that are weakened by post-processing, enhance the robustness of complex texture backgrounds, reduce false positive rates, and improve detection accuracy and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298404A_ABST
    Figure CN120298404A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image robust tampering positioning method based on segmentation all model SAM. The method comprises the following steps: step 1, acquiring an input remote sensing image I and a noise image R; step 2, inputting the remote sensing image I and the noise image R into an optical flow SAM encoder and a noise flow SAM encoder respectively, wherein the optical flow SAM encoder and the noise flow SAM encoder form a double-flow SAM encoder sharing parameters; step 3, performing dynamic calibration on the multi-level semantic features through a dynamic tampering signal enhancement module DTSEM; step 4, inputting the calibration features into a parallel branch counterfeit decoder (PFD); 5, optimizing a loss function; and step 6, reasoning stage processing. According to the method, the high sensitivity of the SAM to the feature distribution difference is utilized, the tampering trace weakened by post-processing can be effectively captured, and the positioning accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image processing, and particularly to a robust tampering localization method for remote sensing images based on the Segment Anything Model (SAM). Background Art

[0002] With the rapid development of remote sensing technology, high-resolution remote sensing images have been widely used in fields such as land resource survey, urban planning, and environmental monitoring. However, with the popularization of image editing technology, the act of tampering with remote sensing images has increased day by day, seriously threatening the reliability and security of remote sensing data. Therefore, it is of great practical significance to develop efficient and accurate remote sensing image tampering detection technology.

[0003] Currently, remote sensing image tampering detection methods are mainly divided into traditional feature-based methods and deep learning-based methods. Traditional methods mainly rely on manually designed features such as noise residuals, JPEG compression artifacts, etc., but these methods perform poorly in the face of complex scenes and diverse tampering types. With the development of deep learning technology, tampering detection methods based on convolutional neural networks (CNNs) and Transformers have gradually become a research hotspot.

[0004] In the field of remote sensing image processing, multi-scale feature extraction and fusion are key technologies to improve detection accuracy. CN117809198A discloses a remote sensing image saliency detection method based on a multi-scale feature aggregation network. This method uses dilated convolutions with different dilation rates and feature attention guidance operations to fuse different-scale feature information of the backbone network branches, and uses deformable convolutions through a feature alignment module for feature alignment aggregation, enhancing the model's perception ability of targets at different scales. This multi-scale feature fusion strategy provides a useful reference for remote sensing image tampering detection.

[0005] The double-branch network structure also shows good performance in remote sensing image processing. CN118314353B proposes a remote sensing image segmentation method based on double-branch multi-scale feature fusion, which uses two parallel branches of CNN and Transformer to extract local features and global features at different resolutions respectively, achieving effective integration of global-local information. Similarly, CN119540558A introduces a Transformer remote sensing semantic segmentation method based on spatial-channel cross-decoding. It decodes spatial features and channel features through a two-stream decoder network, and uses a deformable attention mechanism to capture the context dependencies of spatial and channel features, enhancing the feature representation ability.

[0006] In the field of image forgery detection, CN117893858A discloses an image forgery localization method that fuses multi-level multi-scale and boundary information. This method uses a pyramid vision Transformer backbone network to extract multi-level forgery features, enhances the feature representation through a multi-scale forgery feature enhancement module, and uses a forgery boundary information module to specifically model the boundary information of the forgery area. This method of fusing multi-level features and boundary information is of great significance for improving the accuracy of forgery detection.

[0007] In addition, CN117612029B proposes a remote sensing image object detection method based on progressive feature smoothing and scale-adaptive dilated convolution. By constructing an adaptive feature extraction network, this method effectively reduces the missed detection rate in remote sensing images and improves the object detection accuracy. This scale-adaptive design idea also has important reference value for remote sensing image forgery detection.

[0008] However, the existing remote sensing image forgery detection methods still have the following problems: First, the existing methods have poor perception ability for the distribution differences of subtle features. Especially after the image undergoes post-processing operations (such as Gaussian blur, JPEG compression, etc.), the original forgery traces are weakened, resulting in a significant decline in detection performance. The existing methods lack robustness to such post-processing interference and are difficult to detect the weakened forgery traces.

[0009] Second, the textures of remote sensing images are complex and diverse, including a large number of natural textures and artificial structures. The existing methods are easily interfered by these complex textures and are difficult to effectively distinguish normal texture changes from forgery traces, resulting in a high false alarm rate. Especially in high-resolution remote sensing images, this texture interference problem is more prominent.

[0010] Third, the existing methods lack an explicit forgery trace guidance mechanism and mainly rely on the feature extraction ability of the model itself. When facing post-processed forgery images, they are prone to converge to sub-optimal solutions and are difficult to effectively explore and extract the covered forgery traces. This lack of targeted feature extraction strategy limits the detection performance of the model in complex scenarios.

[0011] Therefore, it is urgent to develop a remote sensing image forgery detection method that can effectively cope with post-processing interference, complex texture backgrounds, and has an explicit forgery trace guidance ability to improve the accuracy and robustness of detection. Summary of the Invention

[0012] Object of the Invention: The technical problem to be solved by the present invention is to provide a robust forgery localization method for remote sensing images based on the Segment Anything Model (SAM) in view of the deficiencies of the prior art, including the following steps: Step 1, Feature Input and Preprocessing: Obtain the input remote sensing image , where H and W represent the height and width of the image respectively, and the number of channels of the image is 3; generate the corresponding noise image through the Steganalysis Rich Model (SRM) filter of the rich steganography analysis model , construct a two-stream input containing optical information and high-frequency noise information, where represents the real number space; Step 2, Two-Stream Feature Encoding: Input the remote sensing image I and the noise image R into the optical flow SAM encoder and the noise flow SAM encoder respectively. The optical flow SAM encoder and the noise flow SAM encoder form a two-stream SAM encoder with shared parameters; The optical flow SAM encoder and the noise flow SAM encoder have the same structure, both containing 4 modules, and the module serial number is represented by i, ; all 4 modules adopt the same structural design, and each module is composed of an Adaptive Multi-Scale Feature Adapter (AMFA) and a Transformer in cascade. Among them, the Adaptive Multi-Scale Feature Adapter AMFA processes the input features through two or more groups of convolutional kernels with different scales to capture feature information at different scales; the Transformer layer globally models the features based on the self-attention mechanism to enhance the feature expression ability; The optical flow SAM encoder and the noise flow SAM encoder respectively output multi-level features and , where is the multi-level semantic feature output after the optical flow (i.e., the input is the remote sensing image I) passes through the i-th module in the optical flow SAM encoder. Optical represents optics, corresponding to the semantic features extracted from the remote sensing image input; is the multi-level tampering noise feature output after the noise flow (i.e., the input is the noise image R) passes through the i-th module in the noise flow SAM encoder, where noise represents noise, corresponding to the tampering noise features extracted from the noise image input; Step 3, Dynamic Tampering Signal Enhancement: Through the Dynamic Tampering Signal Enhancement Module (DTSEM), use the multi-level tampering noise feature to dynamically calibrate the multi-level semantic feature , calculate the feature weight through the gating mechanism, and output the calibrated feature , where Conv represents the convolution operation, is the convolutional feature of the noise flow, realizing the enhancement of the effective tampering signal and the suppression of texture interference; Step 4, Parallel Branch Forgery Decoding: Input the calibrated feature Input to the parallel branch forgery decoder PFD; the parallel branch forgery decoder PFD includes four parallel branches, namely the first branch, the second branch, the third branch, and the fourth branch; Among them, the first branch includes a 1×1 convolution, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The second branch includes a 3×3 dilated convolution with a dilation rate of 6, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The third branch includes a 3×3 dilated convolution with a dilation rate of 12, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The fourth branch includes a 3×3 dilated convolution with a dilation rate of 18, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; Extract multi-scale tampering features through convolutional branches with different dilation rates, and combine the deep supervision mechanism to perform pixel-level classification on the outputs of each layer in the parallel branch forgery decoder PFD, and finally output a tampering region mask that fuses multi-scale information ; Based on steps 1 to 4, the construction of the dual-stream SAM model is completed; Step 5, loss function optimization: To effectively improve the performance of the model in detecting tampering regions in remote sensing images, a combined loss of weighted intersection over union loss (wIoU) and weighted binary cross-entropy loss (wBCE) is adopted.

[0013] In step 6, the dual-stream SAM model performs inference phase processing: In the inference phase, the dual-stream SAM model supports dynamic input sizes. For remote sensing images with different resolutions, first, through the interpolation preprocessing method, the size of the remote sensing image is uniformly adjusted. For images in the Post - FakeV dataset, it is unified to size; for images in the Post - FakeL dataset, it is unified to size; The dual-stream SAM model performs inference calculations based on the adjusted image, and after obtaining the result, the output is restored to the original size of the image through interpolation operations.

[0014] In step 1, apply a 3×3 SRM filter bank to the remote sensing image I to extract the residual signal containing the high-frequency noise pattern introduced by tampering. The formula is: , where the SRM filter bank contains edge detection kernels in the horizontal, vertical, and diagonal directions, which are used to explicitly model the noise inconsistency during image acquisition and tampering.

[0015] In step 2, the adaptive multi-scale feature adapter AMFA includes a multi-scale feature extraction unit and a hierarchical interaction fusion unit; The adaptive multi-scale feature adapter AMFA performs the following operations: Input feature First, it is dimension-reduced by a linear projection layer, and low-dimensional features are obtained through GeLU activation and reshaping operations. The formula is: , where is the dimension reduction matrix, and RS is the reshaping operation; Subsequently, the features pass through 3×3 dilated convolutions with three different dilation rates of 6, 12, and 18 respectively to extract multi-scale information, and at the same time, global context features are obtained through global average pooling, which is expressed as: , where represents a 1×1 convolution, are dilated convolutions with dilation rates of 6, 12, and 18 respectively, corresponding to ; AAP is the adaptive average pooling layer, and Up is the bilinear interpolation upsampling; finally, the multi-scale features and the original features are concatenated in channels to form a feature set containing multi-scale information; The hierarchical interaction fusion unit performs cross-layer fusion on the multi-scale features, and the formula is: , where, represents the j-th feature at the l-th level after hierarchical interaction fusion; is the j-th original multi-scale feature at the l-th level; is the activation function, used to introduce non-linearity; represents a 3×3 convolution operation; The feature aggregation and restoration unit integrates the multi-scale features after hierarchical interaction and restores the dimension, and the formula is: , , where, represents the aggregated feature obtained at the l-th level after a series of feature processing operations (including multi-scale feature extraction, hierarchical interaction fusion, etc.), which is the overall feature representation after comprehensive processing of multiple related features at this level. is the linear projection layer, which restores the feature dimension to the original input dimension C. is the final output feature obtained after multi-step feature processing and dimension restoration at the l-th level.

[0016] In step 2, the Transformer layer adopts a hierarchical feature extraction architecture. In the optical flow SAM encoder and the noise flow SAM encoder, the first module and the second module are bottom-layer modules, and the third module and the fourth module are high-layer modules; The bottom-layer modules focus on local texture differences, and the high-layer modules capture global semantic distributions. The representation migration from semantic features to tampering features is achieved through the Adaptive Multi-Scale Feature Adapter (AMFA).

[0017] In step 3, the gating mechanism includes: Performing convolution (Convolution), batch normalization (BatchNormalization), and PReLU activation (using the parametric rectified linear unit PReLU) on the multi-level tampering noise features to generate transition features ; Building a two-layer feature transformation network with the same structure for both layers, and sequentially performing convolution, batch normalization, and PReLU activation operations respectively; After concatenating the transition features with the multi-level semantic features along the channels and inputting them into the two-layer feature transformation network, after two rounds of convolution, batch normalization, and PReLU activation operations, and then through the global average pooling (GAP, Global AveragePooling) operation, a gating value is generated; The global average pooling operation performs an average calculation on the feature map in the spatial dimension to obtain a fixed-length feature vector, and through such a process, the adaptive selection of tampering-related noise features is achieved.

[0018] In step 4, the first branch, the second branch, the third branch, and the fourth branch are respectively used to capture local detail features, medium-scale features, larger-scale features, and global context features, and through skip connections, the shallow feature maps that retain high-resolution details and the deep feature maps that contain abstract semantic information are fused in a way of channel concatenation (Concat) or weighted summation.

[0019] In step 5, a weighted factor in the loss function is designed based on the local neighborhood mean : , where is the weighted factor of the pixel at the image coordinate in the loss function, used to adjust the weight of this pixel during loss calculation; is the neighborhood mean of the pixel at the coordinate in the ground truth mask ; is the pixel value at the coordinate in the real mask G at the position.

[0020] In step 6, for the input remote sensing images with different resolutions, the dual-stream SAM model will first uniformly adjust the size of the remote sensing images through interpolation preprocessing. For the images in the Post - FakeV dataset, it will be unified to size; for the images in the Post - FakeL dataset, it will be unified to size; the dual-stream SAM model performs inference calculations based on the adjusted images, and after obtaining the results, it restores the output to the original size of the image through interpolation operations.

[0021] The present invention also provides an electronic device, including a processor and a memory. The memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method described above.

[0022] The present invention also provides a storage medium, storing a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method.

[0023] The present invention has the following beneficial effects: High sensitivity of SAM: Utilizing the high sensitivity of SAM to the difference in feature distribution, it can effectively capture the tampering traces weakened by post-processing and improve the positioning accuracy.

[0024] Domain adaptation ability of AMFA: Bridging the gap between natural images and remote sensing images, semantic and non-semantic features, and enhancing the pertinence of feature extraction.

[0025] Anti-interference ability of DTSEM: Dynamically suppressing texture interference based on the gating mechanism, explicitly enhancing the tampering signal, and reducing the false alarm rate.

[0026] Multi-scale modeling of PFD: Combining deep supervision to extract multi-scale features and enhancing the detection ability for tampering regions of different scales.

[0027] Verification of post-processing robustness: Experiments on the Post-FakeV / Post-FakeL datasets show that the method of the present invention is significantly superior to existing methods and has stronger generalization ability. Description of the Drawings

[0028] Figure 1 is a flow chart of the method of the present invention, showing the overall process from creating the dataset to finally verifying the advantages of the method.

[0029] Figure 2It is the architecture diagram of the method described in the present invention, presenting the interconnection relationships among the dual-stream SAM encoder, AMFA, DTSEM, and PFD, as well as the processing paths of the input image, noise image, and mask.

[0030] Figure 3 It is the structural schematic diagram of the Adaptive Multi-scale Feature Adapter (AMFA), which details its internal process of multi-scale feature processing through operations such as convolution with different dilation rates and pooling.

[0031] Figure 4 It is the structural schematic diagram of the Dynamic Tampering Signal Enhancement Module (DTSEM), depicting the gating mechanism and feature fusion process for dynamically calibrating the optical flow features using the noise flow features.

[0032] Figure 5 It is the structural schematic diagram of the Parallel Branch Forgery Decoder (PFD), presenting its process of extracting multi-scale tampering features through multi-branch convolution operations and processing them in combination with the deep supervision mechanism. Specific implementation manners

[0033] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific implementation manners, and the above and / or other advantages of the present invention will become clearer.

[0034] As Figure 1 shown, in the first embodiment of the present invention, a remote sensing image robust tampering localization method based on the Segment Anything Model (SAM) is provided, including the following steps: Including the following steps: Step 1, Feature input and preprocessing: Obtain the input remote sensing image , where H and W respectively represent the height and width of the image, and the number of channels of the image is 3; generate the corresponding noise image through the Steganalysis Rich Model (SRM) filter of the rich steganalysis model , construct a dual-stream input containing optical information and high-frequency noise information, where represents the real number space; Step 2, Dual-stream feature encoding: As Figure 2 shown, input the remote sensing image I and the noise image R into the optical flow SAM encoder and the noise flow SAM encoder respectively. The optical flow SAM encoder and the noise flow SAM encoder form a dual-stream SAM encoder with shared parameters; The optical flow SAM encoder and the noise flow SAM encoder have the same structure, both including 4 modules, and the module numbers are represented by i, , all 4 modules adopt the same structural design. Each module is composed of an Adaptive Multi-scale Feature Adapter (AMFA) and a Transformer in cascade. Among them, the Adaptive Multi-scale Feature Adapter AMFA processes the input features through two or more groups of convolutional kernels with different scales to capture feature information at different scales; the Transformer layer globally models the features based on the self-attention mechanism to enhance the feature expression ability. Although these 4 modules have the same structure, at different levels, they will learn and adjust their respective parameters according to the differences in the resolution and semantic levels of the feature maps they process. As the level i increases, the resolution of the feature map gradually decreases, while the semantic information gradually increases, thus realizing the gradual extraction and conversion from low-level detailed features to high-level semantic features. In this way, the optical flow SAM encoder and the noise flow SAM encoder respectively output multi-level features and , where are the multi-level semantic features output after the optical flow (i.e., the input is the remote sensing image I) passes through the i-th module in the optical flow SAM encoder. Optical represents optics, corresponding to the semantic features extracted from the remote sensing image input; are the multi-level tampering noise features output after the noise flow (i.e., the input is the noise image R) passes through the i-th module in the noise flow SAM encoder. Among them, noise represents noise, corresponding to the tampering noise features extracted from the noise image input, realizing the parallel modeling of semantic features and tampering noise features; Step 3, Dynamic Tampering Signal Enhancement: As Figure 4 shown, through the Dynamic Tampering Signal Enhancement Module (DTSEM), the multi-level tampering noise features are used to dynamically calibrate the multi-level semantic features . The feature weights are calculated through the gating mechanism, and the calibrated features are output, where Conv represents the convolution operation, is the convolutional feature of the noise flow, realizing the enhancement of the effective tampering signal and the suppression of texture interference; Step 4, Parallel Branch Forgery Decoding: As Figure 5 shown, the calibrated features are input into the Parallel Branch Forgery Decoder PFD; the Parallel Branch Forgery Decoder PFD includes four parallel branches, namely the first branch, the second branch, the third branch, and the fourth branch; Among them, the first branch includes a 1×1 convolution, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The second branch includes a 3×3 dilated convolution with a dilation rate of 6, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The third branch includes a 3×3 dilated convolution with a dilation rate of 12, a conventional convolutional layer, a normalization layer, an activation function layer, and a computational layer for deep supervision; The fourth branch includes a 3×3 dilated convolution with a dilation rate of 18, a conventional convolutional layer, a normalization layer, an activation function layer, and a computational layer for deep supervision; Multi-scale tampering features are extracted through convolutional branches with different dilation rates, and a deep supervision mechanism is combined to perform pixel-level classification on the outputs of each layer in the parallel-branch forgery decoder PFD, and finally a tampering region mask that fuses multi-scale information is output. ; Based on steps 1 to 4, the construction of the dual-stream SAM model is completed; Step 5, loss function optimization: To effectively improve the performance of the model in detecting tampering regions in remote sensing images, a combined loss of weighted intersection over union loss (wIoU) and weighted binary cross-entropy loss (wBCE) is adopted.

[0035] Among them, the total loss function comprehensively considers the prediction performance of the model at different levels. Different levels of the parallel-branch forgery decoder (PFD) , and the feature maps output by it reflect the prediction results of this level for the tampering regions in the image.

[0036] When designing the loss function, a weighting factor is introduced. For the weighted binary cross-entropy loss function , the weighting factor adjusts the weights of pixels during loss calculation to alleviate the class imbalance problem and make the model pay more attention to small target tampering regions. It is calculated based on information such as the mean of the pixel neighborhood in the true mask G, and measures the difference degree between the pixel-level probability predicted by the model and the true mask in terms of classification.

[0037] The weighted intersection over union loss function then uses the weighting factor to measure the overlap difference between the model prediction region and the true tampering region. Especially in the case where the tampering region is small and the proportion of the true region is large, it can enhance the model's attention to different regions and improve the ability to identify small-scale tampering regions. The true mask serves as a standard reference accurately representing the actual tampering regions in the image and provides a reliable basis for loss calculation.

[0038] In addition, to enhance the model's ability to capture details and identify forgery traces at different scales, a deep supervision mechanism is introduced for the outputs of the first three layers of the PFD. Through this hierarchical supervision method, the model obtains effective feedback at different levels, thereby enhancing the perception ability of complex tampering patterns and finally better achieving accurate detection of tampering regions in remote sensing images.

[0039] Step 6, Inference Phase Processing: In the inference phase, the dual-stream SAM model supports dynamic input sizes. For remotely sensed images with different resolutions as inputs, first, through the interpolation preprocessing method, the sizes of the remotely sensed images are uniformly adjusted. For images in the Post-FakeV dataset, they are unified to size; for images in the Post-FakeL dataset, they are unified to size; the model performs inference calculations based on the adjusted images. After obtaining the results, the output is restored to the original size of the image through interpolation operations; this processing method can ensure that the model has the ability to accurately locate tampered areas in high-resolution remotely sensed images, effectively handle the diverse resolutions of remotely sensed images in practical applications, and improve the practicality and effectiveness of the model in different scenarios.

[0040] In Step 1, the input remotely sensed image I can be satellite or aerial remotely sensed images of various resolutions, including but not limited to visible light bands, infrared bands, or multispectral remotely sensed images. In this embodiment, the resolution of the input image can be 352×352 pixels (for the Post-FakeV dataset) or 512×512 pixels (for the Post-FakeL dataset) according to the dataset type. Apply a 3×3 SRM filter bank to the remotely sensed image to extract the residual signal containing the high-frequency noise pattern introduced by tampering. The formula is: , where the SRM filter bank contains edge detection kernels in the horizontal, vertical, and diagonal directions, which are used to explicitly model the noise inconsistency during image acquisition and tampering.

[0041] In Step 2, as Figure 3 shown, the Adaptive Multi-Scale Feature Adapter (AMFA) includes a multi-scale feature extraction unit and a hierarchical interaction and fusion unit; The multi-scale feature extraction unit extracts multi-scale features containing local details and global context through 1×1 convolution, dilated convolutions with dilation rates of 6 / 12 / 18, and global average pooling. The input feature is first dimensionally reduced by a linear projection layer, and through GeLU activation and reshaping operations, a low-dimensional feature is obtained. The formula is: , where is the dimensional reduction matrix (dimensional reduction factor r = 4), and RS is the reshaping operation; Subsequently, the feature passes through 3×3 dilated convolutions with three different dilation rates of 6, 12, and 18 respectively to extract multi-scale information, and at the same time, global context features are obtained through global average pooling, expressed as: , wherein represents a 1×1 convolution, and RConv is a dilated convolution, denotes dilated convolutions with dilation rates of 6, 12, and 18 respectively, corresponding to ; AAP is an adaptive average pooling layer, and Up is a bilinear interpolation upsampling; finally, the multi-scale features are concatenated with the original feature channels to form a feature set containing multi-scale information; The hierarchical interaction and fusion unit performs cross-layer fusion on the multi-scale features, and the formula is: , wherein represents the th feature at the l-th level after hierarchical interaction and fusion; is the th original multi-scale feature at the l-th level; is an activation function used to introduce non-linearity; represents a 3×3 convolution operation. When , the fused feature is the original multi-scale feature; when , the fused feature is obtained by first adding the th and the (j - 1)-th original multi-scale features, then performing a 3×3 convolution and ReLU activation processing, realizing the hierarchical interaction of different-scale features and enhancing the sensitivity to subtle tampering traces.

[0042] The feature aggregation and restoration unit integrates the multi-scale features after hierarchical interaction and restores the dimension, and the formula is: , , wherein is a linear projection layer that restores the feature dimension to the original input dimension . Finally, the processed feature is added to the original input feature through a residual connection, which not only retains the multi-scale interaction information but also avoids the degradation of feature representation.

[0043] In step 2, the Transformer layer adopts a hierarchical feature extraction architecture. In the optical flow SAM encoder and the noise flow SAM encoder, the first module and the second module are low-level modules, and the third module and the fourth module are high-level modules; The low-level modules (such as ) focus on pixel-level texture differences through small-scale convolution kernels and local attention mechanisms to capture subtle structural anomalies in the tampered areas; the high-level modules (such as ), leveraging large-scale dilated convolutions and global self-attention mechanisms to model scene-level semantic distributions and identify global semantic inconsistencies caused by tampering. The processing flow of each module is as follows: the features output by the upper layer are first subjected to multi-scale fusion by AMFA (including 1×1 convolution, dilated convolutions with dilation rates of 6, 12, and 18, and global average pooling), and then cross-position dependence modeling is achieved through the Transformer layer, and finally the feature representation of the current layer is output. Among them, the first-layer module directly uses the original image features (optical flow or noise flow ) as input, and gradually constructs a feature pyramid from low-level details to high-level semantics to provide multi-granularity feature support for subsequent tampering localization.

[0044] In step 3, the gating mechanism includes: Performing convolution (Convolution), batch normalization (BatchNormalization), and activation (using the parametric rectified linear unit PReLU) operations on the multi-level tampering noise features to generate transitional features ; this series of operations is simply referred to as convolution, batch normalization, parametric rectified linear unit activation operation (CBPR). Specifically, the convolution operation uses a convolution kernel to extract local feature information from the multi-level tampering noise features ; the batch normalization operation normalizes the convolved features in the batch dimension to make the data distribution more stable, which helps to accelerate model training and alleviate the gradient problem; the parametric rectified linear unit PReLU (ParametricRectified Linear Unit) in the activation operation, as an activation function, introduces learnable parameters to adaptively adjust the function slope, adding non-linearity factors to the model and enhancing its expressive ability.

[0045] Build a two-layer feature transformation (CBPR) network. The two-layer feature transformation network structures are the same, and they sequentially perform convolution, batch normalization, and PReLU activation operations respectively; After concatenating the transitional features with the multi-level semantic features in the channel dimension, input them into the two-layer feature transformation network. After two rounds of convolution, batch normalization, and PReLU activation operations, and then through the global average pooling GAP (Global AveragePooling) operation, generate a gating value ; the global average pooling operation will perform an average calculation on the feature map in the spatial dimension to obtain a fixed-length feature vector, and through such a process, the adaptive selection of tampering-related noise features is realized.

[0046] In step 4, the first branch, the second branch, the third branch, and the fourth branch are respectively used to capture local detailed features (for detecting small-scale tampering, such as pixel-level forgery, slight erasure, etc., and the area of such tampered regions usually accounts for less than 5% of the total area of the image), medium-scale features (able to detect edge tampering, such as image splicing traces and texture anomalies, and the area of the tampered region generally ranges from 5% to 20% of the total area of the image), larger-scale features (suitable for locating medium-sized tampered regions, such as regional replacement, content copy-move, and the area of the tampered region is approximately between 20% and 50% of the total area of the image), and global context features (able to identify large-scale semantic inconsistencies, such as whole-region synthesis, scene tampering, and the area of the tampered region usually accounts for more than 50% of the total area of the image). Through skip connections, shallow feature maps that retain high-resolution details (such as edges, textures, from the early layers of the encoder or the low-level branches of the PFD) and deep feature maps that contain abstract semantic information (such as objects, scenes, from the high-level branches of the PFD) are fused by channel concatenation (Concat) or weighted summation to improve the localization accuracy for complex tampering patterns, enabling the PFD to simultaneously focus on multi-scale tampering features from the pixel level to the scene level and effectively handle diverse forgery methods in remote sensing images.

[0047] In step 5, the weighted binary cross-entropy loss (wBCE) aims to optimize pixel-level classification accuracy and alleviate the class imbalance problem. The binary cross-entropy loss (BCE) is commonly used in pixel-level classification tasks to measure the difference between the predicted probability map and the true mask, and the formula is: , where, represents the binary cross-entropy loss. S represents the prediction result of the model, and the superscript i represents different levels of the parallel branch forgery decoder (PFD), i.e., the prediction result of the th level. G is the mask. To highlight the key regions and alleviate the class imbalance, a weighted factor is innovatively introduced to obtain the weighted binary cross-entropy loss : , The weighted factor is calculated using the following formula , where is the weighted factor of the pixel at the image coordinate in the loss function, used to adjust the weight of this pixel during loss calculation to alleviate the class imbalance problem and enhance the attention to small target tampered regions; is the coordinate The pixel value at a certain position reflects whether the coordinate at this level in the image is a prediction result of whether the position is a tampered area. is the ground truth mask at the coordinate in, the neighborhood mean of the pixel, that is, calculate the average value of the pixel values in the neighborhood of size centered on this pixel; is the pixel value at the coordinate in the ground truth mask G, identifying whether this position belongs to the tampered area. By calculating the weighting factor in this way, the weights can be reasonably allocated during the loss calculation, enabling the model to pay more attention to small target tampered areas and effectively alleviating the impact brought by class imbalance.

[0048] The intersection over union (IoU) is a key metric for measuring the overlap degree between the predicted area and the mask area. The weighted intersection over union loss (wIoU) helps to enhance the model's attention to different areas. Especially when the tampered area is small and the real area occupies most of the image, it can improve the model's ability to identify small-scale tampered areas: , , , where i here represents different levels of the parallel branch forgery decoder (PFD), is for all pixel positions, according to the weighting factor , the model prediction value and the ground truth mask value to calculate an intermediate quantity, measuring the intersection part between the prediction result and the ground truth mask considering the weighting. Similarly, for all pixel positions, based on the weighting factor , the model prediction value and the ground truth mask value to calculate an intermediate quantity, measuring the union part between the prediction result and the ground truth mask considering the weighting. represents calculating the weighted intersection over union loss for the prediction result and the ground truth mask at the i-th level of the parallel branch forgery decoder (PFD).

[0049] The overall loss function is and combined, and the expression is: , In addition, to enhance the model's ability to capture details and identify forgery traces at different scales, a deep supervision mechanism is introduced for the segmentation output of each layer. Through this layer-by-layer guidance, the model obtains effective feedback at different levels, thereby enhancing its perception ability of complex forgery patterns. The final total loss function is defined as: .

[0050] In step 6, the inference stage of the SAM model supports dynamic input sizes and is made compatible with remote sensing images of different resolutions through interpolation preprocessing. This process uses methods such as bilinear interpolation to scale the images, reducing information loss while ensuring feature integrity, ensuring that the model can effectively extract tampering features at different scales, enhancing the ability to capture subtle tampering traces in high-resolution remote sensing images, and enhancing the generalization performance of the model in actual complex scenarios.

[0051] Training and inference strategies: In one specific embodiment of the present invention, during the training process, a progressive enhancement strategy is adopted: first, pre-train on the original dataset without post-processing, and then fine-tune on the Post-FakeV / Post-FakeL dataset to improve the model's robustness to post-processing interference. Specifically, in the pre-training stage, the original remote sensing image tampering dataset is used to learn basic tampering feature representations; in the fine-tuning stage, the constructed post-processing datasets Post-FakeV and Post-FakeL are used, which contain single post-processing operations and comprehensive post-processing operations. Single post-processing operations include Gaussian blur (radius 0.1, 0.3, 0.5, 1), Gaussian noise (variance 3, 9, 15, 20), JPEG compression (quality 30, 50, 70, 90); comprehensive post-processing operations include (blur radius 0.3)+(noise variance 9)+(JPEG quality 70) to simulate multiple interferences in real attack scenarios.

[0052] In the inference stage of the model, dynamic input sizes are supported. Remote sensing images of different resolutions are unified to 352×352 (Post-FakeV) or 512×512 (Post-FakeL) through interpolation preprocessing and restored to the original size during output to ensure the tampering localization ability for high-resolution remote sensing images. Specifically, for any size of input remote sensing image, it is first adjusted to the standard size through bilinear interpolation, then the tampering mask prediction is obtained through model processing, and finally the prediction result is interpolated back to the original size to maintain consistency with the input image.

[0053] Experimental results: The method of this embodiment exhibits excellent tampering localization performance under various post - processing interference conditions. For example, under the Gaussian blur (radius 0.5) condition, the F1 - score of this method reaches 0.712, which is 8.5% higher than the existing state - of - the - art method; under the JPEG compression (quality 50) condition, the IoU of this method reaches 0.658, which is 7.2% higher than the existing method; under the Gaussian noise (variance 15) condition, the precision of this method reaches 0.735, which is 9.1% higher than the existing method. Especially under the comprehensive post - processing condition (blur radius 0.3 + noise variance 9 + JPEG quality 70), the F1 - score of this method reaches 0.679, significantly better than 0.598 of the existing method, proving the robustness of this method in complex interference environments.

[0054] By ablation experiments to analyze the contributions of each component, the results show that: after removing the AMFA module, the F1 - score drops by 4.3%; after removing the DTSEM module, the F1 - score drops by 5.7%; after removing the noise flow, the F1 - score drops by 7.2%; after using a single - branch decoder to replace the multi - branch decoder (PFD), the F1 - score drops by 3.8%. These results prove the important contributions of each component to the model performance, especially the key roles of the noise flow and the DTSEM module in enhancing the model robustness.

[0055] The present invention also provides a second embodiment. Based on the first embodiment, the implementation of the adaptive multi - scale feature adapter (AMFA) is optimized in this embodiment. AMFA includes a multi - scale feature extraction unit and a hierarchical interaction and fusion unit. The multi - scale feature extraction unit extracts multi - scale features containing local details and global context through 1×1 convolution, dilated convolutions with dilation rates of 6, 12, and 18, and global average pooling , where l represents the layer, and l takes values of 1, 2, 3, 4; m is an index used to distinguish different features, .

[0056] In this embodiment, the specific implementation of the multi - scale feature extraction unit is as follows: First, the input feature reduces the channel dimension to 128 through 1×1 convolution to obtain the feature ; then, respectively passes through three 3×3 dilated convolutions with different dilation rates (dilation rates are 6, 12, 18) to obtain the features , , ; at the same time, is globally averaged pooled and then processed through 1×1 convolution to obtain the global context feature ; finally, , , , , Concatenate on the channel dimension to form a multi-scale feature set.

[0057] The hierarchical interaction and fusion unit performs cross-layer fusion on the multi-scale features, and the formula is: , where is the weight coefficient for adaptive learning, which is calculated through the attention mechanism. Specifically, first, the multi-scale features are concatenated and then passed through a 1×1 convolution and the Softmax function to generate the weight coefficient . Then, weighted summation is performed on the features of each scale to obtain the final output feature . This hierarchical interaction and fusion mechanism can adaptively select and combine features of different scales, enhancing the sensitivity to subtle tampering traces.

[0058] Experimental results show that the optimized AMFA module can extract and fuse multi-scale features more effectively, improving the F1 score by 1.2% on the Post-FakeV dataset and the IoU by 1.5% on the Post-FakeL dataset, demonstrating its important contribution to improving the model performance.

[0059] The present invention also provides a third embodiment. On the basis of the first embodiment, the gating mechanism of the dynamic tampering signal enhancement module (DTSEM) is optimized in this embodiment. The gating mechanism of DTSEM includes performing a convolution-batch normalization-activation (CBPR) operation on the noise flow feature to generate a transition feature ; after concatenating with the optical flow feature on the channel dimension, passing through a two-layer CBPR network and global average pooling (GAP) to generate a gating value to achieve adaptive selection of tampering-related noise features.

[0060] In this embodiment, the specific implementation of the CBPR operation is as follows: First, apply a 3×3 convolution to the input feature to adjust the number of channels to half of the input channel number; then, perform batch normalization on the convolution result to stabilize the training process; finally, apply the ReLU activation function to introduce non-linearity. For the noise flow feature , after one CBPR operation, the transition feature is obtained.

[0061] The calculation process of the gating value is optimized as: First, and Concatenate on the channel dimension; then halve the number of channels through the first layer of CBPR, and the second layer of CBPR further reduces the number of channels to 64; then apply global average pooling to obtain a channel-level feature vector; finally, generate a gating value ranging from [0,1] through a 1×1 convolution and a Sigmoid activation function 。

[0062] Calibrated feature The calculation formula is: ,where is the noise flow feature processed by a 3×3 convolution. This gating mechanism can adaptively adjust the weights of the noise feature and the optical feature according to the content of the input feature, enhance the attention to the tampered area, and suppress the background interference.

[0063] The experimental results show that the optimized DTSEM module can more accurately identify and enhance the tampering-related noise features, improving the F1 score by 2.3% under Gaussian noise (variance 20) conditions and the IoU by 2.7% under JPEG compression (quality 30) conditions, demonstrating its important contribution to improving the robustness of the model.

[0064] The present invention also provides a fourth embodiment. On the basis of the first embodiment, this embodiment optimizes the implementation of the parallel branch forgery decoder (PFD). PFD adopts a parallel branch structure, and each branch contains dilated convolutions with different dilation rates for capturing (dilation rate = 1), (dilation rate = 6), (dilation rate = 12), (dilation rate = 18) scale tampering features, and fuses the shallow-layer details and the deep-layer semantic information through skip connections.

[0065] In this embodiment, the specific implementation of PFD is as follows: First, input the calibrated feature into four parallel branches respectively; each branch first adjusts the number of channels to 128 through a 1×1 convolution, and then extracts features of a specific scale through a 3×3 dilated convolution with the corresponding dilation rate (1, 6, 12, 18); then, concatenate the features of the four branches on the channel dimension, and adjust the number of channels to 256 through a 1×1 convolution to obtain the comprehensive feature 。

[0066] To fuse the shallow-layer details and the deep-layer semantic information, PFD adopts a skip connection mechanism: for the feature of the l-th layer, adjust its resolution to the resolution of the feature of the (l - 1)-th layer through an upsampling operation, and then combine it with They are concatenated and fused through 1×1 convolution to obtain a new feature representation. This hierarchical fusion starts from the deepest layer and proceeds layer by layer upward, finally obtaining a feature representation containing multi-level information.

[0067] During the decoding process, an intermediate prediction result is generated after each layer of features is processed , which is used for deep supervision. The final tampered region mask M is determined by the prediction result of the top layer and restored to the original image resolution through bilinear interpolation.

[0068] Experimental results show that the optimized PFD can extract and fuse multi-scale tampering features more effectively, with the tampering detection accuracy in complex texture regions increased by 3.5% and the recall rate in small target tampering regions increased by 4.2%, demonstrating its important contribution to improving the model's detection performance.

[0069] The present invention also provides a fifth embodiment. On the basis of the first embodiment, this embodiment details the construction of the post-processing datasets Post-FakeV and Post-FakeL. These datasets contain single post-processing operations and comprehensive post-processing operations for evaluating the robustness of the model under various interference conditions.

[0070] The single post-processing operations include:

[0071] Gaussian blur: The image is blurred using Gaussian kernels with radii of 0.1, 0.3, 0.5, and 1 respectively to simulate the blurred effect after image defocusing or compression.

[0072] Gaussian noise: Gaussian noise with variances of 3, 9, 15, and 20 is added to the image to simulate the noise interference during image acquisition or transmission.

[0073] JPEG compression: The image is compressed using the JPEG compression algorithm with quality factors of 30, 50, 70, and 90 respectively to simulate the compression distortion during image storage or transmission.

[0074] The comprehensive post-processing operation includes: a combined processing of a blur radius of 0.3 + a noise variance of 9 + a JPEG quality of 70 to simulate multiple interferences in a real attack scenario. Specifically, first, Gaussian blur (radius 0.3) is applied to the image, then Gaussian noise (variance 9) is added, and finally JPEG compression (quality 70) is performed.

[0075] The dataset construction process is as follows: First, collect the original remote sensing image tampering dataset, including real images and their corresponding tampered versions and real masks; then, apply the same post-processing operation to each pair of images (the original image and the tampered image) to generate the post-processing version; finally, pair the post-processed images with the original real masks to form new training and test samples.

[0076] The Post-FakeV dataset is constructed based on medium-resolution remote sensing images, and the image size is uniformly adjusted to 352×352 pixels; the Post-FakeL dataset is constructed based on high-resolution remote sensing images, and the image size is uniformly adjusted to 512×512 pixels. The two datasets contain 5000 and 3000 pairs of training samples, and 1000 and 600 pairs of test samples respectively.

[0077] The experimental results show that the models trained on these post-processed datasets exhibit stronger robustness and can effectively cope with various interference conditions. Especially under the comprehensive post-processing conditions, the F1 score of this method is 12.3% higher than that of the model trained on the original dataset, demonstrating the important role of the post-processed datasets in improving the model's robustness.

[0078] The present invention also provides a sixth embodiment. Based on the first embodiment, this embodiment details the implementation of the Transformer layer of the dual-stream SAM encoder. The Transformer layer of the dual-stream SAM encoder adopts a hierarchical feature extraction architecture. The underlying modules (l = 1, 2) focus on local texture differences, and the high-level modules (l = 3, 4) capture the global semantic distribution. The representation transfer from semantic features to tampered features is achieved through AMFA.

[0079] In this embodiment, each Transformer layer contains a multi-head self-attention mechanism and a feed-forward neural network. Specifically, for the input feature , first, it is processed by layer normalization; then, the attention weights and weighted features are calculated through the multi-head self-attention mechanism; next, the weighted features are added to the input features to form a residual connection; finally, it is processed by another layer normalization and a feed-forward neural network, and the residual connection is applied again to obtain the output feature .

[0080] The calculation process of the multi-head self-attention mechanism is as follows: First, the input feature is linearly projected into three representations: query (Q), key (K), and value (V); then, the dot product similarity between the query and the key is calculated, and the attention weights are obtained through scaling and the Softmax function; finally, the values are weighted and summed using the attention weights to obtain the attention output. To enhance the expressive ability, multiple attention heads are calculated simultaneously, and their outputs are concatenated and linearly projected to obtain the final output.

[0081] The feed-forward neural network contains two linear layers and a ReLU activation function, which are used to introduce non-linear transformations and enhance the feature expression ability. Specifically, the first linear layer expands the feature dimension to 4 times the original, and after ReLU activation, the second linear layer restores the dimension to the original dimension.

[0082] In the two-stream architecture, the optical stream and the noise stream share the parameters of the Transformer layer, but perform domain-specific adaptation through the AMFA module. This parameter sharing strategy not only reduces the number of model parameters, but also maintains the feature consistency between the two streams, which is beneficial to subsequent feature fusion and calibration.

[0083] Experimental results show that the hierarchical feature extraction architecture can effectively capture tampering features at different levels. The bottom-level module is 8.7% more sensitive to local texture differences than a single architecture, and the high-level module’s ability to understand global semantic distribution is 7.5% better, proving the important contribution of this architecture to improving model performance.

[0084] The present invention also provides a seventh embodiment, which, based on the first embodiment, describes in detail the progressive enhancement strategy used in the training process. The progressive enhancement strategy includes: first pre-training on the original data set that has not been post-processed, and then fine-tuning on the Post-FakeV / Post-FakeL data set to improve the robustness of the model to post-processing interference.

[0085] In this embodiment, the training process is divided into two stages: a pre-training stage and a fine-tuning stage.

[0086] Pre-training stage: Use the original remote sensing image tampering dataset to learn the basic tampering feature representation. Specifically, use the AdamW optimizer, with an initial learning rate of 0.001, a weight decay coefficient of 0.0005, a batch size of 12, and 50 epochs. The learning rate adopts a cosine annealing strategy and gradually decreases during the training process.

[0087] Fine-tuning stage: Use the constructed post-processing datasets Post-FakeV and Post-FakeL to adapt the model to various interference conditions. Specifically, use a smaller learning rate of 0.0001, keep other hyperparameters unchanged, and train for 30 epochs. In order to enhance the generalization ability of the model, a hybrid training strategy is adopted in the fine-tuning process, using both original data and post-processed data at a ratio of 1:2.

[0088] In addition, data enhancement techniques were also used during the training process, including random horizontal flipping, random vertical flipping, random rotation (±15 degrees), random brightness adjustment (±0.2), random contrast adjustment (0.8~1.2), etc., to increase the diversity of training data and improve the generalization ability of the model.

[0089] In order to deal with the problem of class imbalance, in addition to using the weighted loss function, a variant of focal loss (FocalLoss) is also used to adjust the weights of difficult and easy samples so that the model pays more attention to samples that are difficult to classify. Specifically, the calculation formula of focal loss is: , where represents the predicted probability, represents the class weight, represents the focusing parameter (set to 2 in this embodiment).

[0090] The experimental results show that the progressive enhancement strategy can effectively improve the robustness of the model to post - processing interference. Compared with the model directly trained on the post - processed dataset, the model adopting the progressive enhancement strategy has an F1 - score increase of 3.8% under the condition of Gaussian blur (radius 1), an IoU increase of 4.2% under the condition of JPEG compression (quality 30), and a precision increase of 5.1% under the condition of Gaussian noise (variance 20), proving the effectiveness of this training strategy.

[0091] The present invention also provides an eighth embodiment. Based on the first embodiment, in this embodiment, the calculation method of the weighting factor in the loss function is described in detail. The weighting factor is calculated based on the local neighborhood mean, and the formula is: , where is the 31×31 neighborhood mean of the true mask . The class imbalance problem is alleviated through the weighting mechanism, and the attention to the tampered regions of small targets is enhanced.

[0092] In this embodiment, the calculation process is as follows: Apply a 31×31 mean filter to the true mask to calculate the local neighborhood mean at each pixel position; Perform element - wise multiplication on the mean result and the original mask to obtain ;

[0093] Subtract from 1 to obtain the weighting factor .

[0094] The core idea of this weighting mechanism is: For the pixels inside the tampered region, most of the pixels in its 31×31 neighborhood are also in the tampered region. Therefore, is close to 1, and the corresponding value is close to 0; For the pixels at the edge of the tampered region, its neighborhood contains some non - tampered regions. Therefore, is between 0 and 1, and the corresponding value is also between 0 and 1; For the background pixels far from the tampered region, is close to 0, and the corresponding The value is close to 1.

[0095] In this way, the loss function will give higher weights to the edges of the tampered area and small-target tampered areas, enhancing the model's attention to these areas and improving the detection accuracy. Specifically, the calculation formulas of the weighted intersection over union loss (wIoU) and the weighted binary cross-entropy loss (wBCE) are as follows: , , where, represents the pixel index, and represent the value of the predicted mask at position i and the value of the ground-truth mask at position i, respectively, represents the weighting factor at position i.

[0096] Experimental results show that this weighting mechanism based on the local neighborhood mean can effectively improve the model's detection performance for small-target tampered areas and tampered edges. Compared with the loss function using fixed weights, the F1 score of the model using this weighting mechanism increases by 6.5% in small-target tampered areas (with an area less than 1% of the image area), and the IoU in the tampered edge area increases by 5.8%, demonstrating the effectiveness of this weighting mechanism.

[0097] The present invention also provides a ninth embodiment. On the basis of Embodiment 1, this embodiment details the generation process of the noise image . The generation process of the noise image includes: applying a 3×3 SRM filter bank to the remote sensing image I to extract the residual signal containing the high-frequency noise patterns introduced by tampering, with the formula: , where the SRM filter bank contains edge detection kernels in horizontal, vertical, and diagonal directions, which are used to explicitly model the noise inconsistency in the image acquisition and tampering processes.

[0098] In this embodiment, the SRM filter bank contains 30 3×3 filters, which are divided into three categories: The first category (No. 1-10): Detect horizontal edges and texture changes; The second category (No. 11-20): Detect vertical edges and texture changes; The third category (No. 21-30): Detect diagonal edges and texture changes. The weights of each filter are designed according to different direction and frequency characteristics, and can effectively capture the high-frequency noise patterns in the image.

[0099] The specific generation process of the noise image R is as follows: Convert the input image I into a grayscale image ; For Applying each SRM filter, 30 filtering results are obtained; Taking the absolute value of the filtering results and concatenating them in the channel dimension to form a 30-channel noise feature map; Adjusting the number of channels to 3 through 1×1 convolution to obtain a noise image R with the same number of channels as the original image.

[0100] The design of the SRM filter is based on the noise residual analysis theory, which can capture the tiny noise fingerprints introduced by image tampering and help the model identify the tampered areas. Experimental results show that after introducing the noise stream, the F1 score of the model increases by 6.2% under the Gaussian blur (radius 0.5) condition, the IoU increases by 5.7% under the JPEG compression (quality 50) condition, and the precision increases by 7.3% under the Gaussian noise (variance 15) condition, proving the important contribution of the noise image to improving the model performance.

[0101] The present invention also provides a tenth embodiment. On the basis of the first embodiment, this embodiment details the inference stage of the model. The inference stage of the model supports dynamic input sizes. The remote sensing images with different resolutions are unified to 352×352 (Post-FakeV) or 512×512 (Post-FakeL) through interpolation preprocessing and restored to the original size during output, ensuring the tampering localization ability for high-resolution remote sensing images.

[0102] The specific inference process is as follows: Preprocessing: For the input image Perform normalization (scaling pixel values to [0,1]) and size adjustment (bilinear interpolation to the standard size); Noise image generation: Generate a noise image through the SRM filter ; Model inference: Input I and R into the model, and through two-stream feature encoding, DTSEM calibration, and PFD decoding, obtain the tampering mask prediction M; Size restoration: Restore M to the original image size through bilinear interpolation.

[0103] Inference optimization strategies: Batch processing inference: Divide large images into overlapping small blocks, and fuse the results with weights after inference to avoid boundary discontinuity; Model quantization: Convert the parameters from 32-bit floating-point numbers to 16-bit / 8-bit to reduce memory occupancy and computational complexity; Self-attention caching: Cache the key-value pairs of the Transformer layer to avoid repeated calculations; Multi-scale testing: Generate 0.75×, 1×, and 1.25× scale versions of the input image, and average the inference results to improve robustness.

[0104] The experimental results show that the optimized inference speed is increased by 2.3 times, the memory occupancy is reduced by 35%, and multi-scale testing improves the F1 score by 1.8%, especially showing remarkable effects on complex scenes and small target detection.

[0105] The present invention provides a robust tampering localization method for remote sensing images based on the Segment Anything Model (SAM). There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A robust tampering localization method for remote sensing images based on the Segment Anything Model (SAM), characterized in that, It includes the following steps: Step 1, Feature Input and Preprocessing: Obtain the input remote sensing image , where H and W respectively represent the height and width of the image, and the number of channels of the image is 3; generate the corresponding noise image through the rich steganography analysis model SRM filter , construct a two-stream input containing optical information and high-frequency noise information, where represents the real number space; Step 2, Dual-stream feature encoding: Input the remote sensing image I and the noise image R into the optical flow SAM encoder and the noise flow SAM encoder respectively. The optical flow SAM encoder and the noise flow SAM encoder form a dual-stream SAM encoder with shared parameters; The optical flow SAM encoder and the noise flow SAM encoder have the same structure and both contain 4 modules, with the module number represented by i. ; All 4 modules adopt the same structural design. Each module is composed of an adaptive multi-scale feature adapter AMFA and a Transformer in cascade. Among them, the adaptive multi-scale feature adapter AMFA processes the input features through 4 groups of convolutional kernels with different scales to capture feature information at different scales; The Transformer layer performs global modeling on the features based on the self-attention mechanism to enhance the feature expression ability; The optical flow SAM encoder and the noise flow SAM encoder respectively output multi-level features and , where is the multi-level semantic feature output after the optical flow passes through the i-th module in the optical flow SAM encoder. "optical" represents optical, corresponding to the semantic feature extracted from the remote sensing image input; is the multi-level tampering noise feature output after the noise flow passes through the i-th module in the noise flow SAM encoder, where "noise" represents noise, corresponding to the tampering noise feature extracted from the noise image input; Step 3, Dynamic Tampering Signal Enhancement: Through the Dynamic Tampering Signal Enhancement Module (DTSEM), utilize multi-level tampering noise features to perform dynamic calibration on multi-level semantic features by calculating feature weights through a gating mechanism and output calibrated features where Conv represents a convolution operation is the convolution feature of the noise stream, achieving the enhancement of effective tampering signals and the suppression of texture interference; Step 4, Parallel Branch Forgery Decoding: Input the calibration feature into the parallel branch forgery decoder PFD; the parallel branch forgery decoder PFD includes four parallel branches, namely the first branch, the second branch, the third branch, and the fourth branch; Among them, the first branch includes a 1×1 convolution, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The second branch includes a 3×3 dilated convolution with a dilation rate of 6, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The third branch includes a 3×3 dilated convolution with a dilation rate of 12, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; The fourth branch includes a 3×3 dilated convolution with a dilation rate of 18, a conventional convolutional layer, a normalization layer, an activation function layer, and a calculation layer for deep supervision; Extract multi-scale tampering features through convolutional branches with different dilation rates, and combine the deep supervision mechanism to perform pixel-level classification on the outputs of each layer in the parallel branch forgery decoder PFD, and finally output a tampering region mask that fuses multi-scale information ; Based on steps 1-4, the construction of the two-stream SAM model is completed; Step 5, Loss function optimization: Adopt a combined loss of weighted intersection over union loss wIoU and weighted binary cross-entropy loss wBCE; Step 6, The dual-stream SAM model performs inference phase processing: In the inference phase, the dual-stream SAM model supports dynamic input sizes.

2. The method according to claim 1, wherein In step 1, apply a 3×3 SRM filter bank to the remote sensing image I to extract the residual signal containing the high-frequency noise pattern introduced by tampering. The formula is: , Among them, the SRM filter bank contains edge detection kernels in the horizontal, vertical, and diagonal directions, which are used to explicitly model the noise inconsistency during image acquisition and tampering.

3. The method according to claim 2, characterized in that, In step 2, the adaptive multi-scale feature adapter AMFA includes a multi-scale feature extraction unit and a hierarchical interaction and fusion unit; The adaptive multi-scale feature adapter AMFA performs the following operations: Input features First, it is dimensionally reduced by a linear projection layer, and low-dimensional features are obtained through GeLU activation and reshaping operations , and the formula is: , wherein is a dimensionality reduction matrix, and RS is a reshaping operation; Subsequently, the features pass through 3×3 dilated convolutions with three different dilation rates of 6, 12, and 18 respectively to extract multi-scale information, and at the same time, global context features are obtained through global average pooling, expressed as: , Among them represents a 1×1 convolution is a dilated convolution with dilation rates of 6, 12, and 18 respectively, corresponding to ; AAP is an adaptive average pooling layer, and Up is a bilinear interpolation upsampling; finally, the multi-scale features are concatenated with the original feature channels to form a feature set containing multi-scale information; The hierarchical interaction and fusion unit performs cross-layer fusion on the multi-scale features. The formula is: , Among them, represents the th feature at the l-th level after hierarchical interaction and fusion; is the j-th original multi-scale feature at the l-th level; is an activation function used to introduce non-linearity; represents a 3×3 convolution operation; The feature aggregation and recovery unit integrates and restores the dimensions of the multi-scale features after hierarchical interaction. The formula is: , , Among them, represents the aggregated feature obtained at the l-th level after the feature processing operation, which is the overall feature representation after comprehensively processing multiple related features at this level; is a linear projection layer that restores the feature dimension to the original input dimension C; is the final output feature obtained after feature processing and dimension restoration at the l-th level.

4. The method according to claim 3, characterized in that, In step 2, the Transformer layer adopts a hierarchical feature extraction architecture. In the optical flow SAM encoder and the noise flow SAM encoder, the first module and the second module are the bottom modules, and the third module and the fourth module are the top modules; The bottom modules focus on local texture differences, and the top modules capture global semantic distributions. The representation migration from semantic features to tampering features is realized through the adaptive multi-scale feature adapter AMFA.

5. The method according to claim 4, wherein In step 3, the gating mechanism includes: For multi-level tampering noise features perform convolution, batch normalization, and PReLU activation operations to generate transitional features ; Build a two-layer feature transformation network. The structures of the two-layer feature transformation networks are the same, and they respectively perform convolution, batch normalization, and PReLU activation operations in sequence. The transition features and the multi-level semantic features are concatenated in channels and then input into a two-layer feature transformation network. After two rounds of convolution, batch normalization, and PReLU activation operations, and then through the global average pooling (GAP) operation, a gating value is generated; the global average pooling operation performs an average calculation on the feature map in the spatial dimension to obtain a feature vector of a fixed length.

6. The method according to claim 5, wherein In step 4, the first branch, the second branch, the third branch, and the fourth branch are respectively used to capture local detail features, medium-scale features, larger-scale features, and global context features. Through skip connections, the shallow feature maps that retain high-resolution details and the deep feature maps that contain abstract semantic information are fused by means of channel concatenation Concat or weighted summation.

7. The method according to claim 6, characterized in that, In step 5, design the weighting factor in the loss function based on the local neighborhood mean : , Among them is the weighting factor of the pixel located at the image coordinate in the loss function, which is used to adjust the weight of this pixel during loss calculation; is the ground truth mask at the coordinate of the pixel neighborhood mean; is the pixel value at the coordinate in the ground truth mask G.

8. The method according to claim 7, wherein In step 6, for the input remote sensing images with different resolutions, the dual-stream SAM model will first uniformly adjust the size of the remote sensing images through interpolation preprocessing. For the images in the Post-FakeV dataset, they will be unified to size; for the images in the Post-FakeL dataset, they will be unified to size; the dual-stream SAM model performs inference calculations based on the adjusted images, and after obtaining the results, the output is restored to the original size of the image through interpolation operations.

9. An electronic device, characterized in that, It includes a processor and a memory. The memory stores program code. When the program code is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 8.

10. A storage medium, characterized in that, Stores a computer program or instruction. When the computer program or instruction runs on a computer, it executes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image tampering detection and positioning method based on multi-scale supervised contrast learning

    CN117541571A

  • Optical remote sensing image change detection method and system based on change perception and semantic guidance, storage medium and electronic equipment

    CN119494830A

  • Multi-modal image-text tampering detection and positioning method based on feature enhancement

    CN119513743A

Cited By

  • PCB surface defect detection method and system based on double-layer SAM model collaboration

    CN120747085A

  • Depth image restoration tampering detection method based on adaptive tampering trace learning

    CN121032864A

  • Image forgery detection method and system based on large model, terminal and storage medium

    CN121120507A