Spectral information guided network for high resolution satellite map tampering localization

By introducing a spectral information guidance network (SIGNet) into satellite map tamper detection, and using spectral information for feature extraction and fusion, the problem of inaccurate tampering positioning in the prior art is solved, and precise positioning of tampering areas and prediction mask generation at clear edges is realized.

CN119963518AInactive Publication Date: 2025-05-09HUNAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510047705.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing satellite map tamper detection and positioning methods are difficult to accurately utilize the spectral characteristics of satellite maps, resulting in inaccurate tampering and positioning results and unclear edges of the generated prediction mask.

Method used

A spectral information guidance network (SIGNet) is proposed to extract multi-scale semantic features and spectral features through RGB branches and spectral information (SI) branches in the encoder, and feature fusion and upsampling are used in the decoder to achieve accurate positioning of the tampered area.

Benefits of technology

By utilizing spectral information, SIGNet can generate prediction masks with complete interiors and clear edges, accurately indicating the position and shape of the spliced ​​object, significantly improving the accuracy of tampering positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005238989310000011
    Figure HDA0005238989310000011
  • Figure HDA0005238989310000012
    Figure HDA0005238989310000012
  • Figure HDA0005238989310000013
    Figure HDA0005238989310000013
Patent Text Reader

Abstract

The invention relates to a technology for tampering positioning of a high-resolution satellite map, and aims to solve the problem of inaccurate positioning result caused by neglecting spectral characteristics in tampering detection and positioning of the satellite map in the prior art. The spectral information guided network (SIGNet) provided by the invention utilizes the spectral information of the satellite image to accurately position the tampered area, and particularly focuses on the tampering of the splicing type. The SIGNet network works through a double-branch encoder structure, one branch processes an original satellite image to extract semantic features, and the other branch processes spectral information obtained through preprocessing operations (including visible vegetation index calculation and Laplacian filtering) to extract spectral features. The features of the two branches are fused in a decoder through a spectrum semantic information aggregation module (SSIAB) and a fusion decoder module (FDB), finally a high-resolution prediction mask is output, and the mask can accurately indicate the position and the shape of a tampered area and has a complete internal structure and a clear edge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence and biomedical imaging technology, and specifically relates to a sparse CT reconstruction method based on complementary priors and implicit neural representation. The present invention relates to the field of satellite map tampering detection technology, and specifically relates to a spectral information guided network (SIGNet) for high-resolution satellite map tampering positioning, which uses spectral information to guide upsampling in a decoder to achieve accurate positioning of the tampered area. Background Art

[0002] With the surge in the number of commercial satellites, the availability of high-resolution satellite maps continues to increase. However, the popularity of image editing software and the maturity of image processing technology make it relatively easy to tamper with satellite maps. The main type of tampering studied in this invention is splicing, that is, covering an area in an image with a spliced ​​object in another image, thereby forging certain objects or concealing their existence. Satellite map forensics can be carried out from two dimensions: tampering detection and tampering positioning. Tampering detection is to detect the authenticity of satellite maps, that is, to detect whether the map has been forged; tampering positioning is to confirm the integrity of satellite maps, that is, to find out the area where tampering has occurred.

[0003] In order to ensure the authenticity and integrity of satellite maps, in the past few years, some researchers have proposed many methods for satellite map tampering detection and positioning. Among them, in 2023, Niloy et al. proposed a supervised method based on segmentation model HRFNet. Specifically, the HRFNet network consists of an RGB branch and an SRM branch. The former distinguishes the manipulated area from the real area by capturing the visual inconsistency in the tampering boundary, while the latter uses the SRM filter to analyze the local noise characteristics in the image. Both branches are composed of shallow and deep parts. The shallow part can globally extract features with a large receptive field and better capture spatial information, while the deep part can effectively extract high-level semantic information. Due to the complementarity of deep and shallow layers, feature fusion can improve segmentation performance. After merging the features of the RGB and SRM branches, the ASPP module is used to capture features at multiple scales to obtain richer contextual information. Finally, the decoder is entered and the final prediction mask is generated, which indicates the tampered area. Experimental results show that this method obtains a higher AUC value, and the prediction mask generated by it has clearer edges than other methods.

[0004] Although existing technologies have achieved certain results, they ignore the inherent spectral characteristics of satellite maps and are difficult to obtain accurate tampering positioning results. For satellite map tampering positioning, many current technologies can only locate the approximate position of the spliced ​​object, and the generated prediction mask does not have complete and clear edges, and cannot clearly depict the shape of the spliced ​​object. Summary of the invention

[0005] The purpose of the present invention is to provide a spectral information guided network (SIGNet) for high-resolution satellite map tampering positioning, which uses spectral information to guide upsampling in the decoder to achieve accurate positioning of the tampered area. The SIGNet proposed in the present invention uses spectral information to guide upsampling in the decoder to achieve accurate positioning of the tampered area. The prediction mask obtained using SIGNet can accurately indicate the position and shape of the spliced ​​object, with a complete interior and clear edges.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] 1. A spectral information guided network (SIGNet) such as Figure 1 As shown, it is characterized in that the SIGNet includes:

[0008] The encoder includes an RGB branch and a spectral information (SI) branch. The input of the RGB branch is the satellite map, and the input of the SI branch is the spectral information obtained by preprocessing the satellite map.

[0009] The decoder, including the spectral semantic information aggregation module (SSIAB), the fusion decoder module (FDB) and the output decoder module (ODB), is used to enhance and integrate the complementary semantic and spectral features of the same scale at each scale, restore the resolution, and output the prediction results.

[0010] 2. The SIGNet as described above is characterized in that the RGB branch and the SI branch in the encoder are both composed of three ResBlocks and a global-local Transformer module (GLTB), the RGB branch extracts multi-scale semantic features, and the SI branch extracts corresponding spectral features.

[0011] 3. The SIGNet as described above, characterized in that the SSIAB module in the decoder includes convolutions with a kernel size of 1×5 and a kernel size of 5×1 to extract features in vertical and horizontal directions, making it easier for the model to notice tampered edges.

[0012] 4. The SIGNet as described above is characterized in that the FDB module in the decoder is deployed to aggregate features of different scales layer by layer, so that the feature map retains rich semantic discriminant information while gradually restoring the resolution.

[0013] 5. The SIGNet as described above is characterized in that the preprocessing operation includes "visible vegetation index calculation + Laplace filtering" for extracting spectral information and capturing possible tampering edges in satellite maps. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is an overview diagram of the SIGNet architecture of the present invention;

[0015] Figure 2 It is a structural diagram of the encoder GLTB of the present invention;

[0016] Figure 3 is a block diagram of the SSIAB of the present invention;

[0017] Figure 4 It is a structural diagram of the FDB of the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose and advantages of the present invention more clearly understood, the present invention is specifically described below in conjunction with embodiments. It should be understood that the following text is only used to describe one or several specific embodiments of the present invention, and does not strictly limit the scope of protection of the specific claims of the present invention.

[0019] Example:

[0020] 1. Preprocessing operations

[0021] In a specific embodiment of the present invention, the satellite map is first preprocessed to extract spectral information. This step is based on the widespread proximity effect in satellite images, and is achieved by calculating the visible vegetation index and applying Laplace filtering. Specifically, we use two vegetation indices: RGBVI (Red-Green-Blue Ratio Vegetation Index) and MGRVI (Modified Green-Red Vegetation Index). By performing Laplace filtering on these indices, we can effectively capture the possible tampering edges in the satellite map and provide key spectral information for subsequent tampering positioning.

[0022] 2. Spectral Information Guided Network (SIGNet)

[0023] The core of this invention is SIGNet, a network with a U-shaped encoder-decoder structure. Figure 1 As shown, it includes an RGB branch for extracting semantic features and an SI branch for extracting spectral features. In the encoder stage, the RGB branch inputs the original satellite map, while the SI branch inputs the preprocessed spectral information. Both branches consist of three ResBlocks and a global-local Transformer module (GLTB), which extract multi-scale semantic features and spectral features respectively.

[0024] The GLTB module in the encoder is inspired by UNetFormer, and GLTB is embedded into the last layer of the encoder's dual branches, such as Figure 2As shown. The output of the third ResBlock of the RGB or SI branch, F3 or S3, is set as input, followed by batch normalization and efficient global-local attention to extract local context and global context information, respectively. Subsequently, the global-local context information and the input features are added, and then BN and a multi-layer perceptron (MLP) with 2 linear layers are performed. Finally, the execution result and the addition result are also added as the output result F4 or S4. In order to avoid information loss, jump connections are performed between the input and parallel processing and between the parallel processing and the output. Efficient global-local attention uses two parallel branches to extract local and global context. For the global branch, the input is first divided into multiple windows, and the information modeling within the window uses a multi-head self-attention mechanism. The horizontal average pooling layer and the vertical average pooling layer are used between the windows to complete the information interaction. For the local branch, two parallel convolution operations with different convolution kernel sizes and batch normalization are used to obtain the local context. In this way, by utilizing local and global information in the network, we can more accurately locate the tampered area without internal holes caused by insufficient receptive field.

[0025] In the decoder stage, we construct a spectral semantic information aggregation module (SSIAB) to enhance and integrate the complementary semantic and spectral features at each scale. Figure 3 As shown in Figure 1, in order to better fuse spectral features and semantic features, we first enhance the features of Fi and Si, that is, integrate the semantic features into the spectral features, integrate the spectral features into the semantic features, and then perform weighted summation to obtain the final fusion feature SFi. The structure of SSIAB is shown in Figure 1. Figure 3 As shown. Convolutions with kernel sizes of 1×5 and 5×1 are designed to extract features in the vertical and horizontal directions, making it easier for the model to notice tampering edges. After concatenating the directional features from Fi and Si, the spatial attention weight SAW is generated using the spatial attention module SAM. Feature enhancement is completed by multiplying Fi and Si by SAW respectively. Finally, according to the contribution of Fi and Si to the network, they are summed to obtain the aggregate feature SFi.

[0026] Subsequently, the fusion decoder module (FDB) uses the channel attention weights to guide the fusion of high-level and low-level features layer by layer to retain more important information. The structure of FDB is shown in Figure 4As shown. First, FFi+1 is upsampled by 2 times to ensure that it has the same resolution as SFi, and then they are cascaded together. Then, the channel attention weights are obtained through global pooling and 1×1 convolution and Sigmoid function, and then multiplied and added with the concatenated features. This process can effectively retain important features while suppressing irrelevant features. Then, 1×1 convolution, BN and ReLU are combined to reduce the number of channels, integrate information from different features to enhance the capabilities of the model, and finally a combination of 3×3 convolution, BN and ReLU is used to output the aggregated feature FFi. Finally, the output decoder module (ODB) restores the resolution and outputs the prediction results.

[0027] 3. Dataset Construction and Evaluation

[0028] Satellite maps of Daxing District, Beijing in 2023 were collected from Amap to create a satellite map tampering dataset (SMTD). The height and width of each satellite map are 1024 pixels, and the spatial resolution is 0.597164 meters / pixel. The dataset contains a total of 1000 satellite maps, of which 750 contain spliced ​​objects from the xView2 / XBD dataset provided by Google Earth. The process of making a fake map is as follows: First, we select some representative pre-disaster satellite maps from the xView2 / XBD dataset and use the labeling tool labelme to mark some objects as splicing objects, such as buildings, roads, clouds, bushes, etc. Then, a complete satellite map without blanks is selected from the satellite map of Daxing District, Beijing as the original satellite map. Finally, some pixels of the original satellite map are covered with specific splicing objects to create a fake satellite map. We divide the SMTD dataset into training set, validation set, and test set in a ratio of 3:1:1. That is, the training set contains 600 satellite maps and their corresponding real masks, of which 450 are forged images; the validation set contains 200 satellite maps and their corresponding real masks, of which 150 are forged images; the test set contains 200 satellite maps and their corresponding real masks, of which 150 are forged images. We use IOU, F1-score, ROC_AUC and P / R_AUC as evaluation indicators. The experimental results show that SIGNet has achieved excellent performance in both forged image detection and locating tampered areas in forged images.

[0029] The above is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention are implemented according to conventional means in the art unless otherwise specified and limited.

Claims

1. A Spectral Information Guided Network (SIGNet) for high-resolution satellite map tampering localization, characterized in that: The network includes: S1: a dual-branch encoder, including an RGB branch for extracting semantic features and an SI branch for extracting spectral features; S2: A decoder, including four spectral semantic information aggregation modules (SSIAB) for enhancing and fusing spectral and semantic features, three fusion decoder modules (FDB) for fusing features at different levels, and an output decoder module (ODB) for restoring the resolution.

2. The network according to claim 1, characterized in that The RGB branch in the encoder inputs the satellite map, and the SI branch inputs the spectral information obtained by preprocessing the satellite map. Both consist of three ResBlocks (layer1-layer3 of ResNet18) and a Global-Local Transformer Block (GLTB).

3. The GLTB constituting the encoder according to claim 2, characterized in that The output of the third ResBlock of the RGB or SI branch is set as the input, followed by batch normalization and efficient global-local attention to extract local context and global context information respectively. Subsequently, the global-local context information and the input features are added, and then BN (batch normalization) and a multi-layer perceptron (MLP) with 2 linear layers are performed. Finally, the execution result and the addition result are also added as the output result. In order to avoid information loss, jump connections are performed between the input and parallel processing and between the parallel processing and the output. Efficient global-local attention uses two parallel branches to extract local and global context. For the global branch, the input is first divided into multiple windows, and the information modeling within the window uses a multi-head self-attention mechanism. The horizontal average pooling layer and the vertical average pooling layer are used between the windows to complete the information interaction. For the local branch, two parallel convolution operations with different convolution kernel sizes and batch normalization are used to obtain the local context. In this way, by utilizing local and global information in the network, the tampered area can be more accurately located without internal holes caused by insufficient receptive field.

4. The network according to claim 1, characterized in that Preprocessing operations include calculating visible vegetation indices and Laplace filtering to extract spectral information that captures possible tampering edges in satellite maps. The visible vegetation indices used include RGBVI (Red-Green-Blue Ratio Vegetation Index) and MGRV (Modified Green-Red Vegetation Index).

5. The network according to claim 1, characterized in that The SSIAB module in the decoder includes: Convolutions with kernel size of 1×5 and 5×1 are designed to extract features in vertical and horizontal directions, making it easier for the model to notice tampered edges; The Spatial Attention Module (SAM) generates spatial attention weights for feature enhancement.

6. Perform weighted summation based on the contribution of semantic features and spectral features to obtain aggregated features. The network according to claim 1, characterized in that The FDB module in the decoder is used to aggregate features of different scales layer by layer to retain rich semantic discriminant information. The specific process is as follows: First, FFi+1 is upsampled by 2 times to ensure that it has the same resolution as SFi, and then they are cascaded together. Then, the channel attention weights are obtained through global pooling and 1×1 convolution and Sigmoid function, and then multiplied and added with the concatenated features. This process can effectively retain important features while suppressing irrelevant features. Then, 1×1 convolution, BN and ReLU are combined to reduce the number of channels, integrate information from different features to enhance the capabilities of the model, and finally a combination of 3×3 convolution, BN and IReLU is used to output the aggregated feature FEi.

7. A method for satellite map tampering positioning using the spectral information guidance network as claimed in any one of claims 1 to 5, characterized in that: The method comprises the following steps: Extract spectral information from satellite maps using preprocessing operations; The semantic and spectral features of satellite maps are extracted through a dual-branch encoder; The semantic features and spectral features are fused and enhanced through the decoder; Outputs a predicted mask indicating the location and shape of the stitched object.