Remote sensing image tampering target detection method and system based on spatiotemporal evolution

By combining high spatiotemporal resolution generation and spatiotemporal inconsistency detection modules, the problem of insufficient spatiotemporal correlation in remote sensing image tampering target detection is solved, efficient and accurate tampering target detection is achieved, the manual identification process is simplified, and the detection speed and accuracy are improved.

CN115457362BActive Publication Date: 2025-09-19SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211149714.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-09-19
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

Existing remote sensing image tampering target detection methods lack the temporal and spatial correlation analysis between time-series remote sensing images, and the insufficient temporal and spatial resolution leads to deviations and missed judgments when the objects evolve. There is also a lack of efficient means of fusion of temporal and spatial dimension features, which affects the detection accuracy.

Method used

A high spatiotemporal resolution generation module, a spatiotemporal inconsistency detection module and a result output module are adopted. By fusing high-temporal-low-space and low-temporal-high-space features and combining spatial and temporal inconsistency analysis, the inconsistency of forged areas is captured, information complementarity is achieved, and the diversity and completeness of the spatiotemporal correlation features of the model are improved.

Benefits of technology

Accurately locate the tampered target area in remote sensing images, simplify the manual identification process, speed up the detection speed and precision, improve the detection accuracy and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457362B_ABST
    Figure CN115457362B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting tampered targets in remote sensing images based on multi-view features and multi-scale supervision. The system comprises: a boundary detection module that detects boundary artifacts resulting from camouflaged and masked tampered targets; a noise detection module that captures the difference in noise distribution between the tampered area and the real area of ​​the remote sensing image; a dual-attention feature fusion module that fuses channel features using a channel attention fusion network; and a multi-scale supervised loss model that combines three scales of loss to improve model generalization performance. The method and system for detecting tampered targets in remote sensing images based on multi-view features and multi-scale supervision, provided by the present invention, constructs an automatic detection model for tampered target anomalies in remote sensing images, which can accurately detect and locate tampered image areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a remote sensing image tampering target detection method and system based on spatiotemporal evolution, and belongs to the fields of computer and remote sensing. Background Art

[0002] Remote sensing image tampering detection technology is a critical intelligent technology that quickly and accurately identifies tampered targets in remote sensing images and locates the tampered targets. Tampered targets in remote sensing images are typically masked using various techniques, such as splicing, duplication, and deletion, to reduce or eliminate the visual distinction between sensitive targets and the surrounding background, effectively concealing the real and revealing the fake.

[0003] Traditional machine learning methods for detecting tampered objects in remote sensing imagery offer strong targeting and flexible design, but they suffer from rigid modeling and poor robustness. In recent years, deep learning technology has achieved remarkable results in remote sensing image processing. Numerous object detection algorithms based on convolutional neural networks (CNNs), a deep learning technique, have been proposed and applied to remote sensing imagery. Compared to traditional algorithms, convolutional neural networks can learn higher-level semantic information from remote sensing imagery, resulting in greater robustness and improved performance in detecting tampered objects in remote sensing imagery.

[0004] In recent years, many works have explored the task of detecting tampered objects in remote sensing images, and proposed many new ideas based on deep learning object detection networks to improve them. However, the existing methods still have the following common problems: (1) they do not effectively utilize the spatiotemporal correlation between time-series remote sensing images; (2) the spatiotemporal resolution of time-series remote sensing images is insufficient, which leads to bias and missed detection when analyzing the evolution of land objects; (3) there is a lack of efficient fusion methods for spatiotemporal dimension features to improve the diversity and completeness of spatiotemporal correlation features. Summary of the Invention

[0005] The purpose of this invention is to improve the visual feature expression of time-series remote sensing images, analyze the evolution of land objects through the spatiotemporal correlation between time-series remote sensing images, mine deep spatiotemporal correlation features through spatiotemporal feature fusion, improve the diversity and completeness of the model for mining the spatiotemporal correlation features of multi-phase remote sensing images, and thus improve the accuracy of model detection.

[0006] In order to achieve the above-mentioned object, a technical solution of the present invention is to provide a remote sensing image tampering target detection system based on spatiotemporal evolution, characterized in that it includes a high spatiotemporal resolution generation module, a spatiotemporal inconsistency detection module and a result output module;

[0007] The high spatiotemporal resolution generation module further includes:

[0008] The high-temporal-low-spatial unit is used to extract low-frequency features from the high-temporal-low-spatial-resolution remote sensing image at the specified time to be generated and the high-temporal-low-spatial-resolution remote sensing image at nearby times near the specified time to be generated, thereby obtaining a final high-temporal-low-spatial-feature map.

[0009] The high-spatial and low-temporal unit is used to extract high-frequency features from the low-temporal and high-spatial resolution remote sensing images at the corresponding moments, and filter out noise information to obtain the final low-temporal and high-spatial feature map;

[0010] A feature fusion unit is used to fuse a high-temporal and low-spatial feature map with a low-temporal and high-spatial feature map of the same dimension and size to generate a high-temporal and spatial resolution remote sensing image at a specified moment;

[0011] The spatiotemporal inconsistency detection module further includes a spatial inconsistency unit, a temporal inconsistency unit, and a spatiotemporal feature fusion unit, wherein:

[0012] The spatial inconsistency unit and the temporal inconsistency unit detect the inconsistency of the forged area from the temporal and spatial perspectives based on high spatiotemporal resolution remote sensing images, and obtain spatial inconsistency feature maps and temporal inconsistency feature maps respectively.

[0013] The spatiotemporal feature fusion unit is used to capture and fuse the inconsistent information of spatially inconsistent feature maps and temporally inconsistent feature maps, achieve information complementarity, obtain the final feature representation, and realize auxiliary detection of forged targets;

[0014] The result output module identifies and outputs the coordinates of the tampered area and the tampering method of the remote sensing image based on the final feature representation output by the spatiotemporal inconsistency detection module.

[0015] Another technical solution of the present invention is to provide a remote sensing image tampering target detection method based on spatiotemporal evolution implemented based on the above remote sensing image tampering target detection system, which is characterized by comprising the following steps:

[0016] Step 1: Build a detection model based on the high spatiotemporal resolution generation module, spatiotemporal inconsistency detection module and result output module.

[0017] Step 2: Train the detection model, including the following steps:

[0018] Step 201: Collect sample data and construct a training data set;

[0019] Step 202: Train the detection model. The implementation of the detection model includes the following steps:

[0020] Step 2021: Generate the specified time to be generated using the high spatiotemporal resolution generation module The high temporal and spatial resolution images further include the following steps:

[0021] Step 20211: The high time and low space units extract the specified time to be generated respectively The low-frequency features in the high-temporal and low-spatial resolution remote sensing images and the features at the specified time to be generated The low-frequency features in the high-temporal and low-spatial resolution remote sensing images at the nearby time t are obtained to obtain the low-frequency feature map, and the size of the low-frequency feature map is expanded to obtain the specified time to be generated. and the final high temporal and low spatial feature map at nearby time t and

[0022] Step 20212: The low-time high-space unit extracts high-frequency features from the low-time, high-spatial resolution remote sensing image at the nearby time t, and filters the noise information to obtain the final low-time high-spatial feature map.

[0023] Step 20213, the feature fusion unit combines the low temporal and high spatial feature maps at the nearby time t Subtract the high temporal and low spatial feature maps at nearby time t As a difference reference, then add the difference to the specified time to be generated High temporal and low spatial feature maps Thus, the high and low spatial frequency information of the two feature maps are fused, and the final result is used to generate the specified time to be generated. Fusion characteristics of remote sensing images It is expressed as the following formula:

[0024]

[0025] Step 2022: Detection of spatiotemporal inconsistency:

[0026] The inconsistency of the forged area is detected from the perspectives of time and space by using the spatial inconsistency unit and the temporal inconsistency unit. Specifically, the following steps are included:

[0027] Step 20221: The spatial inconsistency unit has a three-channel structure. For the input feature map of the spatial inconsistency unit Where T represents the time T, C represents the number of remote sensing image channels, H represents the length of the remote sensing image, and W represents the width of the remote sensing image. The middle channel in the three-channel structure works in a low-resolution manner. The input feature map X1 is average pooled with a kernel size of 2×2 and a stride of 2, and then undergoes two consecutive convolution and bilinear upsampling operations to obtain the output S:

[0028] S=upsampling(K1*K2*(AvgPool2(X1)))

[0029] Where upsampling represents bilinear upsampling; K1 and K2 are convolutions with kernel sizes of 1×3 and 3×1 respectively; AvgPool2(X1) represents average pooling of the input feature map X1;

[0030] The upper path in the three-path structure is a residual connection added to the middle path;

[0031] The output of the upper path and the output of the middle path are fused, and the fusion result is passed through sigmoid to obtain the confidence, which is multiplied with the output of the lower path in the three-path structure after 3x3 convolution to obtain the final distribution confidence score feature map Y1:

[0032] Y1=K4⊙(σ(S+X1)⊙K3(X1))

[0033] Where K3 is the 3×3 convolution for feature extraction, K4 is the 3×3 convolution for post-processing, and σ is the sigmoid function;

[0034] Step 20222: The temporal inconsistency unit observes and models the remote sensing image sequence from the horizontal and vertical directions, respectively, and uses the feature differences of the remote sensing images at adjacent moments along these two orthogonal directions to discover temporally inconsistent regions, including the following steps:

[0035] 1) Input feature map for time-inconsistent units Input in the horizontal and vertical directions respectively, and then undergo convolution, difference and sigmoid operations to obtain two importance weights F with the same shape as the input feature map x2 h and F w , including the following steps

[0036] a) After compressing the input feature map X2 by r times in the channel dimension, perform differential calculations along the vertical and horizontal directions to obtain the vertical slice difference map S h and horizontal slice difference map S w ;

[0037] b) The vertical slice difference map S is enhanced by the vertical time inconsistency enhancement unit VTIE and the horizontal time inconsistency enhancement unit HTIE respectively. h and horizontal slice difference map S w Process and extract vertical slice difference map S h and horizontal slice difference map S w Multi-level representation of

[0038] c) Use element-by-element addition to fuse the multi-level representations extracted by the vertical time inconsistency enhancement unit VTIE and the horizontal time inconsistency enhancement unit HTIE, and then use the sigmoid function to determine the importance weight F h and F w ;

[0039] 2) The importance weight F h and F w Multiply by X2 to obtain the enhanced vertical inconsistency feature map Y2 along the time dimension, as shown in the following formula:

[0040]

[0041] Where, VTIE(S h ) represents the vertical time inconsistency enhancement unit VTIE for the vertical slice difference map S h Operation, HTIE(S w ) represents the horizontal slice difference map S of the horizontal time inconsistency enhancement unit HTIE w Operation;

[0042] Step 20223: The spatiotemporal feature fusion unit fuses the inconsistent information captured by the spatial inconsistency unit and the temporal inconsistency unit, effectively utilizing the complementary effect of the two types of information, including the following steps:

[0043] The spatiotemporal feature fusion unit selects useful channels to supplement the feature map Y1 output by the spatial inconsistency unit through a combination of global average pooling and one-dimensional convolution, as shown in the following formula:

[0044]

[0045] Where GAP represents global average pooling, K5 is a one-dimensional convolution with a kernel size of 3, Represents the feature map output by Y1 after global average pooling and one-dimensional convolution.

[0046] The spatiotemporal feature fusion unit will The final feature representation Y3 is obtained by fusing it with the output of the time-inconsistent unit, as shown in the following formula:

[0047]

[0048] Where K6 is a 3×3 convolution;

[0049] Step 203: The result output module finally identifies and outputs the coordinates of the tampered area and the tampering method of the remote sensing image based on the final feature representation Y3;

[0050] Step 3: Input the remote sensing image data uploaded by the client in real time into the trained detection model. The detection model outputs the detection results of the coordinates of the tampered area and the tampering method of the remote sensing image on the server.

[0051] Preferably, in step 201, a large-scale remote sensing image database is constructed through the remote sensing data service provided by the satellite, sample data is collected from the large-scale remote sensing image database, and a training data set for detection model training is constructed. Each sample data in the large-scale remote sensing image database contains the marked tampered area location and tampering method, and finally a training set, a test set and a validation set are obtained based on the large-scale remote sensing image database.

[0052] Preferably, the vertical time inconsistency enhancement unit VTIE converts S h As input, it extracts multi-level representations through three branches - 3×1 convolution along the horizontal dimension, average pooling operation, 3×1 convolution along the horizontal dimension and upsampling pooling operation, skip connection;

[0053] The horizontal time inconsistency enhancement unit HTIE will S w As input, a multi-level representation is extracted through three branches - 3×1 convolution along the vertical dimension, average pooling operation, 3×1 convolution along the vertical dimension and upsampling pooling operation, skip connection.

[0054] Preferably, in step 3, the client and the server are implemented based on the C / S architecture, the server isolates the network through a firewall, transmits image data through an encrypted communication protocol to prevent illegal client access, and performs large-scale remote sensing image data storage and GPU calculation through distributed nodes; for the result output module, the client is connected to a display and a printer, and the result output module signal is set to be connected to the display and the printer to realize screen display and document printing of the diagnostic report.

[0055] The present invention has a reasonable structural design. It analyzes the evolution of land objects through the spatiotemporal correlation between time-series remote sensing images, mines deep spatiotemporal correlation features through spatiotemporal feature fusion, and improves the diversity and completeness of the model in mining the spatiotemporal correlation features of multi-phase remote sensing images, thereby being able to accurately locate the tampered target area in the remote sensing image, greatly simplifying the manual recognition process, and accelerating the speed and accuracy of detection and recognition, thereby being able to accurately locate the tampered target area in the remote sensing image.

[0056] Compared with the existing technical solutions, the present invention has the following effects:

[0057] (1) The present invention proposes a remote sensing image tampering target detection method and system based on spatiotemporal evolution, which improves the visual feature expression of time-series remote sensing images, analyzes the evolution of ground objects through the spatiotemporal correlation between time-series remote sensing images, and mines deep spatiotemporal correlation features through spatiotemporal feature fusion, thereby improving the diversity and completeness of the model in mining the spatiotemporal correlation features of multi-phase remote sensing images, thereby improving the detection accuracy of the system.

[0058] (2) The present invention proposes a high temporal and spatial resolution remote sensing image generation technology. By fusing the remote sensing image features of low temporal and high spatial resolution with the remote sensing image features of high temporal and low spatial resolution in the same region and time period, remote sensing images with high temporal and low spatial resolution are obtained, thereby avoiding the deviation and omission in the analysis of the evolution of land objects caused by insufficient resolution of time series remote sensing images.

[0059] (3) The temporal inconsistency analysis and spatial inconsistency analysis methods proposed in the present invention can analyze the evolution of land objects through the temporal and spatial correlation between time-series remote sensing images, thereby detecting traces of camouflage and masking processing that are different from the normal evolution process of land objects and locating the corresponding hidden target area.

[0060] (4) The present invention combines the fusion and enhancement of spatial and temporal features through the information supplementation module to obtain a more comprehensive representation of hidden target camouflage and masking traces, which can not only accurately locate the hidden target but also ensure the strong generalization ability of the model.

[0061] (5) The present invention proposes a remote sensing image tampering target detection method and system based on spatiotemporal evolution, which uses a sample database for supervised learning, inputs remote sensing images with multiple temporal resolutions and spatial resolutions into the system, generates high temporal resolution and high spatial resolution images, and finally detects the tampering targets and tampering methods in the remote sensing images through a spatiotemporal correlation evolution analysis model, which greatly simplifies the manual recognition process, speeds up the speed and accuracy of detection and recognition, and reduces the detection cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is the overall framework diagram of the present invention;

[0063] Figure 2 This is a diagram showing a framework for implementing the detection model involved in the present invention;

[0064] Figure 3 Generate a network structure diagram with high spatiotemporal resolution according to the present invention;

[0065] Figure 4 This is a structural diagram of a spatially inconsistent module according to the present invention;

[0066] Figure 5 This is a structural diagram of the time inconsistency module involved in the present invention;

[0067] Figure 6 This is a structural diagram of the information supplement module involved in the present invention;

[0068] Figure 7 This is a diagram of the overall network architecture of the client / server system involved in the present invention. DETAILED DESCRIPTION

[0069] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0070] like Figure 1 As shown, this embodiment proposes a remote sensing image tampering target detection system based on spatiotemporal evolution, including a high spatiotemporal resolution generation module, a spatiotemporal inconsistency detection module, and a result output module.

[0071] The high spatiotemporal resolution generation module further includes:

[0072] The high-temporal-low-spatial unit is used to extract low-frequency features from the high-temporal-low-spatial-resolution remote sensing image at the specified time to be generated and the high-temporal-low-spatial-resolution remote sensing image at nearby times near the specified time to be generated, thereby obtaining a final high-temporal-low-spatial-feature map.

[0073] The high-spatial and low-temporal unit is used to extract high-frequency features from the low-temporal and high-spatial resolution remote sensing images at the corresponding moments, and filter out noise information to obtain the final low-temporal and high-spatial feature map;

[0074] The feature fusion unit is used to fuse the high-temporal and low-spatial feature map with the low-temporal and high-spatial feature map of the same dimension and size to generate a high-temporal and spatial resolution remote sensing image at a specified moment.

[0075] The spatiotemporal inconsistency detection module further includes a spatial inconsistency unit, a temporal inconsistency unit, and a spatiotemporal feature fusion unit, wherein:

[0076] The spatial inconsistency unit and the temporal inconsistency unit detect the inconsistency of the forged area from the temporal and spatial perspectives based on high spatiotemporal resolution remote sensing images, and obtain spatial inconsistency feature maps and temporal inconsistency feature maps respectively.

[0077] The spatiotemporal feature fusion unit is used to capture and fuse the inconsistent information of spatially inconsistent feature maps and temporally inconsistent feature maps, achieve information complementarity, obtain the final feature representation, and realize auxiliary detection of forged targets.

[0078] Considering that in the process of camouflaging and masking hidden targets, the camouflaged and masked remote sensing images generated are often accompanied by subtle artificial traces of the hidden target area, such as mixed boundaries or image quality mismatch during image stitching, the spatial inconsistency of these artificial traces between the original area and the forged area of ​​each image is captured by the spatial inconsistency unit.

[0079] Considering that although spatial inconsistency plays a key role in identifying hidden targets through camouflage and masking, high-precision camouflage and masking methods may produce extremely realistic forged remote sensing images, a temporal inconsistency unit is used to capture the temporal inconsistency of remote sensing images after camouflage and masking in the time series, thereby realizing auxiliary detection of hidden targets from the temporal dimension.

[0080] The spatiotemporal feature fusion unit fuses the inconsistent information captured by the spatial inconsistency unit and the temporal inconsistency unit respectively, effectively utilizing the complementary effect of the two types of information, thereby improving the diversity and completeness of the model in mining the spatiotemporal correlation features of multi-temporal remote sensing images.

[0081] The result output module uses the final feature representation output by the spatiotemporal inconsistency detection module to identify and output the coordinates of the tampered areas in the remote sensing image and the tampering methods, including splicing, copying, and cropping. The result output module also uses red boxes to mark the specific locations of the tampered areas, facilitating further analysis and processing.

[0082] This embodiment further discloses a remote sensing image tampering target detection method based on spatiotemporal evolution implemented based on the remote sensing image tampering target detection system, comprising the following steps:

[0083] Step 1: Build a detection model based on the high spatiotemporal resolution generation module, spatiotemporal inconsistency detection module and result output module.

[0084] Step 2: Train the detection model, including the following steps:

[0085] Step 201: Collect sample data and build a training data set:

[0086] A large-scale remote sensing image database is built through the remote sensing data services provided by the Gaofen, Sentinel and other series of satellites. Sample data is collected from the large-scale remote sensing image database to construct a training data set for detection model training. Each sample data in the large-scale remote sensing image database contains the marked location of the tampered area and the tampering method. Finally, the training set, test set and validation set are obtained based on the large-scale remote sensing image database.

[0087] Step 202: Train the detection model. The implementation of the detection model includes the following steps:

[0088] Step 2021: Generate the specified time to be generated using the high spatiotemporal resolution generation module The high temporal and spatial resolution images further include the following steps:

[0089] Step 20211: The high time and low space units extract the specified time to be generated respectively The low-frequency features in the high-temporal and low-spatial resolution remote sensing images and the features at the specified time to be generated The low-frequency features in the high-temporal and low-spatial resolution remote sensing images at the nearby time t are obtained to obtain the low-frequency feature map, and the size of the low-frequency feature map is expanded to obtain the specified time to be generated. and the final high temporal and low spatial feature map at nearby time t and

[0090] like Figure 3 As shown in the figure, the high-temporal-low-spatial unit mainly consists of two convolutional layers and three deconvolutional layers. Each convolutional layer and deconvolutional layer uses the ReLU activation function. The deconvolution layer expands the input feature map to the same size as the output feature map of the low-temporal-high-spatial unit, so that the two data sources have the same dimension and size when the features are fused.

[0091] Step 20212: The low-time high-space unit extracts high-frequency features from the low-time, high-spatial resolution remote sensing image at the nearby time t, and filters the noise information to obtain the final low-time high-spatial feature map.

[0092] The low-temporal and high-spatial unit adopts a convolutional network. The input feature map passes through two convolution layers to extract the edge features and texture features of the image content in the low-temporal and high-spatial resolution remote sensing image. The maximum pooling layer is then used to reduce the dimension and filter the high-frequency noise information. Finally, two convolution layers are used to adjust the dimension and size of the feature map.

[0093] Step 20213: Each pixel in the high temporal and low spatial resolution remote sensing image is matched to the pixel of an area in the low temporal and high spatial resolution remote sensing image. The value of each pixel in the high temporal and low spatial resolution remote sensing image can be obtained by weighting the pixel of the corresponding area in the low temporal and high spatial resolution remote sensing image. Therefore, the feature map of the same dimension and size can be obtained by using the high temporal and low spatial unit and the low temporal and high spatial unit. The present invention proposes to specify the time to be generated. The difference between the high temporal and low spatial feature maps at the nearby time t is equal to the difference between the corresponding low temporal and high spatial feature maps, that is:

[0094]

[0095] Therefore, the feature fusion unit combines the low temporal and high spatial feature maps at the nearby time t Subtract the high temporal and low spatial feature maps at nearby time t As a difference reference, then add the difference to the specified time to be generated High temporal and low spatial feature maps Thus, the high and low spatial frequency information of the two feature maps are fused, and the final result is used to generate the specified time to be generated. Fusion characteristics of remote sensing images It is expressed as the following formula:

[0096]

[0097] Step 2022: Detection of spatiotemporal inconsistency:

[0098] The inconsistency of the forged area is detected from the perspectives of time and space through the spatial inconsistency unit and the temporal inconsistency unit. Figure 3 , specifically including the following steps:

[0099] Step 20221: Considering that during the camouflage and masking process of a hidden target, the generated camouflaged and masked remote sensing images are often accompanied by subtle artifacts of the hidden target area, such as mixed boundaries or image quality mismatch during image splicing. The present invention captures the spatial inconsistency of these artifacts between the original area and the forged area of ​​each image through a spatial inconsistency unit, such as Figure 4 shown.

[0100] The spatially inconsistent unit consists of a series of two-dimensional operations at the image level, without considering temporal information. In addition, in order to improve the performance of the model, the present invention proposes a three-channel structure of the spatially inconsistent unit. Where T represents the time T, C represents the number of remote sensing image channels, H represents the length of the remote sensing image, and W represents the width of the remote sensing image. The middle channel in the three-channel structure works in a low-resolution manner. The input feature map X1 is average pooled with a kernel size of 2×2 and a stride of 2, and then undergoes two consecutive convolution and bilinear upsampling operations to obtain the output S:

[0101] S=upsampling(K1*K2*(AvgPool2(X1)))

[0102] In the formula, upsampling represents bilinear upsampling; K1 and K2 are convolutions with kernel sizes of 1×3 and 3×1, respectively; and AvgPool2(X1) represents average pooling of the input feature map X1. The shape of the output S obtained by the middle path is the same as that of the input feature map X1. As a result, the middle path achieves a larger receptive field and achieves better performance when combined with other paths with normal receptive fields.

[0103] The upper path in the three-path structure is a residual connection, which is added to the middle path to avoid information loss caused by downsampling. Finally, the output of the upper path is fused with the output of the middle path. The fused result is passed through a sigmoid to obtain the confidence score, which is multiplied with the output of the lower path in the three-path structure after 3x3 convolution to obtain the final distribution confidence score feature map Y1:

[0104] Y1=K4⊙(σ(S+X1)⊙K3(X1))

[0105] Where K3 is the 3×3 convolution for feature extraction, K4 is the 3×3 convolution for post-processing, and σ is the sigmoid function.

[0106] Step 20222: Considering that spatial inconsistency plays a key role in identifying hidden targets through camouflage and masking, high-precision camouflage and masking methods may produce extremely realistic fake remote sensing images. The present invention uses a temporal inconsistency unit to capture the temporal inconsistency of remote sensing images that have been camouflaged and masked in the time series, thereby achieving auxiliary detection of hidden targets from a new perspective. Figure 5 shown.

[0107] The temporal inconsistency unit observes and models the remote sensing image sequence from the horizontal and vertical directions respectively, and uses the feature differences of the remote sensing images at adjacent moments along these two orthogonal directions to find the temporally inconsistent areas. Input in the horizontal and vertical directions respectively, and then undergo convolution, difference and sigmoid operations to obtain two importance weights F with the same shape as the input feature map X2 h and F w .

[0108] For simplicity, to generate F h Taking the vertical path as an example, the input feature map X2 is compressed r times in the channel dimension to obtain:

[0109]

[0110] Where, Indicates that the number of channels, length, and width of the feature map at the jth moment in the vertical direction satisfy H and W.

[0111] The difference along the vertical direction is calculated as:

[0112]

[0113] Where Convl(·) is a 3×1 two-dimensional convolution. Represents the feature difference map in the vertical direction at time t. In addition, since there is no more time difference information at time T, Set to zero, so that the vertical slice difference map is The corresponding horizontal slice difference map is

[0114] The present invention further proposes a multi-view structure, which mainly includes a vertical slice difference map S h Vertical temporal inconsistency enhancement unit VTIE and horizontal slice difference map S w The horizontal time inconsistency enhancement unit HTIE is also used for the sake of simplicity. h As an example, the vertical time inconsistency enhancement unit VTIE will S h As input, a multi-level representation is extracted through three branches - 3×1 convolution along the horizontal dimension, average pooling operation, 3×1 convolution along the horizontal dimension and upsampling pooling operation, skip connection.

[0115] Then, element-by-element addition is used to fuse the features extracted by the vertical time inconsistency enhancement unit VTIE and the horizontal time inconsistency enhancement unit HTIE, and the importance weight F is determined by the sigmoid function. h and F w After that, multiply it by X2 to obtain the enhanced vertical inconsistency feature map Y2 along the time dimension, as shown in the following formula

[0116]

[0117] Where, VTIE(S h ) represents the vertical time inconsistency enhancement unit VTIE for the vertical slice difference map S h Operation, HTIE(S w ) represents the horizontal slice difference map S of the horizontal time inconsistency enhancement unit HTIE w operation.

[0118] Step 20223: The spatiotemporal feature fusion unit fuses the inconsistent information captured by the spatial inconsistency unit and the temporal inconsistency unit, thereby efficiently utilizing the complementary effect of the two types of information.

[0119] Specifically, such as Figure 6 As shown in Figure 1, the spatiotemporal feature fusion unit selects useful channels to supplement the feature map Y1 output by the spatial inconsistency unit through a combination of global average pooling and one-dimensional convolution. Global average pooling is used to obtain a global representation, and one-dimensional convolution on the channel dimension is used to capture cross-channel dependencies, thereby further determining the importance of each channel, as shown in the following formula:

[0120]

[0121] Where GAP represents global average pooling, K5 is a one-dimensional convolution with a kernel size of 3, Represents the feature map output by Y1 after global average pooling and one-dimensional convolution.

[0122] The spatiotemporal feature fusion unit will The final feature representation Y3 is obtained by fusing it with the output of the time-inconsistent unit, as shown in the following formula:

[0123]

[0124] Where K6 is a 3×3 convolution.

[0125] Step 203: The result output module finally identifies and outputs the coordinates of the tampered area and the tampering method of the remote sensing image based on the final feature representation Y3.

[0126] Step 3: Input the remote sensing image data uploaded by the client in real time into the trained detection model. The detection model outputs the detection results of the coordinates of the tampered area and the tampering method of the remote sensing image on the server.

[0127] In this embodiment, the system is implemented as a client-server architecture. The server isolates the network through a firewall and transmits image data via an encrypted communication protocol to prevent unauthorized client access. Large-scale remote sensing image data storage and GPU computing are performed through distributed nodes. For the result output module, each client is connected to a display and printer. By connecting the result output module to a display and printer, the diagnostic report can be displayed on the screen and printed as a document, facilitating subsequent review and verification by image forensics analysts.

Claims

1. A remote sensing image tampering target detection system based on spatiotemporal evolution, characterized by: It includes a high spatiotemporal resolution generation module, a spatiotemporal inconsistency detection module, and a result output module; The high spatiotemporal resolution generation module further includes: The high-temporal-low-spatial unit is used to extract low-frequency features from the high-temporal-low-spatial-resolution remote sensing image at the specified time to be generated and the high-temporal-low-spatial-resolution remote sensing image at nearby times near the specified time to be generated, thereby obtaining a final high-temporal-low-spatial-feature map. The high-spatial and low-temporal unit is used to extract high-frequency features from the low-temporal and high-spatial resolution remote sensing images at the corresponding moments, and filter out noise information to obtain the final low-temporal and high-spatial feature map; A feature fusion unit is used to fuse a high-temporal and low-spatial feature map with a low-temporal and high-spatial feature map of the same dimension and size to generate a high-temporal and spatial resolution remote sensing image at a specified moment; The spatiotemporal inconsistency detection module further includes a spatial inconsistency unit, a temporal inconsistency unit, and a spatiotemporal feature fusion unit, wherein: The spatial inconsistency unit and the temporal inconsistency unit detect the inconsistency of the forged area from the temporal and spatial perspectives based on high spatiotemporal resolution remote sensing images, and obtain spatial inconsistency feature maps and temporal inconsistency feature maps respectively. The spatiotemporal feature fusion unit is used to capture and fuse the inconsistent information of spatially inconsistent feature maps and temporally inconsistent feature maps, achieve information complementarity, obtain the final feature representation, and realize auxiliary detection of forged targets; The result output module identifies and outputs the coordinates of the tampered area and the tampering method of the remote sensing image based on the final feature representation output by the spatiotemporal inconsistency detection module.

2. A remote sensing image tampering target detection method based on spatiotemporal evolution implemented by the remote sensing image tampering target detection system according to claim 1, characterized in that: The following steps are involved: Step 1: Build a detection model based on the high spatiotemporal resolution generation module, the spatiotemporal inconsistency detection module, and the result output module; Step 2: Train the detection model, including the following steps: Step 201: Collect sample data and build a training data set; Step 202: Train the detection model. The implementation of the detection model includes the following steps: Step 2021: Generate the specified time to be generated using the high spatiotemporal resolution generation module The high temporal and spatial resolution images further include the following steps: Step 20211: The high time and low space units extract the specified time to be generated respectively The low-frequency features in the high-temporal and low-spatial resolution remote sensing images and the features at the specified time to be generated The low-frequency features in the high-temporal and low-spatial resolution remote sensing images at the nearby time t are obtained to obtain the low-frequency feature map, and the size of the low-frequency feature map is expanded to obtain the specified time to be generated. and the final high temporal and low spatial feature map at nearby time t and Step 20212: The low-time high-space unit extracts high-frequency features from the low-time, high-spatial resolution remote sensing image at the nearby time t, and filters the noise information to obtain the final low-time high-spatial feature map. Step 20213, the feature fusion unit combines the low temporal and high spatial feature maps at the nearby time t Subtract the high temporal and low spatial feature maps at nearby time t As a difference reference, then add the difference to the specified time to be generated High temporal and low spatial feature maps Thus, the high and low spatial frequency information of the two feature maps are fused, and the final result is used to generate the specified time to be generated. Fusion characteristics of remote sensing images It is expressed as the following formula: Step 2022: Detection of spatiotemporal inconsistency: The inconsistency of the forged area is detected from the perspectives of time and space by using the spatial inconsistency unit and the temporal inconsistency unit. Specifically, the following steps are included: Step 20221, the spatial inconsistency unit has a three-channel structure. For the input feature map of the spatial inconsistency unit Where T represents the time T, C represents the number of remote sensing image channels, H represents the length of the remote sensing image, and W represents the width of the remote sensing image. The middle channel in the three-channel structure works in a low-resolution manner. The input feature map X1 is average pooled with a kernel size of 2×2 and a stride of 2, and then undergoes two consecutive convolution and bilinear upsampling operations to obtain the output S: Where upsampling represents bilinear upsampling; K1 and K2 are convolutions with kernel sizes of 1×3 and 3×1 respectively; AvgPool2(X1) represents average pooling of the input feature map X1; The upper path in the three-path structure is a residual connection added to the middle path; The output of the upper path and the output of the middle path are fused, and the fusion result is passed through sigmoid to obtain the confidence, which is multiplied with the output of the lower path in the three-path structure after 3x3 convolution to obtain the final distribution confidence score feature map Y1: Where K3 is the 3×3 convolution for feature extraction, K4 is the 3×3 convolution for post-processing, and σ is the sigmoid function; Step 20222: The temporal inconsistency unit observes and models the remote sensing image sequence from the horizontal and vertical directions, respectively, and uses the feature differences of the remote sensing images at adjacent moments along these two orthogonal directions to discover temporally inconsistent regions, including the following steps: 1) Input feature map for time-inconsistent units Input in the horizontal and vertical directions respectively, and then undergo convolution, difference and sigmoid operations to obtain two importance weights F with the same shape as the input feature map X2 h and F w , including the following steps a) After compressing the input feature map X2 by r times in the channel dimension, perform differential calculations along the vertical and horizontal directions to obtain the vertical slice difference map S h and horizontal slice difference map S w ; b) The vertical slice difference map S is enhanced by the vertical time inconsistency enhancement unit VTIE and the horizontal time inconsistency enhancement unit HTIE respectively. h and horizontal slice difference map S w Process and extract vertical slice difference map S h and horizontal slice difference map S w Multi-level representation of c) Use element-by-element addition to fuse the multi-level representations extracted by the vertical time inconsistency enhancement unit VTIE and the horizontal time inconsistency enhancement unit HTIE, and then use the sigmoid function to determine the importance weight F h and F w ; 2) The importance weight F h and F w Multiply by X2 to obtain the enhanced vertical inconsistency feature map Y2 along the time dimension, as shown in the following formula: Where, VTIE(S h ) represents the vertical time inconsistency enhancement unit VTIE for the vertical slice difference map S h Operation, HTIE(S w ) represents the horizontal slice difference map S of the horizontal time inconsistency enhancement unit HTIE w Operation; Step 20223: The spatiotemporal feature fusion unit fuses the inconsistent information captured by the spatial inconsistency unit and the temporal inconsistency unit, effectively utilizing the complementary effect of the two types of information, including the following steps: The spatiotemporal feature fusion unit selects useful channels to supplement the feature map Y1 output by the spatial inconsistency unit through a combination of global average pooling and one-dimensional convolution, as shown in the following formula: Where GAP represents global average pooling, K5 is a one-dimensional convolution with a kernel size of 3, Represents the feature map output by Y1 after global average pooling and one-dimensional convolution; The spatiotemporal feature fusion unit will The final feature representation Y3 is obtained by fusing it with the output of the time-inconsistent unit, as shown in the following formula: Where K6 is a 3×3 convolution; Step 203: The result output module finally identifies and outputs the coordinates of the tampered area and the tampering method of the remote sensing image based on the final feature representation Y3; Step 3: Input the remote sensing image data uploaded by the client in real time into the trained detection model. The detection model outputs the detection results of the coordinates of the tampered area and the tampering method of the remote sensing image on the server.

3. The method for detecting tampered objects in remote sensing images based on spatiotemporal evolution according to claim 2, wherein: In step 201, a large-scale remote sensing image database is constructed through the remote sensing data service provided by the satellite, sample data is collected from the large-scale remote sensing image database, and a training data set for detection model training is constructed. Each sample data in the large-scale remote sensing image database contains the marked tampered area location and tampering method. Finally, a training set, a test set, and a validation set are obtained based on the large-scale remote sensing image database.

4. The method for detecting tampered objects in remote sensing images based on spatiotemporal evolution according to claim 2, wherein: The vertical time inconsistency enhancement unit VTIE will S h As input, it extracts multi-level representations through three branches - 3×1 convolution along the horizontal dimension, average pooling operation, 3×1 convolution along the horizontal dimension and upsampling pooling operation, skip connection; The horizontal time inconsistency enhancement unit HTIE will S w As input, a multi-level representation is extracted through three branches - 3×1 convolution along the vertical dimension, average pooling operation, 3×1 convolution along the vertical dimension and upsampling pooling operation, skip connection.

5. The method for detecting tampered objects in remote sensing images based on spatiotemporal evolution according to claim 2, wherein: In step 3, the client and server are implemented based on the C / S architecture. The server isolates the network through a firewall and transmits image data through an encrypted communication protocol to prevent illegal client access. Large-scale remote sensing image data storage and GPU calculation are performed through distributed nodes. For the result output module, the client is connected to a display and a printer. By setting the result output module signal, the display and printer are connected to realize the screen display and document printing of the diagnostic report.

Citation Information

Patent Citations

  • Remote-sensing image perceptual hash authentication method based on Gabor filter bank and DWT converting

    CN104715440A

  • Method for determining spatio-temporal fusion basic image pair of remote sensing image based on cross fusion

    CN110503137A