A Deep Learning-Based Automatic Remote Sensing Target Detection and Hiding Method

By combining semantic segmentation and image repair technology, Inception-v3 U-Net and gated convolution network are used to automatically detect and hide sensitive targets in remote sensing images, solving the problems of insufficient hiding and mask immobilization in the existing technology, and achieving efficient and accurate remote sensing target hiding.

CN117173404BActive Publication Date: 2025-05-27BEIJING RES INST OF TELEMETRY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310930222.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-05-27
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively hide sensitive objects in remote sensing images, and the fixed mask used by the existing networks causes the position and shape of the hidden area to be fixed, which is easy to be discovered.

Method used

Using a deep learning-based method, combined with semantic segmentation and image repair technology, the Inception-v3 U-Net segmentation network is used for object detection, and a dynamic mask is generated by a gated convolutional image repair network for image repair, realizing automatic detection and hiding of sensitive targets.

Benefits of technology

Accurate and complete coverage of sensitive targets in remote sensing images is achieved, reducing the probability of target exposure, and reducing the risk of hidden rules being discovered through the use of dynamic masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173404B_ABST
    Figure CN117173404B_ABST
Patent Text Reader

Abstract

The present invention provides a method for automatic detection and hiding of remote sensing targets based on deep learning, including generating a remote sensing target semantic segmentation data set, constructing an Inception-v3U-Net network for semantic segmentation of remote sensing images, constructing a gated convolutional image inpainting network, training the Inception-v3U-Net network, training the gated convolutional image inpainting network, detecting remote sensing targets, completely covering the area where the original remote sensing targets are located, and hiding the original remote sensing targets. The present invention uses semantic segmentation as the target detection model and the gated convolutional network as the image inpainting model, and combines with the method of automatic detection and hiding of remote sensing targets in traditional image processing. During the process of training the gated convolutional network, the generation of random masks is added, breaking the shape and position limitations of the missing areas, and solving the problems that the existing remote sensing target detection network model structure is relatively complex, the training and inference time is too long, the remote sensing target hiding is insufficient, incomplete, and easy to be exposed, and the hiding rule is easy to be discovered due to the use of fixed masks in the existing network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of physical technologies, and particularly relates to an automatic detection and hiding method for remote sensing targets based on deep learning. Background Art

[0002] Remote sensing images have the characteristics of wide coverage and high timeliness of data acquisition, and can effectively provide a large amount of ground object information. According to the corresponding geographical location, target information can be provided, which is of great significance. To prevent the exposure of targets, a method for reducing and desensitizing sensitive targets is needed. Traditional methods for reducing and decrypting remote sensing images include simple splicing, downsampling, image blurring, mosaics, etc. Among them, methods that simply change the pixel distribution of images, such as splicing, image blurring, and mosaics, are easily recognizable, and downsampling and image blurring can restore the target through transfer learning or digital image processing methods. How to effectively hide sensitive areas without exposing sensitive information has become a difficult problem now.

[0003] The development of deep learning technology provides an effective method for this task. The need to process a large amount of image data requires an automated target detection technology, and semantic segmentation technology meets this need very well. Currently, in engineering practice, deep neural network learning technologies such as U-Net are more commonly used for remote sensing image segmentation. The neural network is used to extract the features of remote sensing images, and the trained network is used to predict the category of each pixel, so as to finally obtain a segmentation map with category labels.

[0004] The task of hiding sensitive targets can be handed over to image inpainting to complete. Image inpainting refers to that the deep learning neural network infers the information content of the missing area based on the surrounding area of the missing area of the image and completes the filling. Since the image inpainting technology can effectively train and learn according to the distribution law of ground objects in the scene, it can simulate the distribution law of ground objects in the scene, so as to achieve the effect that is difficult to distinguish by the human eye. The supplemented content is visually real and semantically correct. By removing the area where the sensitive target is located and performing image inpainting, the hiding of sensitive targets can be better realized.

[0005] Pengjie Lu, Dalu Xu, Fu Ren et al. proposed an automatic detection and hiding technology for sensitive targets in emergency remote sensing mapping in their published paper "Automatic Detection and Hiding Methods for Sensitive Targets in Emergency Remote Sensing Mapping" (Journal of Wuhan University of Technology, Information Science Edition, 1671-8860(2020)08-1263-10). This method trains Mask R-CNN to achieve the automatic detection of sensitive targets on remote sensing images. Then, the detection results obtained by this method are used as a mask to input the Deepfill image inpainting network to achieve the automatic detection and hiding of sensitive targets. Among them, Mask R-CNN is a deep convolutional network model for object detection. It adds pre-segmentation of possible target areas on the basis of Faster R-CNN, and does not refine the calibration and pixel segmentation of the memory of this area. Deepfill, on the other hand, consists of dilated convolution and ordinary convolution to form an image segmentation network with two stages of pre-generation and fine-generation. Its input is the original unoccluded image. During the training process, the original image is occluded with a rectangular box, and the occluded image is used for inference and repair, and the original image is used as a reference to form a loss function to train the model. Although this method has achieved good results, there are limitations. As a network model that can simultaneously achieve object detection and semantic segmentation, Mask R-CNN has problems such as complex model structure and long inference and training time for simple task detection. And if the detection results are not processed, the generated mask is not sufficient to completely cover the sensitive targets, and ultimately the sensitive target information will still be exposed. At the same time, using Deepfill for image inpainting will limit the target area to be hidden, making the mask limited to a rectangular area, which will increase the risk of position exposure.

[0006] Therefore, a means of declassification and desensitization that can effectively hide sensitive areas and not expose sensitive information is needed. Summary of the Invention

[0007] The present invention is to solve the problem of hiding sensitive areas, and provides an automatic detection and hiding method for remote sensing targets based on deep learning, using semantic segmentation as the object detection model, gated convolutional network as the image inpainting model, and cooperating with the automatic detection and hiding method of traditional image processing for remote sensing targets to solve the problems of the existing remote sensing object detection network model with complex structure, long training and inference time, insufficient, incomplete and easy to expose hiding of remote sensing targets, and due to the use of a fixed mask in the existing network, the hidden area has a fixed position and shape, and the hiding rule is easy to be discovered.

[0008] The present invention provides an automatic detection and hiding method for remote sensing targets based on deep learning, including the following steps:

[0009] S1. Select and crop the remote sensing images and the corresponding semantic segmentation label images. After selection, obtain the remote sensing target semantic segmentation dataset, and divide the remote sensing target semantic segmentation dataset into a training set, a validation set, and a test set;

[0010] S2. Construct an Inception-v3 U-Net segmentation network. The Inception-v3 U-Net segmentation network includes an Inception-v3 network and a U-Net network. The Inception-v3 network serves as the encoder of the U-Net network and connects the outputs of each layer of the Inception-v3 network to the decoder of the U-Net. The Inception-v3 U-Net segmentation network extracts image features, predicts pixel categories, and performs classification to output the segmentation result;

[0011] S3. Construct a gated convolutional image inpainting network. The gated convolutional image inpainting network includes a spliced pre-generation network and a fine-generation network. The gated convolutional image inpainting network uses a gated mechanism to inpaint the segmentation result covered by the mask;

[0012] S4. Use the remote sensing target semantic segmentation dataset to train the Inception-v3 U-Net segmentation network until the total cost function converges, and obtain the trained Inception-v3 U-Net segmentation network;

[0013] S5. Use the remote sensing target semantic segmentation dataset to train the gated convolutional image inpainting network. Iteratively update the parameters of each layer in the gated convolutional image inpainting network by the gradient descent method until the total cost function converges, and obtain the trained gated convolutional image inpainting network. The dataset for training the gated convolutional image inpainting network is only the remote sensing image, and the remote sensing image needs to be bilinearly interpolated to the specified size;

[0014] S6. Crop and sort the numbers of the remote sensing image to be predicted, and then input it into the trained Inception-v3 U-Net segmentation network to obtain the segmentation result of the cropped remote sensing image;

[0015] S7. Dilate the segmentation result of the cropped remote sensing image to obtain a mask;

[0016] S8. Use the mask to cover the segmentation result of the cropped remote sensing image and input it into the trained gated convolutional image inpainting network to obtain the hidden result of the sensitive target. Then, splice the hidden results of the sensitive target in sequence according to the numbers to obtain the hidden result of the remote sensing image, and the automatic detection and hiding method of remote sensing targets based on deep learning is completed.

[0017] For an automatic detection and hiding method of remote sensing targets based on deep learning according to the present invention, as a preferred method, step S1 includes:

[0018] S11. Select at least 30 remote sensing images with a relatively balanced foreground-to-background ratio and high resolution and their corresponding label images;

[0019] S12. Crop the remote sensing images and label images to a size of 224×224 pixels to obtain the cropped remote sensing images and cropped label images;

[0020] S13. Select the images with a foreground pixel ratio of more than 10% in the cropped label images and their corresponding cropped remote sensing images to form a remote sensing target semantic segmentation data set;

[0021] S14. Divide the remote sensing target semantic segmentation data set into a training set, a validation set, and a test set.

[0022] For an automatic detection and hiding method of remote sensing targets based on deep learning according to the present invention, as a preferred embodiment, step S2 includes:

[0023] S21. Construct an encoder based on the Inception-v3 network. The Inception-v3 network includes 1 input module connected in series and 3 Inception modules with the same structure;

[0024] S22. Construct a minimum convolution module;

[0025] S23. Construct a decoder, which is an upsampling sub-network composed of 3 upsampling modules with the same structure and 1 CBR module connected in series;

[0026] S24. Connect the inputs of the 3 upsampling modules in the decoder to the outputs of the 3 Inception modules in the encoder respectively and connect the output of the input module to the input of the CBR module through the concatenate method to obtain an Inception-v3 segmentation network with a skip connection structure;

[0027] S25. Connect the encoder with a skip connection structure, the minimum convolution module, and the decoder in series to obtain an Inception-v3 U-Ne segmentation network.

[0028] For an automatic detection and hiding method of remote sensing targets based on deep learning according to the present invention, as a preferred embodiment, in step S21, the input module includes a convolutional layer, a pooling layer, a BatchNorm layer, 2 convolutional layers, a BatchNorm layer, and a pooling layer connected in sequence;

[0029] In step S22, the minimum convolution module includes 3 convolutional layers connected in series;

[0030] In step S23, the upsampling module includes three consecutive convolutional layers, a BatchNorm layer, an activation layer, and an upsampling layer; the structure of the CBR module is, in sequence: a convolutional layer, a BatchNorm layer, and an activation layer;

[0031] In step S24, H(I) = Up i ([Incep i ,Up i-1 );

[0032] where I is the input feature, H is the output, Up i is the i-th upsampling layer, and [Incep i ,Up i-1 is the result of concatenating the i-th Inception-v3 and the (i - 1)-th upsampling layer at the channel level.

[0033] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention, in step S21, the convolutional kernel size of the first convolutional layer of the input module is 7×7, the stride is 2, the convolutional kernel of the second convolutional layer is 1×1, the stride is 1, the convolutional kernel of the third convolutional layer is 3×3, the stride is 1, and the padding for all is set to "same"; both the first pooling layer and the second pooling layer use max pooling layers, the pooling window is 2×2, and the stride is 2;

[0034] In step S22, the convolutional kernels of the three convolutional layers of the minimum convolutional module are all 1×1, the strides are all 1, and the padding parameter is 1;

[0035] In step S23, the convolutional kernel sizes of the first three convolutional layers of the upsampling module are all 3×3, the strides are all 1, the padding is all 1, the negative slope of the activation layer is 0.2, the activation layer uses the LeakyReLU function, and the nearest neighbor upsampling by a factor of two is used for the upsampling layer; the convolutional kernel of the convolutional layer of the CBR module is 3×3, the stride is 1, the padding for all is 1, the negative slope of the activation layer is set to 0.2, and the activation layer is implemented using the LeakyReLU function.

[0036] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention,

[0037] The pre-generation network can simulate the mask to cover the image repeatedly, and the features extracted by the convolution of the mask can be false masks with different pixel values and different regions. The fine-generation network learns the pixel distribution characteristics of the remaining remote sensing image and dynamically updates the network parameters;

[0038] Step S3 includes:

[0039] S31. Construct a pre-generation network, which consists of 12 gated convolutional layers, an upsampling layer, a gated convolutional layer, an upsampling layer, two consecutive gated convolutional layers, and an activation layer in sequence;

[0040] S32. Construct a fine-generation network. The last eight layers of the fine-generation network are the same as the last eight layers of the pre-generation segmentation network. The first to the tenth layers of the fine-generation network are parallel branches one and two;

[0041] S33. Connect the pre-generation network and the fine-generation network in series to obtain a gated convolutional image inpainting network.

[0042] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention, in step S31,

[0043] In step S32, branch one is the same as the first ten layers of the pre-generation network;

[0044] Branch two has one less layer than branch one. The first five layers and the last two layers of branch two are the same as the corresponding layers of the first five layers and the last two layers of the pre-generation network. An activation function Relu is added to the sixth layer of branch two, and a content attention mechanism is added to the seventh layer;

[0045] The outputs of branch one and branch two are concatenated through a concatenate operation.

[0046] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention, in step S31, the convolutional kernel size of the first gated convolutional layer is 5×5, the stride is 1, the strides of the second and fourth gated convolutional layers are both 2, the sizes of the remaining gated convolutional kernels are all 3×3, the strides are all 1, dilated convolutions are used in the seventh to tenth gated convolutional layers, and the activation function of the activation layer is Tanh.

[0047] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention, in step S3, the gating mechanism is:

[0048] G,F=Conv(I);

[0049] H=GConv(I)

[0050] =σ(G)·φ(F);

[0051] Where G and F are two feature matrices obtained by evenly dividing a single convolution result, H is the output of the gated convolution, σ and φ are the sigmoid and Tanh activation functions respectively, and · is the matrix multiplication operation of the two matrices in the element dimension.

[0052] In a preferred embodiment of the automatic remote sensing target detection and hiding method based on deep learning according to the present invention, in step S4, the Inception-v3 U-Net segmentation network is trained using a training set. The optimizer is Adam, the learning rate is 0.0005, and the gradient descent method of the neural network is used to iteratively update the Inception-v3 U-Net network model.

[0053] In step S5, the gated convolutional image inpainting network is trained using the remote sensing images in the training set. The remote sensing images are bilinearly interpolated to 256×256. The optimizer is Adam, the learning rate is 0.0001, and the optimizer parameters beta1 is 0.5 and beta2 is 0.999.

[0054] During the process of training the gated convolutional image inpainting network, the generation of a random mask is added. The random mask is an irregular mask formed by connecting randomly discrete dot matrices. An irregular mask based on a rectangle is generated by superimposing the irregular mask on a regular matrix mask. The random mask is generated by a random generation algorithm.

[0055] In step S6, the remote sensing image to be predicted is cropped to 224×224.

[0056] In step S7, the segmentation result of the cropped remote sensing image is dilated once with a 12×12 kernel function image and then bilinearly interpolated to 256×256 to obtain a mask, which can completely cover the segmentation result of the cropped remote sensing image.

[0057] In step S8, the segmentation result of the cropped remote sensing image covered by the mask is cropped to 224×224 and then input into the trained gated convolutional image inpainting network.

[0058] The present invention is a technology for the automatic detection of sensitive targets in remote sensing mapping and the inpainting of remote sensing images, which can be used for the automatic detection and hiding of sensitive target areas in remote sensing images, and further realize the declassification and desensitization of sensitive targets in automatic remote sensing mapping.

[0059] The present invention provides a method for automatically detecting and hiding remote sensing image targets based on deep learning, which is used for automatically detecting and hiding sensitive targets in remote sensing mapping. The implementation steps are as follows: 1. Generate a remote sensing target semantic segmentation dataset; 2. Construct an Inception-v3 U-Net network for remote sensing image semantic segmentation; 3. Construct a gated convolutional image inpainting network; 4. Train the Inception-v3 U-Net network; 5. Train the gated convolutional image inpainting network; 6. Detect remote sensing targets; 7. Completely cover the area where the original remote sensing target is located; 8. Hide the original remote sensing target. The present invention constructs a training network Inception-v3 U-Net, which utilizes the multi-scale and multi-branch feature extraction ability of Inception-v3 to handle the situation of different target sizes and complex features in remote sensing images. At the same time, with the help of the efficient training ability of U-Net, the training and inference time of the network for detecting remote sensing targets is greatly reduced, and the training efficiency is improved. The present invention adopts a method combining a segmentation network and image dilation to generate a covering mask, realizing the accurate and complete coverage of the mask for remote sensing targets, and reducing the exposure probability of remote sensing targets. Since the dilation operation can expand the target area obtained by the semantic segmentation network and can meet the complete coverage of the target in most cases, a faster and more efficient target detection network can be used in the front-end technology.

[0060] The idea for achieving the object of the present invention is that the present invention constructs a training network Inception-v3 U-Net, which utilizes the multi-scale and multi-branch feature extraction ability of Inception-v3 to handle the situation of different target sizes and complex features in remote sensing images. At the same time, with the help of the efficient training ability of U-Net, the training and inference time of the network for detecting remote sensing targets is greatly reduced, and the training efficiency is improved.

[0061] The present invention adopts a method combining a segmentation network and image dilation to generate a covering mask, realizing the accurate and complete coverage of the mask for remote sensing targets, and reducing the exposure probability of remote sensing targets. Since the dilation operation can expand the target area obtained by the semantic segmentation network and can meet the complete coverage of the target in most cases, a faster and more efficient target detection network can be used in the front-end technology.

[0062] The present invention constructs a gated convolutional image inpainting network, which uses a gating mechanism to provide fake masks with different pixel values and regions in the front-stage network to train the model's ability to handle different missing regions. The latter part can be used as a special attention mechanism to dynamically learn the pixel distribution characteristics of the remaining remote sensing images and improve the ability to forge remote sensing images.

[0063] During the process of training the gated convolutional inpainting network, the present invention adds the generation of random masks, which breaks the shape and position limitations of the missing regions, making the network more suitable for hiding remote sensing targets with diverse positions and shapes. The random hiding method will greatly reduce the problem of the hiding pattern being discovered.

[0064] The present invention combines semantic segmentation and image inpainting technologies. The semantic segmentation technology is used to automatically detect the target to provide the location of the target and its complex shape contour, and then the image inpainting technology is used to hide the remote sensing target. Moreover, the present invention uses artificially provided label samples as the training target of the image inpainting technology. The finally trained network can generate false targets based on the artificially provided mask, and these targets will be more difficult to be detected by the human eye.

[0065] The specific steps of the present invention are as follows:

[0066] Step 1, generate a semantic segmentation dataset for remote sensing image targets:

[0067] Step 1.1, randomly select at least 30 high-resolution remote sensing images with a relatively balanced foreground-to-background ratio and their corresponding label images;

[0068] Step 1.2, crop each high-resolution image and its corresponding label image to a size of 224×224 pixels;

[0069] Step 1.3, select the images with a foreground pixel ratio of more than 10% in the cropped label images and the corresponding remote sensing images to form a dataset;

[0070] Step 2, construct an Inception-v3 U-Net segmentation network:

[0071] The network structure of the present invention adopts an Inception-v3 U-Net network. The Inception-v3 network is used as the encoder of the U-Net network, and its outputs of each layer are connected to the decoder of the U-Net.

[0072] Step 2.1, construct an encoder.

[0073] The encoder is based on the Inception-v3 network. The structure of the Inception-v3 network consists of 1 input module and 3 identical Inception modules connected in series. The structure of the input module is as follows in sequence: the first convolutional layer, the first pooling layer, the first BatchNorm layer, the second convolutional layer, the third convolutional layer, the second BatchNorm layer, and the second pooling layer.

[0074] Set the convolutional kernel size of the first convolutional layer to 7×7, the convolutional kernel of the second convolutional layer to 1×1, and the convolutional kernel of the third convolutional layer to 3×3. Set the stride of the first convolutional layer to 2, and the strides of the second and third convolutional layers to 1. Set the padding for all edges to "same". Both the first and second pooling layers use max pooling layers, with the pooling window size set to 2×2 for both and the stride set to 2 for both.

[0075] Step 2.2, construct the minimum convolutional module:

[0076] Build a convolutional module composed of three convolutional layers connected in series, namely the first convolutional layer, the second convolutional layer, and the third convolutional layer;

[0077] The convolutional kernel sizes of the first, second, and third convolutional layers are 1×1, the strides are all 1, and the padding parameter is 1;

[0078] Step 2.3, construct the decoder;

[0079] Build an upsampling sub-network composed of 3 identical upsampling modules and 1 CBR module connected in series as the decoder. The structure of each upsampling module is as follows in sequence: the first convolutional layer, the second convolutional layer, the third convolutional layer, the BatchNorm layer, the activation layer, and the upsampling layer;

[0080] Set the convolutional kernel sizes of the first to third convolutional layers to 3×3, the strides to 1, and the padding to 1 for all. Set the slope of the negative part of the activation layer to 0.2, and the activation layer is implemented using the LeakyReLU function. The upsampling layer uses bilinear nearest neighbor upsampling;

[0081] The structure of the CBR module is as follows in sequence: the convolutional layer, the BatchNorm layer, and the activation layer;

[0082] Set the convolutional kernel size of the convolutional layer to 3×3, the stride to 1, and the padding for all edges to 1. Set the slope of the negative part of the activation layer to 0.2, and the activation layer is implemented using the LeakyReLU function;

[0083] Step 2.4, in the concatenate manner, connect the inputs of the 3 upsampling modules in the decoder with the outputs of the 3 Inception modules in the encoder respectively, and connect the output of the input module with the input of the CBR module, and then form an Inception-v3 network with a skip connection structure;

[0084] Step 2.5, connect the encoder with a skip connection structure, the minimum convolution module, and the decoder in series to form an Inception-v3 U-Net network;

[0085] Step 3, construct a gated convolution image inpainting network:

[0086] Step 3.1, construct a pre-generation network;

[0087] The pre-generation network consists of 12 gated convolution layers, the thirteenth upsampling layer, the fourteenth gated convolution layer, the fifteenth upsampling layer, two consecutive gated convolution layers, and the eighteenth activation layer.

[0088] Among them, the convolution kernel size of the first convolution layer is 5×5, the stride is 1, the strides of the second and fourth convolution layers are 2, and the convolution kernel sizes of the rest are all 3×3, and the stride is 1. Dilated convolution is used in the seventh to tenth gated convolution layers. The activation function of the eighteenth layer is Tanh activation.

[0089] Step 3.2, construct a fine-generation network;

[0090] The last eight layers of the fine-generation network are the same as the last eight layers of the pre-generation network. The first to tenth layers are two parallel branches.

[0091] Branch one is the same as the first ten layers of the pre-generation network.

[0092] Branch two has one less layer of network than branch one. Its first five layers and the last two layers are the same as the corresponding layers of the first five layers and the last two layers of the pre-generation network. The sixth layer adds the activation function Relu, and the seventh layer adds a content attention mechanism.

[0093] The outputs of branch one and branch two are concatenated through concatenate.

[0094] Step 3.3, connect the pre-generation network and the fine-generation network in series;

[0095] Step 4, train the Inception-v3 U-Net network:

[0096] Input the training set into the Inception-v3 U-Net network. The optimizer is Adam, and the learning rate is 0.0005. Use the gradient descent method of the neural network to iteratively update the Inception-v3 U-Net network model until the total cost function converges, and obtain the trained Inception-v3 U-Net network;

[0097] Step 5, train the gated convolutional image inpainting network:

[0098] Input the training set into the gated convolutional image inpainting network. The optimizer is Adam, the learning rate is 0.0001, and the optimizer parameters beta1 is 0.5 and beta2 is 0.999. Use the gradient descent method to iteratively update the parameters of each layer in the gated convolutional image inpainting network until the total cost function converges, and obtain the trained gated convolutional image inpainting network;

[0099] Step 6, detect remote sensing targets:

[0100] Step 6.1, crop the remote sensing image to be predicted to a size of 224×224, and sort and label the cropped remote sensing image to be predicted;

[0101] Step 6.2, input the images with serial numbers into the trained Inception-v3 U-Net network in sequence to obtain the segmentation result of the cropped remote sensing image;

[0102] Step 7, completely cover the area where the original remote sensing target is located:

[0103] Perform a dilation operation on the remote sensing image segmentation result with a kernel size of 12×12, and the number of dilation times is 1 to obtain a mask, which can completely cover the area where the original remote sensing target is located.

[0104] Step 8, hide the original remote sensing target:

[0105] Step 8.1, use the mask to cover the cropped image and input it into the gated convolutional image inpainting network to obtain the hiding result of the target.

[0106] Step 8.2, splice the target hiding results of the cropped remote sensing images in sequence according to the serial numbers to obtain the final hiding result.

[0107] The present invention has the following advantages:

[0108] (1) The present invention uses the Ineceptionv3 U-Net network to detect remote sensing targets, and uses the Inception-v3 model structure as the backbone network of the encoder part. Compared with the original target detection network, the model structure is simpler. By utilizing the multi-scale and multi-branch feature extraction ability of Inception-v3, it can handle the situation where the target sizes in remote sensing images are different and the features are complex. At the same time, this semantic segmentation network also has the efficient training unique to the U-Net network, which can greatly reduce the training and inference time of the network, reduce the consumption of human and material resources in the target preprocessing stage, and provide a mask template with variable shapes that fits the contours of remote sensing targets for subsequent image restoration techniques, reducing the modification of non-target areas in remote sensing images and increasing the difficulty of detection.

[0109] (2) The present invention adds an image dilation operation after the Ineceptionv3 U-Net detection result. According to the detection result, its edges are expanded. After dilation, the mask can completely cover the area where the original remote sensing target is located and its surroundings. Compared with directly using a rectangular area as the covering mask in the existing image, it can further reduce the hidden danger of target exposure. At the same time, this operation can also make up for the errors in the detected area when there are differences between the detection result and the target.

[0110] (3) The present invention uses a gated convolutional image restoration network to repair the incomplete remote sensing image covered by the mask. This network uses a gating mechanism and also uses the gated convolutional network method in dilated convolution and content attention mechanism.

[0111] The formula for gated convolution can be expressed as:

[0112] [G,F]=Conv(I);

[0113] H=GConv(I)

[0114] =σ(G)·φ(F);

[0115] Where G and F are two feature matrices obtained by evenly dividing the result of a single convolution, H is the output of gated convolution, σ and φ are the sigmoid and Tanh activation functions respectively, and · is the matrix multiplication operation of the two matrices in the element dimension.

[0116] Compared with ordinary convolution, in the front part of the network, gated convolution can simulate the repeated covering of the image by the mask. And since the mask is extracted through convolution, its features may be false masks with different pixel values and regions. Compared with the original mask with only 0 and 1 values, it can train the model's ability to handle different masks and the ability of the network to restore random masks. In the latter part, gated convolution is similar to the attention mechanism and can dynamically update the network parameters. Finally, with a small increase in computational complexity, the restoration ability of the network is improved.

[0117] (4) In the process of training the gated convolutional image inpainting network of the present invention, the generation of a random mask is added. An irregular mask based on a rectangle is generated in the way of forming an irregular mask by connecting random discrete dot matrices and superimposing a regular matrix mask, breaking the limitations on the shape and position of the missing area in the image inpainting technology. Compared with the existing technology of modifying images within a rectangular area, the network obtained by this training method is more adaptable to the situation where the shapes, sizes, and positions of remote sensing targets to be hidden are different.

[0118] (5) The present invention combines semantic segmentation, traditional image processing, and image inpainting technologies to perform declassification and desensitization operations on remote sensing targets, and the processing flow based on code enables this operation to be automated. Since the training time is greatly shortened in the semantic segmentation part and slightly increased in the image inpainting part, the overall time for hiding remote sensing image targets is shortened, and the present invention mainly spends time on the hiding part of remote sensing targets, greatly improving the time utilization rate. Description of the Drawings

[0119] Figure 1 It is a flowchart of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0120] Figure 2 It is a schematic structural diagram of an Inception-v3 U-Net network of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0121] Figure 3 It is a schematic structural diagram of a gated convolutional image inpainting network of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0122] Figure 4a It is a remote sensing image in the test set of Embodiment 1 of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0123] Figure 4b It is a segmentation result map of the test set of the process of an automatic detection and hiding method for remote sensing targets based on deep learning in Embodiment 1;

[0124] Figure 4c It is a result map of directly hiding the target by using the segmentation result as a mask in the process of Embodiment 1 of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0125] Figure 4d It is a segmentation label map corresponding to the remote sensing image in the process of Embodiment 1 of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0126] Figure 4e It is a result map after dilation of the segmentation result in the process of Embodiment 1 of an automatic detection and hiding method for remote sensing targets based on deep learning;

[0127] Figure 4f Result graph of target hiding after dilation in Embodiment 1 of the process of an automatic remote sensing target detection and hiding method based on deep learning. Detailed implementation manners

[0128] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0129] Embodiment 1

[0130] As Figure 1 shown, an automatic remote sensing target detection and hiding method based on deep learning includes the following steps:

[0131] Step 1, generate a remote sensing target semantic segmentation dataset.

[0132] This dataset is mainly a semantic segmentation dataset. The original image and the labeled image are used as the training set during semantic segmentation training, and only the original remote sensing image is used during the training of the image inpainting network. The mask is randomly generated by a random generation algorithm.

[0133] Step 1.1, in the embodiment of the present invention, 30 high-resolution remote sensing images with a size of 1500×1500 and their corresponding label images with a relatively balanced foreground and background ratio are randomly selected from the Massachusetts building dataset, where the foreground is the target of semantic segmentation.

[0134] Step 1.2, each high-resolution image and its corresponding label image are cropped into 224×224 pixel sizes one by one.

[0135] Step 1.3, select the label images with a foreground pixel ratio of more than 10% in the cropped label images, and all the cropped remote sensing images to form a training set.

[0136] Step 2, construct an Inception-v3 U-Net segmentation network:

[0137] As Figure 2 shown, a further detailed description is made of the Inception-v3 U-Net network constructed in the present invention. The network structure of the present invention adopts the Inception-v3 U-Net network. The Inception-v3 network is used as the encoder of the U-Net network, and the outputs of its layers are connected to the decoder of the U-Net.

[0138] Step 2.1, construct an encoder.

[0139] The encoder is based on the Inception-v3 network. The structure of the Inception-v3 network consists of 1 input module and 3 identical Inception modules connected in series. The structure of the input module is as follows in sequence: the first convolutional layer, the first pooling layer, the first BatchNorm layer, the second convolutional layer, the third convolutional layer, the second BatchNorm layer, and the second pooling layer.

[0140] Set the convolutional kernel size of the first convolutional layer to 7×7, the convolutional kernel of the second convolutional layer to 1×1, the convolutional kernel of the third convolutional layer to 3×3. Set the stride of the first convolutional layer to 2, the strides of the second and third convolutional layers to 1, and the padding for all to "same". Both the first and second pooling layers use max pooling layers, with the pooling window size set to 2×2 for both and the stride set to 2 for both.

[0141] Step 2.2, construct the minimum convolutional module:

[0142] Build a convolutional module composed of three convolutional layers connected in series, namely the fourth convolutional layer, the fifth convolutional layer, and the sixth convolutional layer;

[0143] The convolutional kernel sizes of the fourth, fifth, and sixth convolutional layers are 1×1, the strides are all 1, and the padding parameter is 1;

[0144] Step 2.3, construct the decoder;

[0145] Build an upsampling sub-network composed of 3 identical upsampling modules and 1 CBR module connected in series as the decoder. The structure of each upsampling module is as follows in sequence: 3 convolutional layers connected in series, BatchNorm layer, activation layer, and upsampling layer;

[0146] Set the convolutional kernel sizes of the 3 convolutional layers to 3×3, the strides to 1, the padding to 1 for all. Set the slope of the negative part of the activation layer to 0.2, and the activation layer is implemented using the LeakyReLU function. The upsampling layer uses bilinear nearest neighbor upsampling;

[0147] The structure of the CBR module is as follows in sequence: convolutional layer, BatchNorm layer, activation layer;

[0148] Set the convolutional kernel size of the convolutional layer to 3×3, the stride to 1, the padding for all edges to 1, the slope of the negative part of the activation layer to 0.2, and the activation layer is implemented using the LeakyReLU function;

[0149] Step 2.4, in the concatenate manner, connect the inputs of the three upsampling modules in the decoder with the outputs of the three Inception modules in the encoder respectively, and connect the output of the input module with the input of the CBR module, and then form an Inception-v3 network with a skip connection structure. The formula for this part can be expressed as:

[0150] H(I) = Up i ([Incep i ,Up i-1 );

[0151] where I represents the input feature, H is the output, Up i represents the i-th upsampling layer, and [Incep i ,Up i-1 is the result of concatenating the i-th Inception-v3 and the (i - 1)-th upsampling layer at the channel level;

[0152] Step 2.5, connect the encoder with a skip connection structure, the minimum convolution module, and the decoder in series to form an Inception-v3 U-Net network;

[0153] Step 3, construct a gated convolution image inpainting network:

[0154] As Figure 3 shown, a further detailed description of the gated convolution image inpainting network constructed by the present invention is given.

[0155] Step 3.1, construct a pre-generation network;

[0156] The pre-generation network consists of twelve gated convolution layers, an upsampling layer, a gated convolution layer, an upsampling layer, two consecutive gated convolution layers, and an activation layer in sequence.

[0157] Among them, the first convolution layer has a convolution kernel size of 5×5 and a stride of 1, the second and fourth convolution layers have a stride of 2, and the remaining convolution kernels have a size of 3×3 and a stride of 1. The seventh to tenth gated convolution layers use dilated convolution. The activation function of the activation layer is Tanh.

[0158] Step 3.2, construct a fine-generation network;

[0159] The last eight layers of the fine-generation network are the same as the last eight layers of the pre-generation network. The first to tenth layers are two parallel branches.

[0160] Branch one is the same as the first ten layers of the pre-generation network.

[0161] Branch two has one less layer of network than branch one. Its first five layers and the last two layers are the same as the corresponding layers of the first five layers and the last two layers of the pre-generated network. The activation function Relu is added to the sixth layer, and the content attention mechanism is added to the seventh layer.

[0162] The outputs of branch one and branch two are concatenated through a concatenate operation.

[0163] Step 3.3, concatenate the pre-generated network and the fine-generated network in series;

[0164] Step 4, train the Inception-v3 U-Net network:

[0165] Input the training set into the Inception-v3 U-Net network. The optimizer is Adam, the learning rate is 0.0005, and the gradient descent method of the neural network is used to iteratively update the Inception-v3 U-Net network model until the total cost function converges, obtaining the trained Inception-v3 U-Net network. The training set is the original image and the corresponding segmentation labels;

[0166] Step 5, train the gated convolutional image inpainting network:

[0167] Input the training set into the gated convolutional image inpainting network. The optimizer is Adam, the learning rate is 0.0001, the optimizer parameter beta1 is 0.5, and beta2 is 0.999. The gradient descent method is used to iteratively update the parameters of each layer in the gated convolutional image inpainting network until the total cost function converges, obtaining the trained gated convolutional image inpainting network. The training set is only the original image, and the image needs to be bilinearly interpolated to a size of 256×256;

[0168] Step 6, detect remote sensing targets:

[0169] Step 6.1, crop the remote sensing image to be predicted to a size of 224×224 and sort and label the cropped remote sensing image to be predicted;

[0170] Step 6.2, input the images with serial numbers into the trained Inception-v3 U-Net network in sequence to obtain the segmentation results of the cropped remote sensing images;

[0171] Step 7, completely cover the area where the original remote sensing target is located:

[0172] Perform a dilation operation on the remote sensing image segmentation result with a kernel size of 12×12 and a dilation times of 1 to obtain a mask. The dilated result can not only completely cover the target and its area, but also make up for the errors in the detection area in the case where there are differences between the detection results and the target.

[0173] Step 8, Hide the original remote sensing target:

[0174] Step 8.1, Use a mask to cover the cropped image and input it into the gated convolutional image inpainting network to obtain the hiding result of the sensitive target.

[0175] Step 8.2, Stitch the target hiding results of the cropped remote sensing images in sequence according to the serial numbers to obtain the final hiding result.

[0176] The effect of the present invention can be further illustrated by the following simulation experiments.

[0177] 1. Simulation experiment conditions:

[0178] The hardware platform for the simulation experiment of the present invention is: the processor is Intel i9-10940X, the main frequency is 3.30GHz, the memory is 64G, the graphics card is NVIDIA GeForce RTX 2080Ti, and the video memory is 12GB.

[0179] The software platform for the simulation experiment of the present invention is: Windows 10 operating system and python3.6.

[0180] The data used in the simulation experiment of the present invention comes from 30 groups of randomly selected data in the remote sensing building dataset Massachusetts, which consists of 151 aerial images in the Boston area. Each image is 1500×1500 in size. The foreground of this dataset is buildings, including various buildings.

[0181] 2. Simulation experiment content and result analysis:

[0182] The simulation experiment of the present invention is carried out according to the following steps.

[0183] The semantic segmentation method uses the Inception-v3 U-Net network to test the segmentation ability and training time of the network.

[0184] The image inpainting ablation experiment method respectively uses the directly predicted result as a mask for image inpainting and the result after image dilation as a mask for image inpainting.

[0185] Step A, Crop the selected data in the Massachusetts dataset to a size of 224×224 as the standard image size of this experiment. Select the dataset according to the foreground ratio threshold of 10%. Among them, 3960 samples are divided into the training set, 144 samples are divided into the validation set, and 360 samples are divided into the test set.

[0186] Step B: Input the images and corresponding labels of the training set into the Inception-v3 U-Net network for training. A total of 90 epochs are trained, and the network weight parameters are stored every 30 epochs. During training, after each epoch of training, the performance of the current model is evaluated on the validation set using evaluation metrics.

[0187] Step C: Only bilinearly interpolate the images in the training set to a size of 256×256, and input them into the gated convolutional image inpainting network. Train in batches of 16 images, save the model every 4000 epochs, and conduct experiments on the validation set every 2000 epochs.

[0188] Step D: Input the test set images into Inception-v3 U-Net to obtain the segmentation results.

[0189] Step E: Bilinearly interpolate the segmentation results to a size of 256×256 and directly input them into the trained gated convolutional image inpainting network. The final result is then reduced to a size of 224×224.

[0190] Step F: Dilate the segmentation results once using a 12×12 kernel function and then bilinearly interpolate them to a size of 256×256. Then input them into the trained gated convolutional image inpainting network. The final result is also reduced to a size of 224×224.

[0191] The following Figures 4a to 4f simulation diagrams are used to further describe the effects of the present invention.

[0192] Figure 4a is the remote sensing image in the test set. Figure 4b is the segmentation result of the test set. Figure 4c is the result of directly hiding the target using the segmentation result as a mask. Figure 4d is the segmentation label map corresponding to the remote sensing image. Figure 4e is the result after dilating the segmentation result. Figure 4f is the result of hiding the target after dilation.

[0193] From Figure 4c , 4f it can be seen that the hiding result of the present invention is better than directly using the prediction result for hiding. Due to the insufficient coverage area of the region masked by the segmentation result, the gated convolutional image inpainting network fails to achieve the hiding of building targets. However, the result after dilation has a larger coverage area, and its hiding effect as a mask is better, achieving complete hiding of building targets.

[0194] Using the following formulas, calculate the three evaluation metrics of recall, accuracy, and intersection over union (IoU) to evaluate the segmentation results, and plot all the calculation results in Table 1:

[0195] The formula for recall is as follows:

[0196]

[0197] Where TP is the number of samples that are actually positive samples and the predicted results are also positive samples, and FN is the number of samples that are actually positive samples but the predicted results are negative samples.

[0198] The formula for accuracy is as follows:

[0199]

[0200] Where FP is the number of samples that are actually negative samples but the predicted results are positive samples, and TN is the number of samples that are actually negative samples and the predicted results are also negative samples.

[0201] The formula for intersection over union (IoU) is as follows:

[0202]

[0203] Table 1. Quantitative analysis table of the segmentation results of the method of the present invention and the ablation experiment method in the simulation experiment

[0204]

[0205]

[0206] The larger the values of Recall, Acc, and IoU in Table 1, the more accurate the segmentation result. From Table 1, it can be seen that the recall of the present invention is 63.70%, the accuracy is 88.04%, and the intersection over union (IoU) is 54.63%. For simple targets, these metrics are sufficient to detect most of the targets. It can also be seen from Figure 4b that all buildings have been detected. After the subsequent image dilation, Figure 4e the mask area of Figure 4d completely covers the building targets in Figure 4f . The final result

[0207] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered by the protection scope of the present invention.

Claims

1. A method for automatic detection and hiding of remote sensing targets based on deep learning, characterized in that: It includes the following steps: S1. Select remote sensing images and corresponding semantic segmentation label images, perform cropping and selection to obtain a remote sensing target semantic segmentation data set, and divide the remote sensing target semantic segmentation data set into a training set, a validation set, and a test set; S2. Construct an Inception-v3 U-Net segmentation network. The Inception-v3 U-Net segmentation network includes an Inception-v3 network and a U-Net network. The Inception-v3 network serves as the encoder of the U-Net network and connects the outputs of each layer of the Inception-v3 network to the decoder of the U-Net. The Inception-v3 U-Net segmentation network extracts image features, predicts pixel categories, and performs classification to output a segmentation result; S3. Construct a gated convolutional image inpainting network. The gated convolutional image inpainting network includes a pre-generation network and a fine-generation network connected in series. The gated convolutional image inpainting network uses a gated mechanism to inpaint the segmentation result covered by the mask; S4. Use the remote sensing target semantic segmentation data set to train the Inception-v3 U-Net segmentation network until the total cost function converges to obtain a trained Inception-v3 U-Net segmentation network; S5. Use the remote sensing target semantic segmentation data set to train the gated convolutional image inpainting network. Iteratively update the parameters of each layer in the gated convolutional image inpainting network by the gradient descent method until the total cost function converges to obtain a trained gated convolutional image inpainting network. The data set for training the gated convolutional image inpainting network is only the remote sensing image, and the remote sensing image needs to be bilinearly interpolated to a specified size; S6. Crop and sort the numbers of the remote sensing image to be predicted and input them into the trained Inception-v3 U-Net segmentation network to obtain the segmentation result of the cropped remote sensing image; S7. Dilate the segmentation result of the cropped remote sensing image to obtain a mask; S8. Use the mask to cover the segmentation result of the cropped remote sensing image and input it into the trained gated convolutional image inpainting network to obtain a sensitive target hiding result, and then splice the sensitive target hiding results in sequence according to the numbers to obtain a remote sensing image hiding result, and the method for automatic detection and hiding of remote sensing targets based on deep learning is completed.

2. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 1, characterized in that: Step S1 includes: S11. Select at least 30 remote sensing images and corresponding label images with relatively balanced foreground and background ratios and high resolutions; S12. Crop the remote sensing images and the label images to a size of 224×224 pixels to obtain cropped remote sensing images and cropped label images; S13. Select the images with the foreground pixel ratio above 10% in the cropped label images and the corresponding cropped remote sensing images to form the remote sensing target semantic segmentation dataset; S14. Divide the remote sensing target semantic segmentation dataset into a training set, a validation set, and a test set.

3. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 1, characterized in that: Step S2 includes: S21. Construct an encoder based on the Inception-v3 network. The Inception-v3 network includes 1 input module connected in series and 3 Inception modules with the same structure; S22. Construct a minimum convolution module; S23. Construct a decoder, which is an upsampling sub-network composed of 3 upsampling modules with the same structure and 1 CBR module connected in series; S24. Connect the inputs of the 3 upsampling modules in the decoder to the outputs of the 3 Inception modules in the encoder respectively and connect the output of the input module to the input of the CBR module in a concatenate manner to obtain an Inception-v3 segmentation network with a skip connection structure; S25. Connect the encoder with the skip connection structure, the minimum convolution module, and the decoder in series in sequence to obtain an Inception-v3 U-Net segmentation network.

4. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 3, characterized in that: In step S21, the input module includes a convolutional layer, a pooling layer, a BatchNorm layer, 2 convolutional layers, a BatchNorm layer, and a pooling layer connected in sequence; In step S22, the minimum convolution module includes 3 convolutional layers connected in series; In step S23, the upsampling module includes 3 convolutional layers, a BatchNorm layer, an activation layer, and an upsampling layer connected in sequence; the structure of the CBR module is: a convolutional layer, a BatchNorm layer, and an activation layer in sequence; In step S24, H(I) = Up i ([Incep i ,Up i-1 ) Where I is the input feature, H is the output, and Up i is the i-th upsampling layer, [Incep i , Up i-1 is the result of concatenating the i-th Inception-v3 and the (i - 1)-th upsampling layer at the channel level.

5. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 4, characterized in that: In step S21, the convolutional kernel size of the first convolutional layer of the input module is 7×7, the stride is 2, the convolutional kernel of the second convolutional layer is 1×1, the stride is 1, the convolutional kernel of the third convolutional layer is 3×3, the stride is 1, and the padding at the edges is set to "same"; the first pooling layer and the second pooling layer both use a max pooling layer, the pooling window is 2×2, and the stride is 2; In step S22, the convolutional kernels of the 3 convolutional layers of the minimum convolution module are all 1×1, the strides are all 1, and the padding parameter is 1; In step S23, the convolutional kernel sizes of the first 3 convolutional layers of the upsampling module are all 3×3, the strides are all 1, the padding is all 1, the negative slope of the activation layer is 0.2, the activation layer uses the LeakyReLU function, and the upsampling layer uses bilinear nearest neighbor upsampling; The convolution kernel of the convolution layer of the CBR module is 3×3, the stride is 1, and the padding on both edges is 1. The negative slope of the activation layer is set to 0.2, and the activation layer is implemented using the LeakyReLU function.

6. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 1, characterized in that: The pre-generation network can simulate the repeated coverage of the image by the mask, and the features extracted by the convolution of the mask can be false masks with different pixel values and different regions. The fine-generation network learns the pixel distribution characteristics of the remaining remote sensing image and dynamically updates the network parameters; Step S3 includes: S31. Construct the pre-generation network, which consists of 12 gated convolution layers, an upsampling layer, a gated convolution layer, an upsampling layer, two consecutive gated convolution layers, and an activation layer in sequence; S32. Construct the fine-generation network. The last eight layers of the fine-generation network are the same as the last eight layers of the pre-generation segmentation network. The first to tenth layers of the fine-generation network are parallel branches one and two; S33. Connect the pre-generation network and the fine-generation network in series to obtain the gated convolution image inpainting network.

7. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 6, characterized in that: In step S32, branch one is the same as the first ten layers of the pre-generation network; Branch two has one less layer of network than branch one. The first five layers and the last two layers of branch two are the same as the corresponding layers of the first five layers and the last two layers of the pre-generation network. The sixth layer of branch two adds the activation function Relu, and the seventh layer adds the content attention mechanism; The outputs of branch one and branch two are spliced together through a concatenate operation.

8. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 6, characterized in that: In step S31, the convolution kernel size of the first gated convolution layer is 5×5 and the stride is 1. The strides of the second gated convolution layer and the fourth gated convolution layer are both 2. The remaining gated convolution kernels are all 3×3 and the strides are all 1. The seventh to tenth gated convolution layers use dilated convolution, and the activation function of the activation layer is Tanh.

9. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 1, characterized in that: In step S3, the gating mechanism is: [G,F]=Conv(I); H=GConv(I) =σ(G)·φ(F); where G and F are two feature matrices obtained by evenly dividing the result of a single convolution, H is the output of the gated convolution, σ and φ are the sigmoid and Tanh activation functions respectively, and · is the matrix multiplication operation of the two matrices in the element dimension.

10. A method for automatic detection and hiding of remote sensing targets based on deep learning according to claim 1, characterized in that: In step S4, the Inception-v3 U-Net segmentation network is trained using the training set, the optimizer is Adam, the learning rate is 0.0005, and the gradient descent method of the neural network is used to iteratively update the Inception-v3 U-Net network model; In step S5, the gated convolutional image inpainting network is trained using the remote sensing images in the training set, and the remote sensing images are bilinearly interpolated to 256×256; the optimizer is Adam, the learning rate is 0.0001, and the optimizer parameters beta1 is 0.5 and beta2 is 0.999; During the process of training the gated convolutional image inpainting network, the generation of a random mask is added. The random mask is an irregular mask formed by connecting randomly discrete dot matrices. An irregular mask based on a rectangle is generated by superimposing the irregular mask on a regular matrix mask, and the random mask is generated through a random generation algorithm; In step S6, the remote sensing image to be predicted is cropped to 224×224; In step S7, the segmentation result of the cropped remote sensing image is dilated once with a 12×12 kernel function and then bilinearly interpolated to 256×256 to obtain the mask, and the mask can completely cover the segmentation result of the cropped remote sensing image; In step S8, the segmentation result of the cropped remote sensing image covered by the mask is cropped to 224×224 and then input into the trained gated convolutional image inpainting network.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on hierarchical contour cost function

    CN114972759A

  • Training method and device and image restoration method and device

    CN115660062A