Remote Sensing Image Desensitization Method Driven by Text and Historical Images
Through the deep learning model of dual-driven text and historical images, the problem of lack of constraints and insufficient timing consistency in the production of fill content in remote sensing image desensitization technology is solved, and efficient and safe remote sensing image desensitization effect is achieved.
Patent Information
- Application Number
- CN202210945717.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-08-08
AI Technical Summary
The existing intelligent remote sensing image desensitization technology cannot generate additional binding content, and the timing consistency of the desensitization results is insufficient, resulting in an increased risk of leakage in sensitive areas.
Using a dual-driven method of text and historical images, we use deep learning models to guide the generation of sensitive areas to fill content, and ensure the timing consistency of desensitization results based on historical images.
The text content constraints on the results of desensitizing remote sensing images are achieved, which improves the availability and timing consistency of results, and reduces the risk of leaks in sensitive areas.
Smart Images

Figure CN115272059B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing images, and specifically to a method for desensitizing remote sensing images driven by both text and historical images. Background Art
[0002] Remote sensing images are an important type of data for applications such as urban planning and resource surveys, containing rich geographical location information and resource distribution information, which also includes sensitive information that is not suitable for public disclosure, such as important facilities. According to relevant regulations, remote sensing images should be sent to the administrative department of surveying, mapping and geographical information at or above the provincial level for joint review and desensitization processing by relevant departments before being publicly used. With the rapid development of technology, the accuracy and frequency of remote sensing image acquisition are getting higher and higher, and the amount of data obtained is increasing exponentially. Traditional remote sensing image desensitization work often relies on manual processing methods. For example, in image editing software, methods such as image cropping and fusion are used for desensitization. This method is not only time-consuming and laborious, but also has very high requirements for technical personnel, and can no longer meet the rapid desensitization needs of current massive remote sensing images.
[0003] The remote sensing image desensitization technology based on artificial intelligence can generate filling content that is visually consistent with the surrounding scene through big data mining and self-learning mechanisms, greatly improving the automation degree of remote sensing image desensitization. For example, the Chinese patent application with the application number 202110264025.2 and the application name: A method for intelligent local desensitization of geographical grids based on generative adversarial networks, as an existing technology, specifically proposes using an edge discriminator to desensitize the edge map of geographical grid data, and then designing an image completion network. Using the desensitized edge map and the sensitive area mask, the desensitized area is completed. Among them, the image completion network is implemented based on GAN (Generative Adversarial Networks), and uses the features of the non-sensitive area learned to generate the geographical grid image of the missing area from the grayscale image, which to a certain extent solves the problems of low automation degree and lack of usability of the desensitization result in the traditional geographical grid data protection scheme.
[0004] However, the current intelligent remote sensing image desensitization technology still has the following two problems:
[0005] 1) The filling content of the sensitive area is randomly generated by the desensitization algorithm according to the surrounding scene, and no additional constraints can be imposed, such as specifying that the scene of the filling content is grassland, river, square or residential building, etc.;
[0006] 2) The desensitization algorithm only generates filling content based on the current remote sensing image to be desensitized, without taking into account the desensitization results of the previous phase, resulting in drastic changes in the filling content of sensitive areas in remote sensing images of the same region and different phases. This makes it possible to locate the position of sensitive areas by analyzing the differences between the two phases of remote sensing images, which is extremely unfavorable for security. Summary of the invention
[0007] The technical problem to be solved by the present invention is to address the above-mentioned defects in the prior art and to provide a remote sensing image desensitization method driven by both text and historical images. The text is used to guide the generation of filling content in sensitive areas, and the temporal consistency of the desensitization results is ensured based on historical images, thereby solving the problems raised in the above-mentioned background technology.
[0008] According to the present invention, a remote sensing image desensitization method driven by both text and historical images is provided, comprising:
[0009] The first step is to build three deep learning models, which are: a text-generated remote sensing image model based on SSA-GAN, a text-guided remote sensing image desensitization model based on improved ViT and UNet, and a remote sensing image desensitization model driven by both text and historical images;
[0010] The second step: spatially register the remote sensing images of the front and back phases and the sensitive area vectors, cut the remote sensing images of the front and back phases into small images of fixed size, and add text descriptions to the small images to obtain the first front phase training set and the first back phase training set; delete the data in the first front phase training set and the first back phase training set that overlap with the sensitive area vectors in space, so as to obtain the second front phase training set and the second back phase data set that do not contain sensitive areas;
[0011] Step 3: Use the second pre-phase training set and the second post-phase data set to train the SSA-GAN-based text generation remote sensing image model. After the training, the generation model SSA-GAN-TG and the discriminant model SSA-GAN-TD are obtained.
[0012] Step 4: Use the first pre-phase training set and the generative model SSA-GAN-TG to train the text-guided remote sensing image desensitization model based on the improved ViT and UNet. After the training, the desensitization model ViT-UNet-RG and the discriminative model ViT-UNet-RD are obtained.
[0013] Step 5: Use the desensitization model ViT-UNet-RG and the generation model SSA-GAN-TG to process the data in the first pre-phase training set to obtain the output result S1, and combine the first pre-phase training set and the output result into a data set ;
[0014] Sixth step: Use the dataset , the first post-temporal training set and the generative model SSA-GAN-TG to train the remote sensing image desensitization model driven by both training text and historical images. After the training is completed, the desensitization model VU-His-RG and the discriminant model VU-His-RD are obtained;
[0015] Seventh step: Use the desensitization model VU-His-RG, the desensitization model ViT-UNet-RG and the generative model SSA-GAN-TG as the application models of the desensitization inference algorithm. Input the remote sensing images of two consecutive time phases at the same location, the sensitive area mask, the guidance texts of two phases, and the perturbation parameters into the desensitization inference algorithm to obtain the remote sensing image desensitization result that conforms to the temporal consistency of the text content and the result ground objects.
[0016] Preferably, in the second step, the sliding window method is used to cut the remote sensing images of the front and back time phases into small images of a fixed size.
[0017] Preferably, in the second step, the small images are added with text descriptions by manual annotation.
[0018] Preferably, in the second step, the small images in the first pre-temporal training set and the first post-temporal training set are taken out in sequence, and spatial analysis is performed on them with the sensitive area vector. If there is no spatial overlap relationship between the two, the taken-out small images are put into the second pre-temporal training set and the second post-temporal dataset that do not contain the sensitive area.
[0019] Preferably, in the second step, in the QGis software, the geometric correction tool Georeference is used to perform spatial registration on the remote sensing images of the front and back time phases and the sensitive area vector.
[0020] Preferably, the random perturbation during the training in the third step is to randomly extract a set of data from the data conforming to the normal distribution.
[0021] Preferably, the perturbation parameter is a set of data randomly extracted and conforming to the normal distribution.
[0022] Preferably, the process of the desensitization algorithm adopted in the seventh step includes:
[0023] 1) Assume that the remote sensing images X pre 、X aft of two consecutive time phases at the same location, the sensitive area mask M, the guidance texts C pre 、C aft of two phases, and the perturbation parameter N are input;
[0024] 2) When the input data does not contain the pre-temporal data X pre , take X aft 、C aft、Input M and N into the desensitization model ViT-UNet-RG to obtain the desensitization result, and the algorithm ends; otherwise, execute step 3);
[0025] 3) Determine whether the previous time phase used the method of this article for desensitization based on whether there is previous time phase text. If it is determined that the previous time phase used the method of this article for desensitization, then input X pre 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends; otherwise, continue to execute step 4);
[0026] 4) Determine whether the perturbation parameter N is empty to determine whether the previous time phase contains sensitive areas. If N is not empty, clear the text (i.e., C pre =0), input X pre 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends; otherwise, continue to execute step 5);
[0027] 5) Clear the previous time phase text (i.e., C pre =0), clear the previous time phase image (i.e., X pre =0), input X pre 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends.
[0028] The present invention realizes the constraint of using text to desensitize remote sensing images to generate content, and improves the usability of the results. Especially in the consistency of the desensitization results of historical images, the dual-drive image desensitization method using text and historical images is used. The text is used to guide the generation of the filling content of sensitive areas, and the historical images are used to ensure the temporal consistency of the desensitization results, reducing the risk of leakage of sensitive areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Combined with the drawings and by referring to the following detailed description, it will be easier to have a more complete understanding of the present invention and easier to understand its accompanying advantages and features, where:
[0030] Figure 1 Schematically shows the overall flowchart of the text and historical image dual-drive remote sensing image desensitization method according to the preferred embodiment of the present invention.
[0031] Figure 2 Is a schematic diagram of data flow.
[0032] Figure 3 It is the description of the feature recombination module structure.
[0033] Figure 4 It is the description of the variable feature extraction module structure.
[0034] Figure 5 It is the description of the variable feature migration module structure.
[0035] Figure 6 It is the target content selection process based on global structural similarity.
[0036] Figure 7 It is the description of partial Self-Attention calculation.
[0037] Figure 8 It is the sample result.
[0038] Figure 9 It is the desensitization algorithm flow chart.
[0039] It should be noted that the attached drawings are used to illustrate the present invention, rather than limiting the present invention. Note that the attached drawings showing the structure may not be drawn to scale. And in the attached drawings, the same or similar elements are labeled with the same or similar reference numerals. Detailed implementation manners
[0040] In order to make the content of the present invention clearer and easier to understand, the content of the present invention will be described in detail below in conjunction with specific embodiments and the attached drawings.
[0041] Figure 1 Schematically shown is the overall flow chart of the remote sensing image desensitization method driven by text and historical images according to a preferred embodiment of the present invention.
[0042] As Figure 1 shown, the remote sensing image desensitization method driven by text and historical images according to a preferred embodiment of the present invention includes:
[0043] The first step S01: Construct three deep learning models, and the three deep learning models are respectively: a remote sensing image generation model based on SSA-GAN (Semantic-Spatial Aware Generative Adversarial Networks), a text-guided remote sensing image desensitization model based on improved ViT (Vision Transformer) and UNet, and a remote sensing image desensitization model driven by text and historical images;
[0044] Specifically, the SSA-GAN (Semantic-Spatial Aware Generative Adversarial Networks) model is used as the text generation remote sensing image model. SSA-GAN includes a generative model SSA-GAN-TG and a discriminative model SSA-GAN-TD. The input of the model is random perturbation and text sequence, and the output of the model is the generated remote sensing image.
[0045] The text-guided remote sensing image desensitization model based on the improved ViT and UNet consists of a desensitization model ViT-UNet-RG and a discriminative model ViT-UNet-RD.
[0046] The model ViT-UNet-RG is constructed by three modules: ViT-Encoder, Temp, and UNet-Decoder.
[0047] The ViT-Encoder part is improved based on ViT-Base. The specific improvement points are as follows: using PixelwiseNorm to replace the original Norm method, and using partial self-Attention (as shown in Figure 7) to replace the original Multi-head Attention. Among them, X i represents the input image block (i is the subscript, such as X 1 , X 2 , X 3 ), A i represents the result after projection of X i (i is the subscript, such as A 1 , A 2 , A 3 ), Q i , K i , V i , M i represent the query vector, key vector, value vector, and mask vector respectively (i is the subscript, such as Q 1 , K 1 , V 1 , M 1 , Q 2 , K 2 , V 2 , M 2 , Q 3 , K 3 , V 3 , M 3 ). Q i , K i , V i are generated in the same way as in self-Attention, and M iObtained using the calculation method in the following formula (3). A' i represents Q i (i is a subscript, such as A' in the figure 1 、A' 2 、A' 3 ), and the result of calculating with other key vectors is obtained by dot product. All A' i are multiplied by their respective corresponding M i , then softMax normalized, and multiplied by the corresponding V i , and the sum is used as the output X i of the corresponding X iout . Partial self-Attention is an improvement over the original self-Attention. The original self-Attention uses each image block to generate Q, K, and V, and multiplies Q with K of different blocks to obtain the correlation between different blocks. When generating desensitized content, the images in the sensitive areas are removed. In the deep self-Attention, the desensitized areas will contain the features of non-sensitive areas, resulting in checkerboard artifacts in the results and also causing the generated results to deviate from expectations. Partial self-Attention generates Q, K, V, and M at the beginning, where M is the result after projection and sigmoid, used to represent the soft mask in this area. The calculation formula is shown in the following formula (3). After calculating the similarity between Q and K, it is multiplied by M to avoid introducing other features in the missing areas. Finally, after softmax, it is multiplied by the corresponding V, and the sum is obtained as the output. The final calculation formula is formula (2), and the original Attention calculation formula is formula (1).
[0048]
[0049]
[0050] M = sigmeid(W pt WX i ) (3)
[0051] Among them, the parameter W represents the parameter learned during training for projecting X i , the parameter T represents the transpose of the vector, the parameter d k represents the number of channels of the vector K, and the parameter X i represents the input image block.
[0052] In the Temp intermediate layer, there are a change feature migration module, a fusion module, and a recombination module (see Figure 3 ) connected in sequence.
[0053] The change feature migration module only adds the input data in this model.
[0054] The fusion module is composed of two BottleBlocks connected in sequence. The BottleBlock is derived from ResNet.
[0055] The recombination module contains two parallel links: high-dimensional combination and low-dimensional interaction. Let the input feature be X in , in the high-dimensional combination part, X in will first be upsampled by nearest neighbor interpolation to obtain the result X 11 . The result uses a 3×3 convolution to interact the information of each channel and a 1×1 convolution to reduce the number of channels to fuse the features, and then passes through a 3×3 convolution, a 1×1 convolution, and finally a 3×3 convolution with a stride of 2 to downsample back to the original size to obtain X 13 ; in the low-dimensional interaction part, X in is downsampled by a 3×3 convolution with a stride of 2 to obtain X 21 , and then passes through a 3×3 convolution, a 1×1 convolution, and nearest neighbor upsampling to obtain X 23 . Add X 23 , X 13 and the original input and output X out as the final result.
[0056] The UNet-Decoder is the Decoder part of the Unet after discarding the skip connections.
[0057] ViT-UNet-RD adopts the discriminator structure in PatchGAN.
[0058] The data flow process is as follows:
[0059] 1) Let the text be C, the input image be X, the sensitive region mask be M, and the random perturbation be N.
[0060] 2) Convert C into a text sequence and input it into the trained SSA-GAN-TG to obtain the intermediate layer encoded feature F c , the generated artifact X fake , after training, F c already contains the image features related to the text. Cover the image with M to obtain the image X M . Take M as the fourth channel of X M as the input image, convert it into an image sequence using the method in ViT and then input it into the ViT-Encoder to obtain the result feature FE XM ;
[0061] 3) Combine F CAfter being converted into a sequence using the method in ViT, it is input into a single-layer partial self-Attention, and the masking weights M of this layer W are the same as those in the ViT-Encoder of ViT-UNet-RG W and their sum is 1 to obtain F CM ;
[0062] 4) Input FE XM and F CM into the change feature migration module, fusion module, and feature recombination module to obtain the result F E , and finally input it into the UNet-Decoder to obtain the text-driven remote sensing image desensitization result XRG;
[0063] The remote sensing image desensitization driven by both text and historical images consists of a dual-driven desensitization model VU-His-RG (ViT-UNet History Refine Generator) and a discriminant model VU-His-RD (ViT-UNet History Refine Discriminator).
[0064] Among them, VU-His-RG modifies multiple modules based on ViT-UNet-RG. The specific changes are as follows: Duplicate the ViT-Encoder, use two ViT-Encoders to process the data of the previous and current time phases respectively, add a change feature extraction module in the Temp module (see Figure 5 , before the change feature migration module), and modify the algorithm process of the change feature migration module (the process is shown in Figure 6 , and the structure is shown in Figure 4). VU-His-RD has the same structure as ViT-UNet-RD and loads the parameters of the trained ViT-UNet-RD
[0065] For the change feature extraction module, use the feature FE pre of the previous time-phase image and the feature FE aft of the current time-phase image as X now and X history in the figure respectively, connect the input data in the channel dimension, then use a 1×1 convolution to fuse the features of the two-phase images in the insensitive area, and scale the number of channels back to the original size to obtain the feature fusion result X concat . Subsequently, use partial self-Attention to extract the change features in the insensitive area as the new X concat , and then add it to X now as the output
[0066] In the change feature migration module of Temp, first, according to Figure 6Perform the selection process of the target content according to the process, that is, select the filling content X of the sensitive area fill In this module, first calculate the subsequent time-phase mask M 2 Cover the desensitization result XRG of the previous time-phase pre The obtained XRG preM And M 2 The image X after covering the subsequent time-phase aftM Calculate the MI (mutual information) and SSIM (structural similarity) between them, and obtain the global structural similarity P after weighting the results a (P a = 0.5SSIM + 0.5MI), calculate PE pre And FE aft Calculate the cosine similarity between them as the local similarity P f , P a And P f Are weighted by seven to three to obtain the change degree P (P = 1 - (0.7P a + 0.3P f )) Calculate the cosine similarity between the text semantic feature CE pre Of the previous time-phase and the text semantic feature CE aft Of the subsequent time-phase as the relevance P of the text content c . If the change degree P does not exceed the threshold of 0.3, use CE pre As X in the calculation fill ; if it exceeds 0.3 and does not exceed 0.6, weight and fuse CE pre And CE aft As X fill (X fill = P e × CE pre + (1 - P a ) × CE aft ), when it exceeds 0.6, use CE aft As X ftll (For the specific calculation process, see Figure 6 ). Finally, according to Figure 4 Part to achieve the migration of the change feature and text feature of the feature. Among them, X change Is the output result of the change feature extraction module, and X ftll Is determined in the above process. X change Use self-Attention to make the desensitized area contain the change features of the two-phase images to obtain X 2 , then add the result to X ftll , and complete the change feature migration through 3 IRBlocks (the structure proposed in AnimeGAN), and finally restore the data distribution through BN (normalization). Obtain the result F com .
[0067] At this time, the data flow process in VU-His-RG is as follows:
[0068] 1) Let the previous-phase image be X pre , the mask be M 2 , the previous-phase guidance text be C pre , the previous-phase desensitization result be XRG pre , the perturbation parameter be N 2 (extracted from data conforming to a normal distribution), X pre corresponding to the subsequent-phase image X aft , the subsequent-phase guidance text be C aft .
[0069] 2) Cover XRG with M 2 to obtain XRG pre , and use M preM as the fourth channel of XRG 2 as the input image, convert it into an image sequence using the method in ViT and input it into the ViT-Encoder to obtain the previous-phase encoded feature PE preM , convert the previous-phase guidance text C pre into a text sequence and input it together with N pre into SSA-GAN-TG to obtain the text semantic feature CE 2 , the feature P prE after partial self-Attention in the middle layer; use the same method as above to obtain the covered image X preC , the subsequent-phase encoded feature FE aftM , the subsequent-phase text semantic feature CE aft , the feature F aft after partial self-Attention; aftC ;
[0070] 3) Input FE pre and FE aft into the change feature extraction module (the specific structure is as shown in Figure 4 ) to extract the change features FE change of the insensitive region;
[0071] 4) Input FE change IFE pre IFE aft ICE pre ICE aft IXRG perM IX aftM into the change feature migration module to obtain the result F com
[0072] 5) Add F com and FE aft After that, pass them through the fusion module and the feature recombination module, and finally input them into the Decoder to obtain the desensitized remote sensing image XRGH driven by both text and historical images.
[0073] The second step S02: Perform spatial registration on the remote sensing images and the sensitive area vectors of the front and back time phases, cut the remote sensing images of the front and back time phases into small images of a fixed size, and add text descriptions to the small images to obtain the first front-time-phase training set T1 and the first back-time-phase training set T2; delete the data in the first front-time-phase training set T1 and the first back-time-phase training set T2 that has spatial overlap with the sensitive area vectors, so as to obtain the second front-time-phase training set T3 and the second back-time-phase data set T4 that do not contain sensitive areas;
[0074] Among them, for example, the text description can be added to the small images by means of manual annotation.
[0075] For example, in the QGis software, the Georeference geometric correction tool can be used to perform spatial registration on the remote sensing images and the sensitive area vectors of the front and back time phases.
[0076] Preferably, in the second step, the sliding window method is used to cut the remote sensing images of the front and back time phases into small images of a fixed size. For example, the size of the small images is preferably 256×256, and the step size of the sliding window is preferably 192;
[0077] Moreover, preferably, in the second step, the small images in the first front-time-phase training set T1 and the first back-time-phase training set T2 can be taken out in turn, and their spatial analysis is performed with the sensitive area vectors. If there is no spatial overlap relationship between the two, then the taken-out small images are put into the second front-time-phase training set T3 and the second back-time-phase data set T4 that do not contain sensitive areas.
[0078] The third step S03: Use the second front-time-phase training set T3 and the second back-time-phase data set T4 to train the text generation remote sensing image model based on SSA-GAN. After the training is completed, the generation model SSA-GAN-TG and the discriminant model SSA-GAN-TD are obtained;
[0079] Among them, preferably, the random perturbation during the training in the third step is a set of data randomly selected from the data conforming to the normal distribution.
[0080] Preferably, in the third step, the generation method of the text sequence is as follows: Obtain all the texts from the training set, and segment the texts to form a word dictionary; for any input text, segment the input text and obtain the corresponding identification code (Id) from the word dictionary to form a one-hot encoding. The text sequence is the sequence composed of the one-hot encodings of the segmented identification codes.
[0081] Fourth step S04: Use the first pre-phase training set T1 and the generative model SSA-GAN-TG to train the text-guided remote sensing image desensitization model based on the improved ViT and UNet. After training, obtain the desensitization model ViT-UNet-RG and the discriminative model ViT-UNet-RD;
[0082] For example, the mask M of the simulated sensitive area is randomly generated. The generation method is to randomly select a scale from the fixed aspect ratios [1, 0.5, 0.25, 0.3, 0.1, 0.4], randomly generate the coordinate position of the upper left corner, and the length L (L ∈ [0.2, 0.4]) to obtain the final mask. The text C is randomly drawn without replacement from the training set in a complete data iteration. The perturbation parameter is a set of data randomly drawn from a normal distribution.
[0083] For example, during the training in the fourth step, the generative model is optimized using multiple loss functions, including the adversarial loss calculated by extracting the filled content in the remote sensing image desensitization result XRG using the mask M of the sensitive area and the artifacts X generated by SSA-GAN-TG fake in the content, and then calculating the domain-specific perceptual loss (VGG-19), reconstruction loss, and adversarial loss between the remote sensing image desensitization result XRG and the input image X.
[0084] Fifth step S05: Use the desensitization model ViT-UNet-RG and the generative model SSA-GAN-TG to process the data in the first pre-phase training set T1 to obtain the output result S1, and combine the first pre-phase training set T1 and the output result S1 into a data set
[0085] Among them, for example, when identifying the first pre-phase data set T1, the input image is X, the randomly generated mask is M, the text sequence randomly drawn without replacement is C, the random perturbation is N', and the ViT-UNet-RG generates the result set S1[XRG fix ,....]. S1 and T1 form a data set
[0086] Sixth step S06: Use the data set The first post-phase training set T2 and the generative model SSA-GAN-TG to train the text and historical image dual-driven remote sensing image desensitization model. After training, obtain the desensitization model VU-His-RG and the discriminative model VU-His-RD;
[0087] For example, within a complete data iteration cycle, obtain a set of data D (D contains the pre-phase image X from prt , the mask M 2 , the pre-phase guidance text Cpre , the previous-phase desensitization result XRG pre , the perturbation parameter N 2 ), obtain the post-phase image X with the same upper-left corner coordinates as X from the post-phase dataset T2 pre aft , randomly and without replacement extract the post-phase guidance text C aft .
[0088] For example, during training, fix the ViT-Encoder parameters of VU-His-RG and no longer train them. Use M 2 Extract the corresponding positions in the desensitization results generated in XRGH and the generation results of SSA-GAN-TG to calculate the content L2 loss (when P exceeds 0.3 and does not exceed 0.6, calculate the L2 loss between the generation results of the post-phase and pre-phase SSA-GAN-TG and XRGH respectively, and then calculate the weighted sum as the total loss. In other cases, calculate the corresponding part as the loss), calculate the Garm loss for the feature map after feature recombination, and then calculate the domain-specific perception loss (VGG-19), reconstruction loss, and adversarial loss of XRGH and X aft . The sum of the above losses is used as the total loss of the generation model.
[0089] The seventh step S07: Use the desensitization models VU-His-RG, ViT-UNet-RG, and the generation model SSA-GAN-TG as the application models of the desensitization inference algorithm. Input the remote sensing images, sensitive area masks, guidance texts of two phases, and perturbation parameters at the same position in the front and back phases into the desensitization inference algorithm to obtain the desensitization result of the remote sensing image that meets the temporal consistency of the text content and the result ground objects.
[0090] Figure 8 is an example result of the above operations.
[0091] Moreover, for example, as Figure 9 shown, preferably, the process of the desensitization algorithm adopted in the seventh step S07 includes:
[0092] 1) Assume that the remote sensing images of the front and back phases at the same position (X pre , X aft ), the sensitive area mask M, the guidance texts C pre , C aft , and the perturbation parameter N are input;
[0093] 2) When the input data does not contain the previous-phase data X pre , input X aft , C aft , M, and N into the desensitization model ViT-UNet-RG to obtain the desensitization result, and the algorithm ends; otherwise, execute step 3);
[0094] 3) Determine whether the method of this article is used for desensitization in the previous phase based on whether there is previous-phase text. If it is determined that the method of this article is used for desensitization in the previous phase, then X pre 、X aft 、C pre 、C aft 、M and N are input into the desensitization model VU-His-RG model to obtain the desensitization result, and the algorithm ends; otherwise, continue to execute step 4);
[0095] 4) Determine whether the perturbation parameter N is empty to determine whether the previous phase contains sensitive areas. If N is not empty, clear the text (i.e., C pre =0), and input X pre 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends; otherwise, continue to execute step 5);
[0096] 5) Clear the previous-phase text (i.e., C pre =0), clear the previous-phase image (i.e., X pre =0), and input X pre 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends.
[0097] The present invention realizes the constraint of using text to desensitize remote sensing images to generate content, and improves the usability of the results. Especially in terms of the consistency of historical image desensitization results, the dual-drive image desensitization method using text and historical images is used. The text is used to guide the generation of the filling content of sensitive areas, and the historical images are used to ensure the temporal consistency of the desensitization results, reducing the risk of leakage of sensitive areas.
[0098] More specifically, the present invention has at least the following advantages over the prior art:
[0099] 1. The partial Self-Attention structure proposed by the present invention avoids introducing other features in the missing part of the deep features, thereby avoiding the generation of checkerboard artifacts in the results.
[0100] 2. The present invention adopts target content selection based on the degree of change, ensuring the temporal consistency of the desensitization results.
[0101] 3. The remote sensing image desensitization model of the present invention is driven by both text and historical images. By using text and historical image data, it solves the problems that the desensitization result cannot be constrained, resulting in random generation of desensitization results, and the desensitization result cannot take into account the previous-phase ground objects, reducing the risk of leakage of sensitive areas.
[0102] It should be noted that unless otherwise specified, the terms "first", "second", "third", etc. in the specification are only used to distinguish each component, element, step, etc. in the specification, rather than to represent the logical relationship or sequential relationship, etc. between each component, element, step.
[0103] It can be understood that although the present invention has been disclosed above with preferred embodiments, the above embodiments are not intended to limit the present invention. For any person skilled in the art, without departing from the scope of the technical solution of the present invention, many possible changes and modifications can be made to the technical solution of the present invention by using the above-disclosed technical content, or it can be modified into equivalent embodiments with equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A desensitization method for remote sensing images driven by both text and historical images, characterized in that it includes: The first step: Construct three deep learning models, which are respectively: a text generation remote sensing image model based on SSA-GAN, a text-guided remote sensing image desensitization model based on improved ViT and UNet, and a remote sensing image desensitization model driven by both text and historical images; The second step: Perform spatial registration on the remote sensing images and sensitive area vectors of the front and back time phases, cut the remote sensing images of the front and back time phases into small images of a fixed size, and add text descriptions to the small images to obtain the first front time phase training set and the first back time phase training set; Delete the data in the first front time phase training set and the first back time phase training set that has spatial overlap with the sensitive area vector, so as to obtain the second front time phase training set and the second back time phase data set that do not contain sensitive areas; The third step: Use the second front time phase training set and the second back time phase data set to train the text generation remote sensing image model based on SSA-GAN. After the training is completed, the generation model SSA-GAN-TG and the discriminant model SSA-GAN-TD are obtained; The fourth step: Use the first front time phase training set and the generation model SSA-GAN-TG to train the text-guided remote sensing image desensitization model based on improved ViT and UNet. After the training is completed, the desensitization model ViT-UNet-RG and the discriminant model ViT-UNet-RD are obtained; Fifth step: Use the desensitization model ViT-UNet-RG and the generation model SSA-GAN-TG to process the data in the first pre-phase training set to obtain the output result S1, and combine the first pre-phase training set and the output result into a data set Sixth step: Use the dataset The first post-temporal training set and the generative model SSA-GAN-TG are used to train the remote sensing image desensitization model driven by both text and historical images. After the training is completed, the desensitization model VU-His-RG and the discriminant model VU-His-RD are obtained; The seventh step: Use the desensitization model VU-His-RG, the desensitization model ViT-UNet-RG and the generation model SSA-GAN-TG as the application models of the desensitization inference algorithm, input the remote sensing images of the front and back two time phases at the same location, the sensitive area mask, the guidance texts of the two phases, and the perturbation parameters into the desensitization inference algorithm to obtain the remote sensing image desensitization result that conforms to the text content and the temporal consistency of the result ground objects.
2. The desensitization method for remote sensing images driven by both text and historical images according to claim 1, characterized in that, In the second step, the sliding window method is used to cut the remote sensing images of the front and back time phases into small images of a fixed size.
3. The desensitization method for remote sensing images driven by both text and historical images according to claim 1 or 2, characterized in that, In the second step, the text description is added to the small images by means of manual annotation.
4. The desensitization method for remote sensing images driven by both text and historical images according to claim 1 or 2, characterized in that, In the second step, the small images in the first front time phase training set and the first back time phase training set are taken out in turn, and their spatial analysis is carried out with the sensitive area vector. If there is no spatial overlap relationship between the two, the taken out ones are put into the second front time phase training set and the second back time phase data set that do not contain sensitive areas.
5. The desensitization method for remote sensing images driven by both text and historical images according to claim 1 or 2, characterized in that, In the second step, in the QGis software, the geometric correction tool Georeference is used to perform spatial registration on the remote sensing images and sensitive area vectors of the front and back time phases.
6. The desensitization method for remote sensing images driven by both text and historical images according to claim 1 or 2, It is characterized in that , during the training of the third step, the random perturbation is a set of data randomly selected from data conforming to a normal distribution.
7. The remote sensing image desensitization method driven by text and historical images according to claim 1 or 2, It is characterized in that The perturbation parameter is a set of data randomly selected and conforming to a normal distribution.
8. The remote sensing image desensitization method driven by text and historical images according to claim 1 or 2, It is characterized in that The process of the desensitization algorithm adopted in the seventh step includes: 1) Assume that remotely sensed images X pre and X aft at the same location in two consecutive time phases, sensitive area mask M, and guiding texts C pre and C aft in two phases are input, and disturbance parameter N; 2) When the input data does not contain the pre-phase data X pre , input X aft , C aft , M, and N into the desensitization model ViT-UNet-RG to obtain the desensitization result, and the algorithm ends; otherwise, execute step 3); 3) Determine whether the pre-phase uses the method of this article for desensitization based on whether there is pre-phase text. If it is determined that the pre-phase uses the method of this article for desensitization, then input X pre 、X aft 、C pre 、C aft 、M, and N into the desensitization model VU-His-RG model to obtain the desensitization result, and the algorithm ends; otherwise, continue to execute step 4); 4) Determine whether the perturbation parameter N is empty to determine whether the previous phase contains a sensitive area. If N is not empty, clear the text and input X pre , X aft , C pre , C aft , M, and N into the de-identification model VU-His-RG model to obtain the de-identification result, and the algorithm ends; otherwise, continue to execute step 5); 5) Clear the previous phase text, clear the previous image, and input X prt 、X aft 、C pre 、C aft 、M and N into the desensitization model VU-His-RG to obtain the desensitization result, and the algorithm ends.
Citation Information
Patent Citations
An intelligent local desensitization method for geographic grids based on generative adversarial networks
CN113066094B
Medical image safety multi-stage desensitization method and system
CN113051600A
Remote sensing data processing system
CN113112421A