Remote Sensing Image Sensitive Target Elimination Method, Device, Equipment and Storage Medium

By blocking preprocessing of high-resolution remote sensing images and deep learning models to identify sensitive targets, combined with adaptive expansion and automatic repair of AOT-GAN models, the efficiency and quality problems of eliminating sensitive targets in the existing technology are solved, and efficient and accurate removal of sensitive information is achieved.

CN119850418BActive Publication Date: 2025-05-27NANJING NORMAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336575.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-27
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

When the prior art eliminates sensitive targets in high-resolution remote sensing images, it is difficult to achieve efficient and accurate identification and elimination of sensitive targets while retaining the application value of the image, and the calculation efficiency and detailed reconstruction effect are poor.

Method used

By blocking preprocessing of high-resolution remote sensing images, using deep learning models to identify sensitive targets and generate masks, combining adaptive expansion strategies to optimize mask quality, and automatically repair them through the AOT-GAN model, ultimately achieving efficient elimination of sensitive targets.

Benefits of technology

It realizes efficient and accurate detection and removal of sensitive information in high-resolution remote sensing images in complex scenarios, reducing residual and edge transition unnatural phenomena in target elimination, and ensuring the integrity and authenticity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850418B_ABST
    Figure CN119850418B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, device and storage medium for eliminating sensitive targets in remote sensing images, which relates to the technical field of remote sensing image processing. The method includes: performing block processing on a high-resolution large-size image to obtain an area of interest containing sensitive targets, and converting the geographical coordinates of the area of interest into pixel coordinates; based on the target area, using a trained instance segmentation model to perform target detection on each target area, generating a corresponding segmentation mask, dynamically adjusting the dilation degree of the segmentation mask according to the target area, and determining the mask area; generating filling content and filling it into the mask area to obtain an image after target elimination, and splicing the image after target elimination to the original image to obtain a spliced image. The present application meets the requirements of efficient, accurate and seamless elimination of sensitive targets in large-scale and high-resolution remote sensing images, and provides a reliable technical solution for the secure sharing of remote sensing image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of remote sensing image processing, and particularly relates to a method, device, equipment and storage medium for eliminating sensitive targets in remote sensing images. Background Art

[0002] With the rapid development of remote sensing technology and satellite observation technology, the acquisition of high-resolution remote sensing images has become increasingly convenient. Due to their rich spatial texture features and semantic information, these images have shown great application potential in multiple fields. Large-scale, high-resolution satellite remote sensing images have a wide spatial coverage, rich details, and accurate geographical information. However, these images often contain sensitive information, such as national key facilities or content information related to citizens' privacy. Remote sensing image data containing sensitive information should be declassified and desensitized according to national standards to meet the needs of public sharing. Therefore, the problem of removing sensitive information in large-scale high-resolution remote sensing images has gradually become the focus of research. How to effectively eliminate sensitive targets while retaining the application value of the images has become an urgent problem to be solved.

[0003] In the field of eliminating sensitive targets in high-resolution remote sensing images, traditional methods mainly rely on image processing and machine learning technologies. These methods played a certain role in early research and were able to achieve the task of eliminating sensitive targets to a certain extent. However, with the increasing complexity of the application scenarios of high-resolution images, their limitations have gradually emerged. High-resolution images contain rich details and complex scene information. When traditional methods process such images, it is often difficult to accurately identify the target edges, resulting in serious edge blurring. At the same time, due to the lack of effective extraction ability for complex texture features, key details are easily lost during the target elimination process, greatly reducing the quality of the processed image and making it difficult to achieve a balance between elimination accuracy and image quality.

[0004] In recent years, the remarkable progress of deep learning technology in the field of computer vision has brought new opportunities to the fields of object detection and image inpainting. Currently, commonly used object detection models include the YOLO series, Faster R-CNN, and Transformer-based object detection networks. These models can effectively extract complex texture features in images, thus significantly improving the detection accuracy and efficiency. For example, the literature "CAO Changqing, WANG Bo, ZHANG Wenrui, et al. An improved faster R-CNN for small object detection[J]. Ieee Access, 2019, 7:106838-106846." proposed an improved Faster R-CNN algorithm, which is significantly superior to traditional models in small object detection by optimizing the IoU loss function, enhancing the RoI pooling operation, and multi-scale feature fusion. The literature "JIE Xue, ZHENGYongguo, DONG-YE Changlei, et al. Improved YOLOv5 network method for remote sensing image-based ground objects recognition[J]. Soft Computing, 2022, 26(20): 10879-10889." significantly improved the recognition accuracy of large objects in high-resolution images by improving the backbone network and attention mechanism of the YOLOv5 model. However, in the continuous optimization of the model structure, the increase in computational complexity has limited the real-time performance of the algorithm when processing large-scale high-resolution images, which is particularly significant in practical applications. Therefore, how to effectively reduce the computational cost while ensuring the detection accuracy and achieve efficient and real-time sensitive object detection and elimination is a difficult problem that needs to be overcome in current research.

[0005] The application of deep learning in the field of image inpainting has also made remarkable progress. Models such as Generative adversarial network (GAN), Convolutional Neural Networks (CNN), and Transformer have shown great potential in improving the quality of image inpainting and handling complex scenes. Taking the context encoder proposed in the literature "Pathak D, Krahenbuhl P, Donahue J, et al. Context encoders: Feature learning by inpainting. Proceedings of the IEEE conference on computer vision and pattern recognition[C], 2016:2536-2544." as an example, by generating the content of any region in the image, it has learned to capture the appearance and semantic structure of the image, and is not only used for inpainting tasks, but also as an effective pre-training method for CNN models. The deep generative model proposed in the literature "YU Jiahui, ZHE Lin, YANG Jimei, SHEN Xiaohui, et al. Huang; Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition[C], 2018:5505-5514." with context attention mechanism has significantly improved the inpainting quality and training efficiency. However, although these inpainting models perform well in general image inpainting tasks, when faced with the task of removing sensitive targets in high-resolution remote sensing images, they expose many problems such as computational efficiency and detail reconstruction.

[0006] In summary, existing methods usually focus on improving the accuracy of target recognition in the improvement of deep learning models. However, in the practical application of large-scale high-resolution remote sensing images, real-time performance and computational efficiency are equally important, and achieving a balance between recognition accuracy and efficiency has become a key consideration. In addition, in current research, target recognition and image inpainting are rarely combined for sensitive target elimination, and this combination is of great significance in ensuring the usability and security of images. Summary of the Invention

[0007] The present application provides a method, apparatus, device, and storage medium for eliminating sensitive targets in remote sensing images, aiming to accurately and efficiently detect and remove sensitive content in images. First, the application preprocesses the image by dividing it into blocks to reduce the computational burden and improve processing efficiency. Subsequently, a deep learning model is used to identify sensitive targets in the image and generate corresponding masks to block the sensitive areas. To ensure that the masks accurately cover the targets and reduce interference with the background, the algorithm adopts an adaptive dilation strategy to adjust the masks. In response to the phenomena of missed detection and false detection in the detection, a manual review link is introduced to ensure that all sensitive targets are accurately marked. After manual confirmation, the elimination operation of the sensitive targets is automatically executed, and the processed image blocks are seamlessly stitched back to the original image. By combining automated detection with manual-assisted detection, this method aims to improve the accuracy and efficiency of sensitive target elimination and provide a feasible solution for the secure application of remote sensing images.

[0008] In a first aspect, the present application provides a method for eliminating sensitive targets in remote sensing images, including:

[0009] Performing block processing on a high-resolution large-size image to obtain regions of interest containing sensitive targets, and converting the geographic coordinates of the regions of interest into pixel coordinates;

[0010] Based on the target regions, using a trained instance segmentation model to perform target detection on each target region and generate corresponding segmentation masks, dynamically adjusting the dilation degree of the segmentation masks according to the target area, and determining the mask regions; wherein, the target regions are regions containing sensitive targets manually drawn in the regions of interest;

[0011] Generating filling content and filling it into the mask regions to obtain an image after target elimination, and stitching the image after target elimination to the original image to obtain a stitched image.

[0012] In a possible design, the geographic coordinates of the regions of interest are converted into pixel coordinates through the following formula:

[0013] ;

[0014] In the formula, x and y respectively represent the abscissa and ordinate of the image pixels; X and Y respectively represent the geographic abscissa and geographic ordinate of the region of interest; a , b , c , d , e and f all represent affine transformation parameters.

[0015] In a possible design, the trained instance segmentation model.

[0016] In a possible design, the expansion degree of the segmentation mask is dynamically adjusted according to the target area, and the calculation process for determining the mask area is expressed as:

[0017] ;

[0018] In the formula, is the expansion kernel size; area is the mask area; 500 is the scaling factor used to adjust the expansion degree; max is the maximum value function; min is the minimum value function; int is the nearest distance in the specified edge set.

[0019] In a possible design, generating filling content and filling it into the mask area to obtain an image after target elimination, and splicing the image after target elimination to the original image to obtain a spliced image, including:

[0020] Generating filling content through the following formula and filling it into the mask area to obtain an image after target elimination:

[0021] ;

[0022] In the formula, is the filling content generated by the generator, is the original image; is the image after target elimination; is the mask area;

[0023] Splicing the image after target elimination to the original image through the following formula to obtain a spliced image:

[0024] ;

[0025] In the formula, is the spliced image; is the mask of the segmentation area; is the number of the segmentation area.

[0026] In a possible design, after obtaining the spliced image, the method further includes:

[0027] Evaluating the spliced image using evaluation metrics; the evaluation metrics include edge deviation, Hausdorff distance, and IoU; among them, the edge deviation and Hausdorff distance are used to evaluate the matching degree between the mask edge and the true target boundary, and the IoU is used to evaluate the regional overlap consistency;

[0028] Calculating the edge deviation through the following formula:

[0029] ;

[0030] In the formula, is the total number of real edge pixels; and respectively represent the pixel positions of the real edge and the mask edge; represents and the Euclidean distance between; Dedge represents the edge deviation; represents the real target edge;

[0031] The Hausdorff distance is calculated by the following formula:

[0032] ;

[0033] In the formula, sup represents the maximum distance in the specified edge set; inf represents the nearest distance to the specified edge set, represents the real target edge and the generated mask edge of the Hausdorff distance;

[0034] The IoU is calculated by the following formula:

[0035] ;

[0036] In the formula, represents the number of intersection pixels of the predicted mask and the real mask ; represents the number of union pixels of the predicted mask and the real mask ;

[0037] In a second aspect, the present application provides a remote sensing image sensitive target elimination device, and the device includes:

[0038] An image block module, configured to perform block processing on a high-resolution large-size image to obtain an interested area containing sensitive targets, and convert the geographical coordinates of the interested area into pixel coordinates;

[0039] A detection and correction module, configured to perform target detection on each target area based on the target area by using a trained instance segmentation model, generate a corresponding segmentation mask, and dynamically adjust the dilation degree of the segmentation mask according to the target area to determine the mask area; wherein, the target area is an area containing sensitive targets manually drawn in the interested area;

[0040] The splicing elimination module is configured to generate filling content and fill it into the mask area to obtain an image after target elimination, and splice the image after target elimination to the original image to obtain a spliced image.

[0041] In a third aspect, an embodiment of the present application provides an electronic device, including: at least one processor and a memory; the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the remote sensing image sensitive target elimination method described in the first aspect and various possible designs of the first aspect above.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the remote sensing image sensitive target elimination method described in the first aspect and various possible designs of the first aspect above is implemented.

[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the remote sensing image sensitive target elimination method described in the first aspect and various possible designs of the first aspect above is implemented.

[0044] The remote sensing image sensitive target elimination method, device, equipment and storage medium provided by the present application have at least the following beneficial effects:

[0045] The present application demonstrates excellent detection and elimination performance in complex scenarios, can effectively reduce phenomena such as residues and unnatural edge transitions in target elimination, and ensure the integrity and authenticity of the image. The present application provides a reliable technical solution for the protection of sensitive information in high-resolution remote sensing images and is applicable to a variety of actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0047] Figure 1 It is a flowchart of a remote sensing image sensitive target elimination method provided by an embodiment of the present application;

[0048] Figure 2 It is a mask dilation comparison diagram provided by an embodiment of the present application;

[0049] Figure 3 It is an AOT-GAN model structure diagram provided by an embodiment of the present application;

[0050] Figure 4 It is a target elimination effect comparison diagram provided by an embodiment of the present application;

[0051] Figure 5 This is the experimental result diagram of processing efficiency and resource consumption provided by the embodiments of this application; among them, (a), resource occupancy for 1.5G image processing (5000 frames); (b), resource occupancy for 1.5G image processing (10000 frames); (c), resource occupancy for 5G image processing (5000 frames); (d), resource occupancy for 5G image processing (10000 frames);

[0052] Figure 6 This is the experimental result diagram of processing efficiency and resource consumption provided by the embodiments of this application; among them, (a), resource occupancy for 10G image processing (5000 frames); (b), resource occupancy for 10G image processing (10000 frames); (c), resource occupancy for 23G image processing (5000 frames); (d), resource occupancy for 23G image processing (10000 frames);

[0053] Figure 7 This is the structural schematic diagram of a remote sensing image sensitive target elimination device provided by the embodiments of this application.

[0054] Through the above-mentioned drawings, the clear embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0055] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application. On the contrary, they are only examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0056] In the technical solution of this application, the processing of collection, storage, use, processing, transmission, provision, and disclosure of information such as financial data or user data all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0057] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0058] The following uses specific embodiments to elaborate in detail on the technical solutions of this application and how the technical solutions of this application solve the above technical problems. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below in conjunction with the accompanying drawings.

[0059] Currently, there is still a lack of efficient, accurate, and intelligent methods for eliminating sensitive targets in large-scale, high-resolution remote sensing images. Due to the cumbersome operations such as cropping, splicing, and calibration required by traditional methods, it is difficult to adapt to the characteristics of large amounts of data and multi-scale targets in remote sensing images, and the elimination effect is not good. Therefore, the embodiments of this application provide a method for eliminating sensitive targets in remote sensing images. This method first performs image segmentation to locate the area where the sensitive target is located, then uses the YOLOv8 instance segmentation model to identify the sensitive target and generate a mask, and proposes a mask adaptive dilation method to optimize the mask quality. Finally, the AOT-GAN model is used to fill the background of the mask area uniformly, achieving the seamless elimination of sensitive targets. Experimental results show that this algorithm meets the requirements of efficient, accurate, and seamless elimination of sensitive targets in large-scale, high-resolution remote sensing images, providing a reliable technical solution for the secure sharing of remote sensing image data.

[0060] Figure 1 It is a flowchart of a method for eliminating sensitive targets in remote sensing images provided by the embodiments of this application. As Figure 1 shown, this method for eliminating sensitive targets in remote sensing images aims to effectively process sensitive targets in large-size, high-resolution remote sensing images, taking into account both detection accuracy and processing efficiency to meet the actual application requirements. To achieve this goal, the algorithm designs modules such as image segmentation preprocessing, manual-assisted mask correction, sensitive target recognition and elimination, etc., to ensure the precise elimination of sensitive targets.

[0061] First, optimize the detection process through image segmentation preprocessing. Before detection, for remote sensing images with extremely large sizes, the algorithm will select the region of interest (ROI) and perform regional segmentation to narrow the processing range, thereby effectively reducing the computational load and improving the processing efficiency. In the target recognition stage, the algorithm uses the YOLOv8 instance segmentation model to automatically detect the sensitive target area.

[0062] The quality of target elimination highly depends on the quality of the masks generated by instance segmentation. If the mask coverage is insufficient, the exposure of the edges of the target area will result in residues or a harsh transition of the target after elimination, affecting the elimination effect; while an overly large mask area may obscure the background, affecting the realism of the area after elimination. Therefore, the algorithm performs adaptive dilation on the mask area to ensure complete coverage of the target while minimizing the occlusion of the background, so as to maintain the visual effect of restoration. In addition, the algorithm also provides an artificial correction function, which can manually adjust the mask area after detection to ensure the accurate marking of all sensitive targets, providing a reliable basis for subsequent elimination steps.

[0063] In the target elimination stage, the algorithm introduces the AOT-GAN model for automatic restoration. AOT-GAN has advantages in processing complex scenes in high-resolution images and can generate filled areas with high realism and rich details, thus ensuring the visual consistency and usability of the restored image.

[0064] Generally speaking, through the organic combination of image block preprocessing, automatic detection, artificial correction, and automatic restoration, this method provides a feasible technical solution for the elimination of sensitive targets in large-size high-resolution remote sensing images, and is suitable for performing efficient and accurate sensitive information protection tasks in practical applications.

[0065] Specifically, the method for eliminating sensitive targets in remote sensing images includes steps S10 to S30.

[0066] S10: Perform block processing on the high-resolution large-size image to obtain the region of interest containing sensitive targets, and convert the geographic coordinates of the region of interest into pixel coordinates.

[0067] Before target detection, first perform block processing on the high-resolution large-size image to reduce the computational complexity and improve the processing efficiency. For high-resolution remote sensing images with a large size (above ten billion pixels), first determine the region of interest (ROI) containing sensitive targets, and then convert its geographic coordinates into pixel coordinates to ensure spatial reference consistency during the detection and elimination of sensitive targets. Assume that there is the following affine transformation relationship between the geographic coordinates and pixel coordinates of the image:

[0068] ;

[0069] In the formula, x and y respectively represent the abscissa and ordinate of the image pixel; X and Y respectively represent the geographic abscissa and geographic ordinate of the region of interest; a , b , c , d ,e and f both represent affine transformation parameters, including the geographical origin of the image, the pixel resolution, and the image rotation angle.

[0070] S20: Based on the target regions, use the trained instance segmentation model to perform target detection on each target region, generate corresponding segmentation masks, and dynamically adjust the dilation degree of the segmentation masks according to the target area to determine the mask regions; wherein, the target regions are regions manually drawn in the region of interest that contain sensitive targets.

[0071] Step S20 is the sensitive target recognition stage. To balance detection accuracy and processing efficiency while retaining more background textures and context information, this paper adopts a sensitive target detection method based on YOLOv8 instance segmentation and combines the adaptive dilation of masks and artificial auxiliary correction techniques to improve the target elimination effect.

[0072] First, manually draw regions containing sensitive targets in the ROI region, and perform sensitive target mask generation and recognition operations within these regions. This strategy aims to ensure that the detection regions meet the best input size requirements of the YOLO model, thereby maximizing the recognition performance. After determining the target detection regions, use the trained YOLOv8 instance segmentation model to perform target detection on each region and generate corresponding segmentation masks.

[0073] In some embodiments, to optimize the quality of the segmentation masks, an adaptive dilation method based on the target area is proposed. Through empirical formula (2), this method can dynamically adjust the dilation degree according to the target area: larger targets use a larger dilation range, while smaller targets are moderately dilated, making the dilation effect vary non-linearly with the target area. This design balances the target coverage rate and the background retention rate, effectively improving the quality of subsequent repair processing. In addition, the generated masks can be manually corrected to further ensure the accuracy of the masks. The comparison before and after mask dilation is Figure 2 as shown.

[0074] During the dilation process, the minimum value of the dilation kernel is set to 3 (for small targets), and the maximum value is 7 (for large targets). Through this adaptive dilation strategy, while completely covering the targets, the information of the background region can be retained to the greatest extent, improving the effect and naturalness of sensitive target elimination. The formula for adaptive dilation is as follows:

[0075] ;

[0076] In the formula, is the kernel size of dilation; areais the mask area; 500 is the scaling factor used to adjust the degree of dilation; max is the maximum value function; min is the minimum value function; int is the nearest distance in the specified edge set.

[0077] S30: Generate filling content and fill it into the mask area to obtain the image after target elimination, and splice the image after target elimination to the original image to obtain the spliced image.

[0078] After confirming the mask area, in step S30, the AOT-GAN (Aggregated Contextual-Transformation GAN) model is introduced to automatically fill and repair the sensitive target area. AOT-GAN is a high-resolution image repair model based on the generative adversarial network, which performs particularly well in the filling task of large-scale free-form missing areas. The model structure of AOT-GAN is as Figure 3 shown.

[0079] In some embodiments, the filling content generated by AOT-GAN is defined as follows:

[0080] ;

[0081] In the formula, is the filling content generated by the generator, is the original image; is the image after target elimination; is the mask area.

[0082] After the elimination of sensitive targets is completed, the algorithm reassembles the filled segmentation area back to the original image and adjusts the color difference to ensure that the splicing area is visually Figure 1 consistent with the original so as to obtain a high-quality elimination effect without seamless color difference. The image after reassembly

[0083] ;

[0084] In the formula, is the spliced image; is the mask of the segmentation area to ensure that the splicing result is visually Figure 1 consistent with the original is the number of the segmentation area.

[0085] Through the organic combination of the above-mentioned image block division, automatic detection, mask correction and AOT-GAN elimination, this embodiment realizes efficient and accurate elimination of sensitive targets in the processing of high-resolution images of sensitive targets, and is applicable to privacy protection and information security tasks in practical applications.

[0086] In this embodiment, an aircraft is used as the sensitive target for the experiment, and the dataset used is the fine-grained target recognition dataset (FAIR1M) of high-resolution remote sensing images independently developed by Aerospace Information Research Institute. 3000 images containing aircraft were selected from this dataset to construct the training set of the YOLOv8 instance segmentation model. The annotation of the training set adopted the semi-automatic image annotation tool ISAT proposed in the literature "JI Shuwei, ZHANG Hongyuan. 'ISAT with Segment Anything: An Interactive Semi-Automatic Annotation Tool, 2023.' URL: https: / / github.com / yatengLG / ISAT_with_segment_anything, updated on (2023): 06-03." to complete the preliminary instance segmentation mask annotation, and the incomplete annotation areas caused by reasons such as the similarity of the target object's color to the background or shadows were manually corrected or re-annotated to ensure the accuracy and integrity of the annotation.

[0087] In the selection of the instance segmentation model, the YOLOv8-seg series provides multiple pre-trained models, and their complexities increase from simple to complex as YOLOv8n-seg.pt, YOLOv8s-seg.pt, YOLOv8m-seg.pt, YOLOv8l-seg.pt, and YOLOv8x-seg.pt. The higher the model complexity, the more parameters and the stronger the feature extraction ability, thus improving the detection and segmentation accuracy. In this embodiment, the most complex YOLOv8x-seg.pt is selected as the pre-trained model and trained for 100 rounds to obtain higher detection accuracy and segmentation quality.

[0088] After the training is completed, the obtained optimal model (best.pt) is used as the final model for instance segmentation, and target detection and mask generation are performed on the above 3000 remote sensing images to provide the mask data required for AOT-GAN training. Based on these 3000 remote sensing images and their corresponding segmentation masks, the AOT-GAN model completed the filling training of the sensitive target area after 12500 iterations. By combining the repair ability of the YOLOv8 instance segmentation and the AOT-GAN model, the experimental process in this paper achieved the accurate recognition and elimination of sensitive targets such as aircraft in high-resolution remote sensing images.

[0089] In the experimental stage, to comprehensively evaluate the performance of the algorithm on data of different scales, high-resolution remote sensing images with sizes of 1.5GB, 5GB, 10GB, and 23GB were selected as test samples. The test focuses include mask quality evaluation, target elimination quality evaluation, and analysis of processing efficiency and resource consumption. By testing different data volumes and complexity scenarios, the performance of the algorithm in all aspects can be comprehensively understood, providing key support for the feasibility in practical applications and data basis for subsequent optimization directions.

[0090] In some embodiments, the stitched image is evaluated using evaluation metrics; the evaluation metrics include edge deviation, Hausdorff distance, and IoU; wherein, the edge deviation and Hausdorff distance are used to evaluate the matching degree between the mask edge and the true target boundary, and the IoU is used to evaluate the regional overlap consistency.

[0091] Edge detail accuracy mainly focuses on the matching degree between the mask edge and the true target boundary, that is, the boundary accuracy of the mask. This metric is particularly important because the deviation of the mask edge may lead to an obvious transition between the filled content and the background, thus affecting the naturalness of the elimination. Specifically, the edge deviation and Hausdorff distance can be used to quantify the difference between the mask boundary and the true target boundary.

[0092] The edge deviation is calculated by finding the minimum distance from each true edge point to the mask edge and taking the average of these minimum distances, so as to obtain the average deviation between the mask and the true target edge. The smaller the edge deviation value, the more consistent the mask edge is with the true edge. Let the true target edge be , and the generated mask edge be , the edge deviation is , and the formula is as follows:

[0093] ;

[0094] In the formula, is the total number of true edge pixels; and respectively represent the pixel positions of the true edge and the mask edge; represents and the Euclidean distance between; Dedge represents the edge deviation; represents the true target edge.

[0095] The Hausdorff distance is used to measure the maximum deviation between two boundaries, which is a measure of the farthest distance in two point sets, defined as the larger of the maximum of all minimum distances from the true edge to the mask edge and the maximum of all minimum distances from the mask edge to the true edge. The formula is as follows:

[0096] ;

[0097] In the formula, sup represents the maximum distance in the specified edge set; inf represents the nearest distance to the specified edge set. represents the true target edge and the generated mask edge of the Hausdorff distance;

[0098] The region overlap consistency can be measured by IoU (Intersection over Union). IoU represents the degree of overlap between the predicted mask region and the true target region, and is defined as the ratio of the intersection area to the union area of the two, which is a commonly used standard for evaluating the mask quality in object detection and instance segmentation tasks. The value range of IoU is between 0 and 1. The closer the value is to 1, the higher the coincidence degree between the predicted mask and the true mask, and the better the matching effect. The IoU formula is as follows:

[0099] ;

[0100] In the formula, represents the predicted mask and the true mask of the number of intersection pixels; represents the predicted mask and the true mask of the number of union pixels.

[0101] This embodiment aims to quantitatively evaluate the quality of the generated mask, using edge detail accuracy (including edge deviation and Hausdorff distance) and mean intersection over union (IoU) as evaluation indicators, and systematically comparing and evaluating the quality changes of the mask before and after dilation. The quality of the mask directly affects the naturalness and accuracy of the sensitive target removal effect. Therefore, edge precision and region overlap are two key evaluation criteria.

[0102] This embodiment selects four pictures in the test set as samples, uses the trained YOLOv8 instance segmentation model to perform object detection on the images, generates a preliminary mask, and further generates a dilated mask. To evaluate the edge detail accuracy of the mask, we compare the generated mask edge with the true target edge manually annotated, and calculate its edge deviation and Hausdorff distance.

[0103] When calculating the edge deviation, according to formula (5), for each real edge point, calculate its minimum distance to the mask edge and find the average value. The Hausdorff distance is calculated according to formula (6) to obtain the maximum deviation distance between the mask edge and the real target edge. For the intersection over union (IoU), according to formula (7), calculate the overlapping degree between the mask area and the real target area to obtain the IoU value of each sample, as shown in Table 1.

[0104] Table 1 Experimental results table of mask quality assessment

[0105]

[0106] The experimental results show that the mask after adaptive dilation is significantly superior to the originally generated mask in terms of the accuracy of edge details. Specifically, both the edge deviation and the Hausdorff distance of the dilated mask are smaller, indicating a higher matching degree between the edge of the dilated mask and the real target edge. In addition, the average intersection over union (IoU) of the dilated mask shows a higher coverage rate on all targets compared to the original mask, indicating that the dilated mask can better cover the real target area. The above results show that mask dilation significantly improves the quality of the mask, providing a more accurate mask basis for subsequent target elimination operations, thus effectively reducing the risk of poor repair effects caused by mask quality problems.

[0107] In this embodiment, from the perspective of the visual effect of target elimination, the quality of the repaired images before and after mask dilation is compared for qualitative evaluation. By observing the fusion effect between the repaired area and the surrounding background, the quality of the target elimination operation is evaluated. In the experiment, first, target detection is performed on the original image to obtain the target area, and the target area is covered with a mask. Then, the original mask is adaptively dilated to obtain a more accurate mask. The original mask and the dilated mask are used as inputs respectively, and after passing through the AOT-GAN model for target elimination, the repaired images are generated and compared, as Figure 4 shown.

[0108] The experimental results show that when using the dilated mask as the input, the visualization effect of the target elimination task is significantly better than that when using the original mask. The elimination result under the input of the dilated mask shows higher coherence and naturalness visually, and the transition between the repaired area and the background area is smoother and there are no obvious splicing traces. In addition, while maintaining a high coverage rate, the dilated mask effectively avoids the unnatural transition phenomenon that may occur during the elimination process, further improving the quality of image repair.

[0109] This embodiment also evaluates the processing efficiency of each image, including the total time required for operations such as mask generation, target repair, and stitching. By comparing the processing speed, memory consumption, and CPU usage of images of different scales (1.5GB, 5GB, 10GB, 23GB), the impact of different image sizes and tile sizes on processing efficiency and resource occupancy is measured.

[0110] The device parameters used in this embodiment are the Windows 10 system, equipped with an i7-13700 CPU and 32G of running memory. The processing efficiency and resource consumption of images of different sizes are as Figure 5 and Figure 6 shown. Figure 5 and Figure 6 In the line charts shown in, the blue line represents the CPU occupancy rate (unit: %), and the orange line is the memory occupancy number (unit: MB). The horizontal axis is the processing time (unit: seconds), the left vertical axis is the CPU occupancy rate, and the right vertical axis is the memory occupancy. Among them, the CPU occupancy rate exceeding 100% is due to the device having a multi-core architecture (the full load of a single core is 100%). Figure 5 and Figure 6 The stage division in is as follows: Stage A is the image loading process; the section with an occupancy rate close to 0 between A and B is the manual inspection stage; Stage B is the target detection process; Stage C is the target elimination process; Stage D is the image saving process.

[0111] This embodiment evaluates the performance of remote sensing images of different sizes in terms of processing efficiency and resource consumption, aiming to verify the actual performance of the proposed algorithm in various scenarios. From Figure 5 and Figure 6 It can be seen that in terms of processing efficiency, when the overall size of the image is fixed, the loading time increases with the increase of the tile size, and the time of other processing stages remains unchanged; when the tile size is fixed, the output time increases with the increase of the image size, and the time of other stages is also not affected. In terms of resource consumption, the peak value of the CPU occupancy rate is stable at about 600% and is not affected by the image size and tile size; the peak value of the memory occupancy rate increases with the increase of the tile size but has nothing to do with the overall size of the image. The experimental results show that the processing efficiency and resource utilization rate of the two core stages of target recognition and target elimination are not affected by the overall size or tile size of the image. The affected ones are mainly the loading and output stages. This method shows stable processing efficiency and resource consumption in different application scenarios, demonstrating good performance and resource adaptability, and can maintain the efficient use of computing resources while ensuring high processing accuracy.

[0112] In summary, the method for eliminating sensitive targets in remote sensing images proposed according to the embodiments of the present application can efficiently and accurately detect and remove sensitive information in high-resolution remote sensing images to meet the requirements of privacy protection and information security. Through the combination of modules such as image block preprocessing, sensitive target recognition, mask correction, and automatic filling and repair, the algorithm reduces the computational cost while improving the detection efficiency, achieving high-precision recognition and natural elimination of sensitive targets in large-size remote sensing images.

[0113] In specific implementation, the YOLOv8 instance segmentation model is used to automatically detect and mark sensitive target areas. At the same time, combined with the mask adaptive dilation and manual correction functions, the mask coverage effect is further optimized to ensure the accurate annotation of sensitive areas. Subsequently, through the context aggregation ability of the AOT-GAN model, automatic repair of the target area is realized, and filling content seamlessly integrated with the surrounding environment is generated, thereby effectively improving the quality and visual consistency of the repaired image.

[0114] Experimental results show that the method proposed in this embodiment exhibits excellent detection and elimination performance in complex scenarios, can effectively reduce phenomena such as residues and unnatural edge transitions in target elimination, and ensure the integrity and authenticity of the image. This algorithm provides a reliable technical solution for the protection of sensitive information in high-resolution remote sensing images and is applicable to a variety of practical application scenarios.

[0115] The embodiments of the present application also provide a device for eliminating sensitive targets in remote sensing images, as Figure 7 shown. The device for eliminating sensitive targets in remote sensing images includes:

[0116] An image block module 701, configured to perform block processing on a high-resolution large-size image to obtain an area of interest containing sensitive targets, and convert the geographical coordinates of the area of interest into pixel coordinates;

[0117] A detection and correction module 702, configured to perform target detection on each target area based on the target area using a trained instance segmentation model, generate a corresponding segmentation mask, and dynamically adjust the dilation degree of the segmentation mask according to the target area to determine the mask area; wherein, the target area is an area containing sensitive targets manually drawn in the area of interest;

[0118] An elimination and stitching module 703, configured to generate filling content and fill it into the mask area to obtain an image after target elimination, and stitch the image after target elimination to the original image to obtain a stitched image.

[0119] In some embodiments, the image block module is further configured to convert the geographical coordinates of the area of interest into pixel coordinates through the following formula:

[0120] ;

[0121] In the formula, x and y respectively represent the abscissa and ordinate of the image pixel; X and Y respectively represent the geographical abscissa and geographical ordinate of the region of interest; a , b , c , d , e and f all represent affine transformation parameters.

[0122] In some embodiments, the trained instance segmentation model.

[0123] In some embodiments, the calculation process of determining the mask region by dynamically adjusting the dilation degree of the segmentation mask according to the target area is expressed as:

[0124] ;

[0125] In the formula, is the dilation kernel size; area is the mask area; 500 is the scaling factor used to adjust the dilation degree; max is the maximum value function; min is the minimum value function; int is the nearest distance to the specified edge set.

[0126] In some embodiments, the splicing elimination module is further configured to:

[0127] Generate filling content through the following formula and fill it into the mask region to obtain the image after target elimination:

[0128] ;

[0129] In the formula, is the filling content generated by the generator, is the original image; is the image after target elimination; is the mask region;

[0130] Splice the image after target elimination to the original image through the following formula to obtain the spliced image:

[0131] ;

[0132] In the formula, is the spliced image; is the mask of the segmentation region; is the number of the segmentation region.

[0133] In some embodiments, the apparatus further comprises an evaluation module, wherein the evaluation module is configured to:

[0134] The spliced ​​image is evaluated using evaluation indicators; the evaluation indicators include edge deviation, Hausdorff distance and IoU; wherein the edge deviation and Hausdorff distance are used to evaluate the matching degree between the mask edge and the real target boundary, and the IoU is used to evaluate the consistency of regional overlap;

[0135] The edge deviation is calculated by the following formula:

[0136] ;

[0137] In the formula, is the total number of true edge pixels; and Represent the pixel positions of the true edge and mask edge respectively; express and The Euclidean distance between them; Dedge represents the edge deviation; Indicates the true target edge;

[0138] The Hausdorff distance is calculated using the following formula:

[0139] ;

[0140] In the formula, sup represents the maximum distance in the specified edge set; inf represents the shortest distance to the specified edge set. Represents the true target edge Generate mask edges Hausdorff distance;

[0141] IoU is calculated using the following formula:

[0142] ;

[0143] In the formula, Represents the prediction mask and the true mask The number of intersection pixels; Represents the prediction mask and the true mask The number of pixels in the union.

[0144] An embodiment of the present application provides an electronic device, which may include: a processor and a memory, wherein the processor and the memory may communicate with each other; illustratively, the processor and the memory communicate with each other via a communication bus.

[0145] The processor executes computer-executable instructions stored in the memory, causing the processor to execute the solutions in the above embodiments. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a network processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0146] The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. The transceiver is used to implement communication between the database access device and other computers (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory.

[0147] The electronic device provided in the embodiments of the present application can be the terminal device in the above embodiments.

[0148] The embodiments of the present application also provide a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are run on a computer, the computer is caused to execute the technical solutions of the remote sensing image sensitive target elimination method in the above embodiments.

[0149] The embodiments of the present application also provide a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and when at least one processor executes the computer program, the technical solutions of the remote sensing image sensitive target elimination method in the above embodiments can be implemented.

[0150] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0151] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0152] In addition, each functional module in various embodiments of the present application can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware, or in the form of a combination of hardware and software functional units.

[0153] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in various embodiments of the present application.

[0154] It should be understood that the above processor can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0155] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.

[0156] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0157] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0158] An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic control unit or a master control device.

[0159] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks or optical discs that can store program codes.

[0160] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for eliminating sensitive targets in remote sensing images, characterized in that: The method comprises: The high-resolution large-size image is processed in blocks to obtain an area of ​​interest containing sensitive targets, and the geographic coordinates of the area of ​​interest are converted into pixel coordinates; Based on the target area, the trained instance segmentation model is used to perform target detection on each target area, and a corresponding segmentation mask is generated. The expansion degree of the segmentation mask is dynamically adjusted according to the target area to determine the mask area; wherein the target area is an area containing sensitive targets manually drawn in the region of interest; Generate a filling content and fill it into the mask area to obtain an image after the target is eliminated, and splice the image after the target is eliminated to the original image to obtain a spliced ​​image, including: The filling content is generated by the following formula and filled into the mask area to obtain the image after the target is eliminated: ; In the formula, The filling content generated by the generator, is the original image; The image after the target is eliminated; is the mask area; The image after the target is eliminated is spliced ​​to the original image using the following formula to obtain a spliced ​​image: ; In the formula, is the stitched image; is the mask of the segmented region; The number of the partition area.

2. The remote sensing image sensitive target elimination method according to claim 1, characterized in that: The geographic coordinates of the area of ​​interest are converted to pixel coordinates using the following formula: ; In the formula, x and y Represent the horizontal and vertical coordinates of the image pixels respectively; X and Y Respectively represent the geographic abscissa and ordinate of the area of ​​interest; a , b , c , d , e and f Both represent affine transformation parameters.

3. The remote sensing image sensitive target elimination method according to claim 1, characterized in that: The expansion degree of the segmentation mask is dynamically adjusted according to the target area, and the calculation process of determining the mask area is expressed as: ; In the formula, is the expanded core size; area is the mask area; 500 is the scaling factor used to adjust the degree of expansion; max is the maximum function; min is the minimum function; int is the closest distance to the specified edge set.

4. The remote sensing image sensitive target elimination method according to claim 1, characterized in that: After obtaining the spliced ​​image, the method further includes: The spliced ​​image is evaluated using evaluation indicators; the evaluation indicators include edge deviation, Hausdorff distance and IoU; wherein the edge deviation and Hausdorff distance are used to evaluate the matching degree between the mask edge and the real target boundary, and the IoU is used to evaluate the consistency of regional overlap; The edge deviation is calculated by the following formula: ; In the formula, is the total number of true edge pixels; and Represent the pixel positions of the true edge and mask edge respectively; express and The Euclidean distance between them; Dedge represents the edge deviation; Indicates the true target edge; The Hausdorff distance is calculated using the following formula: ; In the formula, sup represents the maximum distance in the specified edge set; inf represents the shortest distance to the specified edge set. Represents the true target edge Generate mask edges Hausdorff distance; IoU is calculated using the following formula: ; In the formula, Represents the prediction mask and the true mask The number of intersection pixels; Represents the prediction mask and the true mask The number of pixels in the union.

5. A remote sensing image sensitive target elimination device, characterized in that: The device comprises: An image segmentation module is configured to segment high-resolution large-size images to obtain an area of ​​interest containing sensitive targets, and convert geographic coordinates of the area of ​​interest into pixel coordinates; The detection and correction module is configured to perform target detection on each target area based on the target area using the trained instance segmentation model, generate a corresponding segmentation mask, dynamically adjust the expansion degree of the segmentation mask according to the target area, and determine the mask area; wherein the target area is an area containing sensitive targets manually drawn in the region of interest; The elimination and splicing module is configured to generate a filling content and fill it into the mask area to obtain an image after the target is eliminated, and splice the image after the target is eliminated to the original image to obtain a spliced ​​image; The de-joining module is further configured to: The filling content is generated by the following formula and filled into the mask area to obtain the image after the target is eliminated: ; In the formula, The filling content generated by the generator, is the original image; The image after the target is eliminated; is the mask area; The image after the target is eliminated is spliced ​​to the original image using the following formula to obtain a spliced ​​image: ; In the formula, is the stitched image; is the mask of the segmented region; The number of the partition area.

6. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the remote sensing image sensitive target elimination method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the remote sensing image sensitive target elimination method according to any one of claims 1 to 4.

8. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for eliminating sensitive targets in remote sensing images as claimed in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Quick high-resolution remote sensing image segmentation method

    CN105574887A

  • Method for removing cloud and cloud shadows in remote sensing image

    CN111899194A