Remote sensing target detection method based on area positioning
Through the remote sensing object detection method based on region positioning, combined with high-resolution feature extraction and adaptive clustering area mining technology, the accuracy and efficiency problems of remote sensing small object detection in complex backgrounds are solved, and more efficient target recognition and positioning are achieved.
Patent Information
- Application Number
- CN202510299049.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
AI Technical Summary
Remote sensing small target detection is difficult to accurately identify and locate under complex backgrounds. Traditional methods have problems such as low detection accuracy, high computational complexity, and serious resource waste, especially in dense target scenarios, overlap between targets and serious background interference.
Using a remote sensing object detection method based on region positioning, combined with high-resolution feature extraction, target area positioning map generation and adaptive cluster area mining technology, the positioning map is generated and cluster area mining is performed through the CenterNet object detector with HRNet as the backbone network, and the loss function is optimized to improve detection accuracy and efficiency.
The accuracy and efficiency of remote sensing small object detection is improved, background noise interference is reduced, and different target densities and scales are adapted to, and the robustness and computing efficiency of the detection algorithm are improved.
Smart Images

Figure CN120259876A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to a remote sensing target detection method based on regional positioning. Background Art
[0002] With the continuous development of remote sensing technology, remote sensing images are increasingly widely used in many fields such as urban planning, disaster response, military defense, and environmental monitoring. In these applications, remote sensing small target detection, as a key technology, plays a crucial role. Remote sensing small target detection not only accurately identifies small targets in images but also precisely locates them in complex backgrounds, so it has important practical significance in aspects such as monitoring, resource management, and emergency response. However, remote sensing small target detection faces a series of technical challenges, mainly reflected in the following aspects:
[0003] First, the information of small targets is scarce. Small targets in remote sensing images usually occupy fewer pixels in the overall image, and their appearances are often blurred and difficult to distinguish from the background area. Due to the small area and unobvious features of small targets, traditional pixel-based feature extraction methods are difficult to accurately capture their details, resulting in low detection accuracy of small targets. In addition, small targets in remote sensing images often have strong similarity with background noise, which makes traditional methods vulnerable to background interference, thus affecting the target recognition effect.
[0004] Second, the target distribution is uneven. In high-resolution remote sensing images, targets are usually sparse and concentrated in certain areas, while the background area occupies most of the image. This distribution characteristic poses challenges to traditional region proposal-based methods. Especially in dense target scenarios, the distance between targets is small, which easily leads to overlapping of target detection results. In addition, traditional target detection methods often cannot effectively distinguish dense areas from the background, resulting in low detection efficiency and waste of computing resources.
[0005] To address these problems, traditional image processing methods such as cropping and blocking techniques are widely used. These methods process the image by dividing it into small blocks, but in many cases, this method will result in a large amount of background information being processed ineffectively, thus increasing the computational complexity and resource consumption. Furthermore, as the image scale increases, traditional detection methods often perform inaccurately when dealing with small targets, and the detection effect is poor.
[0006] In order to solve these problems, some methods based on density maps or region proposals have been proposed in recent years. These methods assist detection by generating density maps of targets, but in the case of dense small targets and complex backgrounds, the generation of density maps still has ambiguity and position deviation problems, resulting in reduced target positioning accuracy, which in turn affects the detection effect. In addition, these methods often rely on a large amount of computing resources, are inefficient, and cannot adaptively handle targets of different sizes. Therefore, how to improve detection accuracy and reduce computational complexity has become a key issue that needs to be urgently addressed in remote sensing small target detection. Summary of the invention
[0007] In order to solve the problems of insufficient large-format image resources, diverse target scales, uneven data distribution and overlapping dense targets in remote sensing aerial image target detection tasks, the present invention proposes a remote sensing image target detection method based on regional positioning. The method integrates high-resolution feature extraction, target region positioning map generation and adaptive clustering region mining technology, aiming to more accurately describe the target position and improve the detection performance of dense small targets.
[0008] In order to achieve the above object, the technical solution of the present invention is achieved as follows:
[0009] A remote sensing target detection method based on regional positioning, the steps are as follows:
[0010] Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing images in the dataset;
[0011] Step 2: Build a CenterNet target detector with HRNet as the backbone network and train the CenterNet target detector;
[0012] Step 3: Input the input image into the trained CenterNet object detector to obtain the rough detection result;
[0013] Step 4: Input the input image to the localization map generation module to generate localization maps of each category; merge the localization maps of each category and input them to the clustering region mining module to find the dense area, crop the dense area, and splice and enlarge it to obtain a spliced image; input the spliced image to the trained CenterNet target detector to obtain a fine detection result;
[0014] Step 5: Input the coarse detection results and the fine detection results into the refinement-coarse merging module to merge them to obtain the final result.
[0015] Preferably, the method for preprocessing the remote sensing images in the dataset is as follows: Each remote sensing image is cropped into image patches of a unified size, and data augmentation processing is performed on each cropped remote sensing image, including random scaling, random cropping, and random arrangement and splicing, to generate a diverse remote sensing image dataset.
[0016] Preferably, the method for inputting the input image into the localization map generation module to generate localization maps for each category is as follows:
[0017] Calculate the average distance D between targets i :
[0018]
[0019] where d(C i , C j ) is the spatial distance between target i and target j, C i = (x i , y i ) is the center coordinate of target i, and k is the number of neighbors;
[0020] Perform normalization on D i :
[0021]
[0022] where is the normalized average distance between targets, and max(D) is the maximum value of the distances between all targets;
[0023] Construct an inverse square distance transformation function based on the density of targets:
[0024]
[0025] where G(x, y) is the response value of target i at the image coordinates (x, y), D(x, y) is the Euclidean distance from the pixel point (x, y) in the image to the target center, and γ is the attenuation parameter;
[0026] The attenuation parameter γ is adjusted according to the density of targets :
[0027]
[0028] where α is the adjustment factor, max_γ is the maximum attenuation parameter; γ i represents the attenuation parameter under different densities; is the normalized average distance between targets;
[0029] The precise calibration of the target position is achieved by calculating the center coordinates of each target and using the distance transformation method to calculate the response value of the target; for each target i, its center coordinates are:
[0030]
[0031] where (x min , y min ) are the coordinates of the upper left corner of the target bounding box, and (x max , y max ) are the coordinates of the lower right corner of the target bounding box;
[0032] The distance transformation method is used to generate the target response map. For each pixel point (x, y) in the image, the distance d(x, y) to the target center C i is calculated, and then the response value of the target is generated according to the Gaussian decay formula;
[0033] A threshold θ of the response value is introduced. After calculating the response value of each target, if the response value is less than the set threshold θ, it is set to zero;
[0034] By generating a multi-channel localization map, where each channel corresponds to a target of a different category; for the target of category c, its corresponding response map G c is:
[0035]
[0036] where Γ c represents the set of targets of category c, and G i (x, y) is the response value of target i at the image position (x, y).
[0037] Preferably, the method for processing the localization map by the clustering region mining module to obtain the stitched image includes:
[0038] Extracting the target region from the response map:
[0039] Given a target response map Extract the maximum response value at each pixel position:
[0040]
[0041] where R c,h,w represents the response value at the position (h, w) on channel c;
[0042] Adaptive grid division:
[0043] Assume the image size is W×H and the number of targets is N t ; According to the number of targets in each image, dynamically adjust the grid width w sand the grid height h s Specifically, if the target number is greater than 60, it is considered dense, and the number of grids is adjusted to 64×40. The number of grids M is defined as follows according to the number of object instances c: Under this rule, the image grid division is adjusted according to the target density;
[0044] Density map analysis and target clustering: Calculate the density map based on the sum of target pixels in each grid:
[0045]
[0046] where D p,q represents the density value in the (p,q) grid, and I(x,y) is the pixel value at the position (x,y) in the image; First, generate a density map by calculating the density of each grid; Then, based on the density information in the density map, select the top k grid regions with the highest density for target clustering.
[0047] Preferably, the specific implementation method of the density map analysis and target clustering is as follows:
[0048] Calculate the density map:
[0049]
[0050] where binmap[x,y] is the pixel value in the binary image, representing the presence of target pixels;
[0051] Region screening and target clustering: Sort each grid of the density map and select the top k regions with higher density for target clustering; The following formula is used to select the top k regions:
[0052] topk_idx = argsortR[-k:];
[0053] By sorting the density of the regions, filter out the regions with the most concentrated density and the most targets;
[0054] Region coordinate adjustment and visualization: Adjust the upper-left and lower-right coordinates of each clustering region:
[0055]
[0056] Connect adjacent regions in the density map into a connected region, and the eight-connected region labeling method is carried out through the following recursive method:
[0057] top = min(top,x curr ), bottom = max(bottom,x curr ), left = min(left,ycurr ),
[0058] right = max(right, y curr );
[0059] Region coordinate adjustment: An adaptive coordinate adjustment mechanism is adopted to correct the offset of the coordinates coord i = (x i , y i , w i , h i ) of the detected target region; Specifically, the boundary information of the image is extracted from the image metadata imag_metas, denoted as bord_pixs. This information includes the cropping or scaling information of the image, in the form of border_pixs = (x border , y border , w border , h border ), where x border , y border are the upper left coordinates, and w border , h border are the width and height; For the coordinates of each dense region, the following correction formula is used for offset correction:
[0060]
[0061] Target region cropping: After obtaining the coordinates of the dense regions, the image is divided into multiple sub-regions; The size of each sub-region is dynamically adjusted according to the density of the target. The image cropping and annotation update process is completed through the following steps:
[0062] box i = (x i , y i , w i , h i );
[0063] where (x i , y i , w i , h i ) are the coordinates and dimensions of the i-th target region;
[0064] Each target region determines whether to be selected for cropping by calculating its area A i : A i = w i × h i ; Sort by area and select the top t largest regions for further cropping; t = 1, 2, 3;
[0065] Image Cropping and Scaling: For each selected target area, crop out the sub-image containing the target and perform scaling; the specific cropping area is calculated by the following formula:
[0066]
[0067] where WW and HH are the width and height of the original image respectively;
[0068] Stitching of Target Areas: Stitch multiple cropped image patches. The stitching method is to arrange the multiple image patches horizontally or vertically to form a new image; during the stitching process, keep the relative positions and sizes of the targets consistent; assuming the number of cropped image patches is M, the width and height of the final stitched image are:
[0069]
[0070] If the size of the stitched image exceeds the preset maximum width or height, scale it proportionally while keeping the proportion of the image content unchanged;
[0071] Image Padding and Saving: There may be some blank areas in the stitched image, and the image size needs to be adjusted to the specified size through padding operations; the padding operation is carried out by the following formula:
[0072]
[0073] Fill the blank areas with zeros to obtain the final stitched mosaic image.
[0074] Preferably, the loss function adopted during the training of the CenterNet target detector includes focal loss, normalized Wasserstein distance loss, and L1 loss. The final loss function is defined as:
[0075] L det = L k + λ nwd · L nwd + λ l1 · L1;
[0076]
[0077]
[0078] where λ nwd 、λ l1 are both coefficients; L k is the focal loss, p i is the prediction probability, α is the balance factor, and μ is the hyperparameter that adjusts the influence degree of easy and difficult samples; L nwd is the normalized Wasserstein distance loss, W(Bpred , B gt ) represents the predicted bounding box B pred and the ground truth bounding box B gt The Wasserstein distance between them, w gt and h gt are the width and height of the ground truth bounding box respectively; L1 is the L1 loss, and represent the predicted coordinates and the ground truth coordinates respectively.
[0079] Advantages of the present invention:
[0080] 1) The present invention proposes an adaptive adjustment mechanism based on the density between objects: When generating the object localization map, considering the relative positions between objects, by calculating the average distance between objects and normalizing it, the attenuation intensity of each object is dynamically adjusted. This method can ensure that objects in dense regions obtain smaller attenuation parameters, and objects in sparse regions obtain larger attenuation parameters.
[0081] 2) The present invention introduces a threshold for the response value. After calculating the response value of each object, if the response value is less than the set threshold, it is set to zero, thereby filtering out invalid response regions. This mechanism effectively reduces the interference of background noise and improves the robustness of the object detection algorithm.
[0082] 3) By generating a multi-channel localization map, the present invention can process objects of different categories simultaneously, and the localization maps of each category are independent in space, avoiding interference between categories.
[0083] 4) The present invention uses the NWD loss combined with the 2D Gaussian function to measure the similarity between object bounding boxes. By adjusting the weight coefficient of NWD, the detection performance can be optimized for objects of different scales. For medium and large objects, by combining the NWD loss with the focal loss, the accuracy of object detection is further improved. Description of the Drawings
[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0085] Figure 1 is the network structure flowchart of the present invention.
[0086] Figure 2 is the network architecture diagram of CenterNet.
[0087] Figure 3 It is the network architecture diagram of the improved CenterNet.
[0088] Figure 4 It is the structural diagram of the clustering region mining module.
[0089] Figure 5 It is the result diagram of remote sensing target detection. Specific implementation manners
[0090] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0091] As Figure 1 shown, the embodiments of the present invention provide a remote sensing target detection method based on region positioning, and the specific steps are as follows:
[0092] Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing images in the dataset; download multiple remote sensing images and crop each remote sensing image into image patches of a unified size. Then, perform data augmentation processing on each cropped remote sensing image, including random scaling, random cropping, and random arrangement and splicing, etc., to generate a diversified remote sensing image dataset. These processes help to increase the diversity of training data and improve the generalization ability of the model for different target sizes and positions.
[0093] Step 2: Construct a CenterNet target detector with HRNet as the backbone network and train the CenterNet target detector; specifically, HRNet contains multiple stages, and each stage consists of several residual blocks. By inputting low-resolution feature maps and through repeated information exchange between multi-resolution parallel subnets, HRNet can always maintain high-resolution representations and generate high-quality feature maps. These feature maps are fed into the center point detection module of CenterNet to accurately locate the target center and regress the size of the bounding box. As Figure 2 shown, the present invention adopts an architecture based on CenterNet in the target detection network to achieve target detection by regressing information such as the center point, size, direction, and pose of the target. To better adapt to the characteristics of remote sensing images, the present invention optimizes the backbone network of CenterNet, selects HRNet as the backbone network to enhance the feature extraction ability. In addition, the density map is replaced with a positioning map, and the loss function is improved, and the normalized Wasserstein distance (NWD) loss is adopted. As Figure 3As shown. These optimizations help improve the accuracy of the target position, avoid the blurring and position overlap problems in traditional Gaussian kernel density maps, and thus significantly enhance the detection effect.
[0094] During the training process, an improved focal loss was adopted to address the problem of class imbalance in object detection. For bounding box regression, the normalized Wasserstein distance (NWD) was used as a new loss function to solve the limitations of the traditional L1 loss in small object regression. The NWD loss can better measure the distance between small objects and the ground truth bounding boxes, thereby improving the detection accuracy. Specifically, the NWD loss combines a 2D Gaussian function to measure the similarity between object bounding boxes. By adjusting the weight coefficient of NWD, the detection performance can be optimized for objects of different scales. For medium and large objects, the detection accuracy of object detection is further improved by combining the NWD loss with the focal loss.
[0095] The final loss function is defined as:
[0096] L det =L k +λ nwd ·L nwd +λ l1 ·L1;
[0097]
[0098] Where λ nwd 、λ l1 are both coefficients. L k is the focal loss, p i is the predicted probability, α is the balance factor, and μ is the hyperparameter that adjusts the influence degree of easy and hard samples; L nwd is the normalized Wasserstein distance loss, W(B pred ,B gt ) represents the Wasserstein distance between the predicted bounding box B pred and the ground truth bounding box B gt , w gt and h gt are the width and height of the ground truth bounding box respectively; L1 is the L1 loss, and represent the predicted coordinates and the ground truth coordinates respectively. In this embodiment, λ nwd , μ are set to 2, λ l1 is set to 0.5, and α is set to 0.25.
[0099] Step 3: Input the input image into the trained CenterNet object detector to obtain a rough detection result.
[0100] Step 4: Input the input image into the localization map generation module to generate localization maps for each category; merge the localization maps for each category, input them into the clustering region mining module to find dense regions, crop the dense regions, splice and enlarge them to obtain a spliced image; input the spliced image into the trained CenterNet object detector to obtain a refined detection result.
[0101] Localization map generation module: First, define the Euclidean distance change map formula: Among them, S represents the set of marked center points, and D(x, y) represents the distance from the pixel point (x, y) to its nearest target center. The Euclidean distance has a scale invariance problem, which may cause the distances between nearby and far targets to be similar, thus affecting the localization accuracy. In addition, the Euclidean distance performs poorly in high-dimensional spaces, especially when the distance changes greatly, which makes direct regression challenging. Therefore, in this embodiment, the inverse distance transform (IDT) is first used to generate the localization map of the target, and the IDT map generates a response map by calculating the distance from each pixel to the nearest target. The IDT process is expressed as: The IDT map can represent the position corresponding to the local maximum more accurately than the widely used density map. However, ideally, the IDT map decays slowly around the target center (foreground), while the response converges rapidly away from the target center, making the IDT map inapplicable to large-scale remote sensing images. Therefore, the inverse square distance transform map (ISDT) is used to accelerate the response decay, so that the targets in the image can be more accurately located. The process is
[0102] During the localization map generation process, the present invention particularly focuses on the density of the targets. According to the distribution of the targets in the image, the response decay factor γ in the ISDT map is dynamically adjusted to enhance the target response in the dense region and suppress the interference of background noise. Specifically, the value of γ is adaptively adjusted according to the target density change and background complexity. Through the KD-Tree algorithm, the crowding degree between adjacent targets is calculated to evaluate their spatial density. In the aggregation region, γ is increased to enhance the target response, while in the relatively sparse region, γ is decreased to suppress the background interference. The combination of KD-Tree clustering and spatial distance is used to perform clustering analysis on the spatial relationship between targets through the KD-Tree to evaluate the density of the targets. The KD-Tree is an efficient spatial data structure that can quickly find neighboring targets in large-scale data. In the present invention, the KD-Tree is used to calculate the distance between targets and adjust the influence of Gaussian decay through these distances.
[0103] Through this combination, the generation of the positioning map can better adapt to the changes in the target density, improve the detection accuracy, especially in dense scenes. Specifically as follows:
[0104] Calculate the average distance D between targets i :
[0105]
[0106] where d(C i , C j ) is the spatial distance between target i and target j, and C i =(x i , y i ) is the center coordinate of target i, and k is the number of neighbors.
[0107] To fully reflect the density between targets, normalize D i :
[0108]
[0109] where is the average distance between targets after normalization, and max(D) is the value of the maximum distance between all targets. In this way, the density between targets will be quantified, and the attenuation parameter γ will be adjusted according to the density of each target.
[0110] The present invention uses the inverse square distance transformation function to calculate the response value of the target position. For each target, the intensity of the attenuation of its response value is controlled by the adaptively adjusted attenuation parameter γ. Construct the inverse square distance transformation function based on the density of the target:
[0111]
[0112] where G(x, y) is the response value of target i at the image coordinates (x, y), D(x, y) is the Euclidean distance from the pixel point (x, y) in the image to the target center, and γ is the attenuation parameter.
[0113] The attenuation parameter γ is adjusted according to the density between targets :
[0114]
[0115] where α is the adjustment factor, taking α = 10 to control the overall attenuation rate; max_γ is the maximum attenuation parameter; γ i represents the attenuation parameter under different densities; is the average distance between targets after normalization; to prevent excessive attenuation, the present invention sets max_γ = 5. Ensure that the attenuation factor γ iIt will not be too large to avoid the target response being too weak due to excessive attenuation. γ i Controls the response attenuation rate from the target center to other regions. In the target-dense region, γ i is larger, which can strengthen the response in these regions, thereby enhancing the positioning accuracy of targets in the dense region; while in the sparse region, γ i is smaller, thus suppressing unnecessary responses.
[0116] The precise calibration of the target position is achieved by calculating the center coordinates of each target and using the distance transformation method to calculate the response value of the target; for each target i, its center coordinates are:
[0117]
[0118] where, (x min , y min ) is the upper-left coordinate of the target bounding box, and (x max , y max ) is the lower-right coordinate of the target bounding box.
[0119] The distance transformation method is used to generate the target response map. For each pixel point (x, y) in the image, the distance d(x, y) to the target center C i is calculated, and then the response value of the target is generated according to the Gaussian attenuation formula.
[0120] To avoid the influence of low response values, the present invention introduces a threshold θ for the response value. After calculating the response value of each target, if the response value is less than the set threshold θ, it is set to zero, thereby filtering out the invalid response regions. This mechanism effectively reduces the interference of background noise and improves the robustness of the target detection algorithm.
[0121] The present invention generates a multi-channel localization map, where each channel corresponds to a target of a different category; for the target of category c, its corresponding response map G c is:
[0122]
[0123] where, Γ c represents the set of targets of category c, and G i (x, y) is the response value of target i at the image position (x, y).
[0124] In this way, targets of different categories can be processed simultaneously, and the localization maps of each category are independent in space, avoiding interference between categories.
[0125] Such as Figure 4As shown, in the clustering region mining module, an accurate target localization map is obtained through the localization map generation module (LMG). Subsequently, the inverse square distance transform (ISDT) is used to generate a more accurate target position map to enhance the response of the target in the dense region. By superimposing the localization maps of each target category, a full-category localization map is generated, and the localization map is binarized to obtain a binary mask. The binarized image is evenly divided into multiple grids, and K most dense grids are selected as the centers of the clustering regions, and the surrounding grids are merged using the 8-neighborhood method. All regions are enlarged by 1.2 times to prevent objects from being truncated. Finally, small cropped regions are integrated into a mosaic image. The specific steps are as follows:
[0126] Extract the target region from the response map:
[0127] The present invention performs target localization by extracting the maximum response region from the target response map. Specifically, given a target response map Extract the maximum response value at each pixel position:
[0128]
[0129] where R c,h,w represents the response value at position (h, w) on channel c; through this method, the most prominent region of the target can be quickly and accurately extracted.
[0130] Adaptive grid division: Taking a 16×10 grid as a reference and making adaptive adjustments according to the number of objects in each image. Calculate the total number of object instances in the dataset, and then set a threshold to control the number of grids.
[0131] Target quantity statistics and grid adjustment: Assume the image size is W×H and the number of targets is N t ; According to the number of targets in each image, dynamically adjust the grid width w s and the grid height h s ; Specifically, if the number of targets is greater than 60, it is considered dense, and the number of grids is adjusted to 64×40, and the number of grids M is defined according to the number of object instances c as follows: Under this rule, the image grid division is adjusted according to the target density; thus effectively avoiding excessive redundant calculations in the dense or sparse target regions.
[0132] Density map analysis and target clustering: To further optimize the target detection process, the density map is used to analyze the image and judge the target density in each grid. Calculate the density map according to the sum of target pixels in each grid:
[0133]
[0134] Among them, D p,q represents the density value in the (p, q) grid, and I(x, y) is the pixel value at the position (x, y) in the image; first, a density map is generated by calculating the density of each grid; then, based on the density information in the density map, the top k grid regions with the highest density are selected for target clustering. The specific steps are as follows:
[0135] Calculate the density map:
[0136]
[0137] Among them, binmap[x, y] is the pixel value in the binary image, indicating the presence of the target pixel.
[0138] Region screening and target clustering: By sorting each grid in the density map, the top k regions with higher density are selected for target clustering; the following formula is used to select the top k regions:
[0139] topk_idx = argsortR[-k:];
[0140] By sorting the density of the regions, the regions with the most concentrated density and the most targets are screened out.
[0141] Region coordinate adjustment and visualization: To ensure the accurate positioning of the target regions after clustering, the upper left and lower right coordinates of each clustering region are adjusted:
[0142]
[0143] Through these formulas, the coordinates of the clustering regions can be adjusted to appropriate bounding boxes to ensure the high accuracy of the detected target regions.
[0144] After clustering, to improve the accuracy of target segmentation, the eight-connected region labeling algorithm is used to further process the target regions. Through this algorithm, the adjacent regions in the density map are connected into a connected region. The eight-connected region labeling method is carried out through the following recursive method:
[0145] top = min(top, x curr ), bottom = max(bottom, x curr ), left = min(left, y curr ),
[0146] right = max(right, y curr );
[0147] Region Coordinate Adjustment: When processing remote sensing images, the boundary information of the image may cause the target region to shift. Therefore, an adaptive coordinate adjustment mechanism is adopted to correct the coordinates of the detected target region coord i =(x i ,y i ,w i ,h i ); Specifically, the boundary information of the image is extracted from the image metadata imag_metas, denoted as bord_pixs. This information includes the cropping or scaling information of the image, in the form of border_pixs=(x border ,y border ,w border ,h border ), where x border ,y border are the upper left coordinates, and w border ,h border are the width and height; for the coordinates of each dense region, the following correction formula is used for offset correction:
[0148]
[0149] Only x and y are adjusted during the correction process, and the width and height remain unchanged.
[0150] Target Region Cropping: After obtaining the coordinates of the dense regions, the image is divided into multiple sub-regions; the size of each sub-region is dynamically adjusted according to the density of the target. The image cropping and annotation update process is completed through the following steps:
[0151] box i =(x i ,y i ,w i ,h i );
[0152] Among them, (x i ,y i ,w i ,h i ) are the coordinates and dimensions of the i-th target region.
[0153] Each target region decides whether to be selected for cropping by calculating its area A i : A i =w i ×h i ; Sort according to the area and select the top t regions with the largest area for further cropping; t = 1, 2, 3.
[0154] Image Cropping and Scaling: For each selected target region, a sub-image containing the target is cropped and scaled; to avoid the target being cropped off or the boundary being truncated, the cropping region is expanded by a certain ratio (the ratio coefficient is 1.2). The specific cropping region is calculated by the following formula:
[0155]
[0156] where WW and HH are the width and height of the original image respectively;
[0157] Stitching of Target Regions: Multiple cropped image patches are stitched together. The stitching method is to arrange multiple image patches horizontally or vertically to form a new image; during the stitching process, the relative positions and sizes of the targets are kept consistent; assuming the number of cropped image patches is M, the width and height of the final stitched image are:
[0158]
[0159] If the size of the stitched image exceeds the preset maximum width or height (such as 1024×640), it is scaled proportionally while keeping the proportion of the image content unchanged.
[0160] Image Padding and Saving: There may be some blank areas in the stitched image, and the image size needs to be adjusted to a specified size (such as 1024×960) through padding operations; the padding operation is carried out by the following formula:
[0161]
[0162] The blank areas are filled with zeros to obtain the final stitched mosaic image.
[0163] Step Five: The rough detection result and the fine detection result are input into the refinement-rough merging module for merging. Specifically, the fine detection result and the rough detection result of the entire image are merged with the standard non-maximum suppression (NMS) to obtain the final result, that is, the detection boxes are sorted according to the confidence level, the box with the highest confidence level is preferentially retained, and the intersection over union (IoU) with other candidate boxes is calculated. If the IoU exceeds the set threshold, it is regarded as redundant and removed. This process is iterated until all candidate boxes are processed.
[0164] After obtaining the cropped dense region image and the mosaic image, in order to ensure that the input image can match the input requirements of the detector and effectively utilize the feature extraction ability of the network, all images are scaled up by a factor of 1.5 and padded, and finally adjusted to a size of 1024×960. This adjustment helps to improve the image resolution, making the target features more obvious, and at the same time ensuring that the network can handle targets of different sizes, especially in the case of dense target regions and large targets. The processed images are then input into the detector, and through the aforementioned steps, more discriminative small target detection results are obtained.
[0165] To take into account the detection requirements of medium and large targets, the entire image is input into the network to obtain global detection results. By fusing local and global detection information, more accurate detection results are finally obtained. Specifically, first, fine detection results of local regions are obtained through the clustering region mining module, and then these local results are fused with the coarse-grained detection results of the global image. This process effectively improves the detection accuracy, especially in the case of large differences in target sizes, and can meet the detection requirements for both large and small targets.
[0166] To verify the effectiveness of the improved strategy, the present invention further conducts ablation experiments, as shown in Table 1. Among them, the baseline model is set as CenterNet, and on this basis, the proposed improvements are verified respectively. Three widely used evaluation metrics (i.e., AP, AP 50 and AP 75 ) are used to verify the performance. In addition, another three metrics (i.e., AP s , AP m and AP l ) are used to measure the performance of different target scales.
[0167] Table 1 Ablation Experiments
[0168]
[0169] Table 2 shows the result comparison between the method of the present invention and other methods. These methods include: ClusDet, DMNet, CDMNet, CEASC, AMRNet, GLSAN, PRDet, HRDNet, QueryDet, CRENet. It can be found from the table that the method of the present invention shows the best results in remote sensing target detection. In addition, the frames per second (FPS) is also reported to measure the efficiency of the proposed model. At the same time, it also shows that the method of the present invention has good generalization ability.
[0170] Table 2 Comparative Experiments
[0171]
[0172] In addition, the present invention also conducts a visualization experiment on the detection results. The visualization results are as Figure 5 shown. It can be seen that, compared with the baseline method, the model proposed by the present invention shows better detection performance for complex scenarios with a large number of tiny targets and uneven data distribution, especially for small targets. At the same time, the inference speed is relatively fast (s / img), showing better ability to obtain location information.
[0173] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A remote sensing target detection method based on regional positioning, characterized in that, The steps are as follows: Step 1: Obtain a remote sensing image dataset and preprocess the remote sensing images in the dataset; Step 2: Construct a CenterNet object detector with HRNet as the backbone network and train the CenterNet object detector; Step 3: Input the input image into the trained CenterNet object detector to obtain a rough detection result; Step 4: Input the input image into the localization map generation module to generate localization maps for each category; Merge the localization maps for each category, input them into the clustering region mining module to find dense regions, crop the dense regions, splice and enlarge them to obtain a spliced image; input the spliced image into the trained CenterNet object detector to obtain a refined detection result; Step 5: Input the rough detection result and the refined detection result into the refinement-rough merging module for merging to obtain the final result.
2. The remote sensing target detection method based on regional positioning according to claim 1, wherein The method for preprocessing the remote sensing images in the dataset is: Crop each remote sensing image into image patches of a unified size, and perform data augmentation processing on each cropped remote sensing image, including random scaling, random cropping, and random arrangement and splicing methods, to generate a diversified remote sensing image dataset.
3. The remote sensing target detection method based on regional positioning according to claim 1, characterized in that The method for inputting the input image into the localization map generation module to generate localization maps for each category is: Calculate the average distance D between the targets i : where d(C i , C j ) is the spatial distance between target i and target j, C i = (x i , y i ) is the central coordinate of target i, and k is the number of neighbors; Normalize D i as follows: Among them, is the average distance between targets after normalization, and max(D) is the value of the maximum distance among all targets; Construct an inverse square distance transformation function based on the density of the target: where G(x,y) is the response value of target i at the image coordinates (x,y), D(x,y) is the Euclidean distance from the pixel point (x,y) in the image to the target center, and γ is the attenuation parameter; The attenuation parameter γ is adjusted according to the density between targets as follows: where α is an adjustment factor, max_γ is the maximum attenuation parameter; γ i represents the attenuation parameter at different densities; is the average distance between normalized targets; The accurate calibration of the target position is achieved by calculating the center coordinates of each target and using the distance transformation method to calculate the response value of the target; for each target i, its center coordinates are: Among them, (x min , y min ) is the upper left coordinate of the target bounding box, and (x max , y max ) is the lower right coordinate of the target bounding box; Use the distance transformation method to generate the target response map. For each pixel point (x, y) in the image, calculate the distance d(x, y) to the target center C i and then generate the response value of the target according to the Gaussian decay formula; Introduce a threshold θ for the response value. After calculating the response value of each target, if the response value is less than the set threshold θ, set it to zero; By generating a multi-channel localization map, where each channel corresponds to a target of a different category; for a target of category c, its corresponding response map G c is as follows: Among them, Γ c represents the target set of category c, and G i (x, y) is the response value of target i at the image position (x, y).
4. The remote sensing target detection method based on regional positioning according to claim 3, characterized in that The method for processing the localization map through the clustering region mining module to obtain a spliced image includes: Extract the target region from the response map: Given a target response map Extract the maximum response value at each pixel position: Among them, R c,h,w represents the response value at the position (h, w) on channel c; Adaptive grid division: Assume the image size is W×H and the number of targets is N t ; Dynamically adjust the grid width w according to the number of targets in each image s and the grid height h s ; Specifically, if the number of targets is greater than 60, consider it dense and adjust the number of grids to 64×40, and define the number of grids M according to the number of object instances c as follows: Under this rule, the image grid division is adjusted according to the target density; Density map analysis and target clustering: Calculate the density map according to the total number of target pixels in each grid: Among them, D p,q represents the density value in the (p, q) grid, and I(x, y) is the pixel value at the position (x, y) in the image. First, a density map is generated by calculating the density of each grid. Then, based on the density information in the density map, the top k grid regions with the highest density are selected for target clustering.
5. The remote sensing target detection method based on regional positioning according to claim 4, wherein The specific implementation method of the density map analysis and target clustering is: Calculate the density map: where binmap[x,y] is the pixel value in the binary image, indicating the presence of target pixels; Region screening and target clustering: By sorting each grid of the density map, select the top k regions with higher density for target clustering; use the following formula to select the top k regions: topk_idx = argsortR[-k:]; By sorting the density of the regions, screen out the regions with the most concentrated density and the most targets; Region coordinate adjustment and visualization: Adjust the upper left and lower right coordinates of each clustering region: Connect adjacent regions in the density map into a connected region. The eight-connected region labeling method is carried out through the following recursive method: top = min(top, x curr ), bottom = max(bottom, x curr ), left = min(left, y curr ), right = max(right, y curr ); Region coordinate adjustment: Adopt an adaptive coordinate adjustment mechanism to perform offset correction on the coordinates coord i = (x i , y i , w i , h i ) of the detected target region; Specifically, extract the boundary information of the image from the image metadata imag_metas, denoted as bord_pixs. This information includes the cropping or scaling information of the image, in the form of border_pixs = (x border , y border , w border , h border ), where x border , y border are the upper left coordinates, and w border , h border are the width and height; For the coordinates of each dense region, use the following correction formula to perform offset correction: Target region cropping: After obtaining the dense region coordinates, divide the image into multiple sub-regions; the size of each sub-region is dynamically adjusted according to the density of the target, and the image cropping and annotation update process is completed through the following steps: box i =(x i , y i , w i , h i ); Among them, (x i , y i , w i , h i ) are the coordinates and dimensions of the i-th target region; Each target region determines whether it is selected for cropping by calculating its area A i : A i = w i × h i ; Sort by area and select the top t largest regions for further cropping; t = 1, 2, 3; Image cropping and scaling: For each selected target region, crop out the sub-image containing the target and perform scaling; the specific cropping region is calculated by the following formula: where WW and HH are the width and height of the original image respectively; Stitching of target regions: Stitch multiple cropped image patches. The stitching method is to arrange multiple image patches horizontally or vertically to form a new image; during the stitching process, keep the relative positions and sizes of the targets consistent; assuming the number of cropped image patches is M, the width and height of the final stitched image are: If the size of the stitched image exceeds the preset maximum width or height, scale it proportionally while keeping the proportion of the image content unchanged; Image padding and saving: There may be some blank areas in the stitched image, and the image size needs to be adjusted to the specified size through padding operations; the padding operations are carried out by the following formula: Fill the blank areas with zeros to obtain the finally stitched mosaic image.
6. The remote sensing target detection method based on regional positioning according to claim 1, wherein, The loss functions used during the training of the CenterNet target detector include focal loss, normalized Wasserstein distance loss, and L1 loss. The final loss function is defined as: L det = L k + λ nwd · L nwd + λ l1 · L1; Among them, λ nwd , λ l1 are both coefficients; L k is the focal loss, p i is the predicted probability, α is the balance factor, and μ is the hyperparameter that adjusts the influence degree of easy and hard samples; L nwd is the normalized Wasserstein distance loss, and W(B pred , B gt ) represents the predicted bounding box B pred and True box B gt The Wasserstein distance, w, between gt and h gt are the width and height of the true box respectively; L1 is the L1 loss, and represent the predicted coordinates and the true coordinates, respectively.