Wide-view-field high-resolution image target intelligent detection method based on multi-scale progressive cutting strategy

Through a multi-scale progressive cropping strategy, the density regression network and detector are built, and the detection area is dynamically covered, solving the efficiency and accuracy problems of object detection in wide field of view and high-resolution imaging, and achieving efficient and accurate object detection.

CN120355883APending Publication Date: 2025-07-22ZHENGZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510404926.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional methods are difficult to effectively detect targets of different sizes in wide field of view high resolution imaging, resulting in insufficient computing resources, reduced detection speed or inability to process in real time, and there is a problem of incomplete target cropping.

Method used

By adopting a multi-scale progressive cropping strategy, by constructing a density regression network and detector of multiple downsampling scales, dynamically cover the detection area, reducing redundant calculations, improving detection efficiency, and ensuring accurate detection of targets of different sizes.

Benefits of technology

It significantly improves the object detection performance of wide field of view and high resolution images, achieves efficient and accurate detection of targets of different sizes, and reduces the demand for computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355883A_ABST
    Figure CN120355883A_ABST
Patent Text Reader

Abstract

The invention aims to provide a wide-field-of-view high-resolution image target intelligent detection method based on a multi-scale progressive cutting strategy, which is particularly suitable for a wide-field-of-view high-resolution imaging scene. According to the method, the multi-scale down-sampling images are constructed, and the density regression network and the detector are combined, so that graded detection of targets with different sizes in the images is realized. According to the method, the detection area is dynamically covered, redundant calculation is reduced, the detection efficiency is improved, and meanwhile, the detection precision of targets of different sizes is guaranteed. Compared with a traditional method, the target detection performance of the wide-view-field high-resolution image is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an intelligent target detection method for wide-field high-resolution images based on a multi-scale progressive cropping strategy. Background Art

[0002] In wide-field high-resolution imaging, due to the different distances between the targets and the gigabit cameras, the sizes of the targets in the images vary greatly and the variation range is wide. The targets in the distance appear extremely small, while the targets nearby occupy a very large position, and targets of different sizes coexist at the same time. In such a complex scenario where the target size distribution is extremely uneven, the performance of traditional methods for target detection decreases.

[0003] To solve this problem, as mentioned in the patent of BOE Technology Group Co., Ltd., "A target detection method and a target detection device for wide-field images (Publication No. CN117079217A)", sliding windows of at least two sizes are used to cut the wide-field image to be detected respectively to obtain sub-images of different sizes. After using a preset target detection model to detect each sub-image and output the corresponding sub-image detection results, the duplicate target boxes in the sub-image detection results are removed and the sub-images of different sizes are globally fused to obtain the detection result of the wide-field image to be detected. However, when using a fixed-size sliding window to cut the wide-field image, problems such as incomplete cutting of over-sized targets and undetected targets are likely to occur, and it is still difficult to adapt to the current situation of target size diversity in wide fields. Moreover, facing wide-field images requires a larger amount of data and higher computational complexity. Processing such a large amount of data requires more computing resources, including memory, CPU, and GPU, etc. If the computing resources are insufficient, it may lead to a decrease in detection speed or inability to process in real time.

[0004] Therefore, there is an urgent need for an innovative solution that can effectively improve the computing power utilization efficiency without reducing the detection accuracy and achieve fast and accurate detection of targets of different sizes in wide-field high-resolution images. Summary of the Invention

[0005] In order to overcome the problems in the prior art, the purpose of the present invention is to provide an intelligent target detection method for wide-field high-resolution images based on a multi-scale progressive cropping strategy, which can dynamically cover the detection area, reduce redundant calculations, improve the detection efficiency, and at the same time ensure the detection accuracy of targets of different sizes and enhance the target detection performance of wide-field high-resolution images.

[0006] To achieve the above purpose, the present invention provides an intelligent target detection method for wide-field high-resolution images based on a multi-scale progressive cropping strategy, including the following steps:

[0007] Step S1: Preset multiple downsampling scales according to the image resolution. The multiple downsampling scales are set in ascending order of resolution. Construct and train a density regression network and a detector, and obtain an initial image.

[0008] Step S2: Input the image. According to the preset order of downsampling scales, perform a single downsampling process on the image, output the downsampled image to the density regression network to obtain a heat map of the center points of the target objects, and process the heat map to obtain a key area containing the target objects.

[0009] Step S3: Obtain the coordinates of the key area, map them to the original resolution image, mark the key area mapped in the original resolution image as the detection area, extract the detection area image and output it to the detector. After the detector recognizes it, output the bounding box coordinates and the target category labels, and save the information output by the detector.

[0010] Step S4: Determine whether all the operations of downsampling at the preset downsampling scales have been completed. If so, execute Step S5; if not, obtain the coordinates of the detection area, cover the detection area by generating a binary mask, and return the covered image to execute Step S2.

[0011] Step S5: Obtain all the bounding box coordinates and target category labels saved in Step S3, map them all again in the original resolution image, and output the image.

[0012] Further, in Step S1: The specific specifications of the preset multiple downsampling scales of the image are as follows: including four scales V1, V2, V3, and V4. Among them,

[0013] V1 is the coarsest scale, and the resolution is reduced to 1 / 8 of the initial image resolution.

[0014] V2 is the medium-coarse scale, and the resolution is reduced to 1 / 4 of the initial image resolution.

[0015] V3 is the medium-fine scale, and the resolution is reduced to 1 / 2 of the initial image resolution.

[0016] V4 is the finest scale, and the initial image resolution is retained.

[0017] Further, in Step S1, the method for constructing and training the density regression network is as follows:

[0018] Step S11: Construct a density regression network. Collect a large number of images and divide them into a training set and a test set. Manually annotate the centers and categories of the objects in the training set images, input the annotated images into the density regression network to generate a real heat map, and input the test set images into the density regression network to generate a predicted heat map.

[0019] Step S12: Calculate the loss function between the predicted heatmap and the ground truth heatmap;

[0020] Step S13: Use the gradient descent method to continuously adjust the parameters of the density regression network to minimize the mean squared error loss function. Set a preset number of training epochs. When the preset number of training epochs is reached, or when the decrease in the loss function value for consecutive epochs is less than the convergence threshold, it is determined that the training process ends;

[0021] Furthermore, in step S1, the method for constructing and training the detector is as follows:

[0022] Step S11’: Construct a detector that includes a classifier and a regressor. Collect a large number of images and divide them into a training set and a test set. Label the target objects in the training set images, including the class labels of the objects and the target position coordinates and bounding box data for the regression task;

[0023] Step S12’: Input the training set and test set images into the detector in sequence. By comparing with the annotation information, calculate the classifier loss and the regressor loss. Set the weights of the classifier loss and the regressor loss, and use the weighted sum of the classifier loss and the regressor loss as the total loss function of the detector;

[0024] Step S13’: Set a preset number of training epochs. When the preset number of training epochs is reached, or when the decrease in the loss function value for consecutive epochs is less than the convergence threshold, it is determined that the training process ends.

[0025] Furthermore, in step S12, the method for calculating the loss function between the predicted heatmap and the ground truth heatmap is specifically: Use the mean squared error (MSE) loss function to measure the difference between the predicted heatmap and the ground truth heatmap;

[0026] The calculation formula of the MSE loss function is:

[0027]

[0028] where N is the total number of pixels in the heatmap, Y i is the i-th pixel value of the ground truth heatmap, and is the i-th pixel value of the predicted heatmap:.

[0029] Furthermore, in step S12’, the method for calculating the classifier loss is specifically:

[0030] Use the cross-entropy loss function to measure the difference between the predicted class probability distribution and the ground truth class label during the training of the classifier. By minimizing the cross-entropy loss function, complete the training of the classifier to distinguish the target classes; The expression of the cross-entropy loss function is:

[0031]

[0032] Among them, C is the total number of categories, yc is the c-th element in the true category label (taking values of 0 or 1), and pc is the c-th element in the predicted category probability distribution.

[0033] In step S12’, the method for calculating the regressor loss is specifically as follows:

[0034] The Smooth L1 Loss function is adopted for the regressor to measure the difference between the predicted bounding box and the true bounding box; its expression is as follows:

[0035]

[0036] Among them, x is the error between the predicted value and the true value.

[0037] Furthermore, in step S12’, the weighted sum of the classifier loss and the regressor loss is used as the total loss function of the detector, and the expression is:

[0038] L total =λ CE L CE +λ Smooth L1 L Smooth L1

[0039] Among them, λ CE and λ Smooth L1 are respectively the preset weight coefficients of the classifier loss and the regressor loss.

[0040] Furthermore, in step S2, the specific method for outputting the downsampled image to the density regression network to obtain the heat map of the center points of the target objects and processing the heat map to obtain the key regions containing the target objects is as follows:

[0041] Step S21: Input the image and perform a downsampling operation on the image once according to the preset downsampling scale order;

[0042] Step S22: The density regression network recognizes the input image to obtain the heat map of the center points of the target objects, and the confidence level of the existence of the target center point at each pixel position is marked in the heat map;

[0043] Step S23: Set the confidence level threshold, and mark the pixels greater than the confidence level threshold as key pixels;

[0044] Step S24: Perform dilation operations and connected component analysis on the key pixels to generate key regions.

[0045] Furthermore, the method of step S3 is specifically as follows:

[0046] Step S31: Obtain the coordinates of the key regions in step S2, map them in the original resolution image and mark them as detection regions;

[0047] Step S32: Input the image of the detection area into the detector. The detector performs object recognition on the detection area and generates a bounding box and an object classification result.

[0048] Step S33: Use non-maximum suppression technology to filter out redundant detection results.

[0049] Furthermore, the method of mapping the key area coordinates in step S2 to the original resolution image in step S31 is specifically as follows:

[0050] Specifically, step S311: Obtain the current image size, with width W1 and height H1; obtain the original resolution image size, with width W2 and height H2; obtain the coordinates of the key area at the current scale, and define the upper left corner coordinates of the key area as (X1, Y1) and the lower right corner coordinates as (X2, Y2);

[0051] Step S312: Calculate the scaling ratio in the horizontal direction (width direction) and the scaling ratio in the vertical direction (height direction)

[0052] Step S313: According to the scaling ratio, calculate the coordinates of the rectangular box of the key area after mapping, and define the upper left corner coordinates of the key area after mapping as (X1', Y1'), and the calculation formula is:

[0053] Abscissa: X1' = X1 × r w ;

[0054] Ordinate: Y1' = Y1 × r h ;

[0055] Define the lower right corner coordinates (X2', Y2') of the key area after mapping, and the calculation formula is:

[0056] Abscissa: X2' = X2 × r w ;

[0057] Ordinate: Y2' = Y2 × r h ;

[0058] Thus, the coordinates of the key area after mapping are obtained as (X1', Y1') to (X2', Y2'), thereby determining the position of the key area in the original resolution image.

[0059] Through constructing multiple-scale downsampled images and combining a density regression network with a detector, the present invention realizes efficient and accurate hierarchical detection of objects of different sizes in an image. By dynamically covering the detection area, redundant calculations are reduced, the detection efficiency is improved, and at the same time, the detection accuracy for objects of different sizes is ensured. Compared with traditional methods, the present invention significantly improves the object detection performance of wide-field high-resolution images. Description of the Drawings

[0060] Figure 1 It is a flowchart of the steps of the intelligent detection method for wide-field high-resolution image targets based on the multi-scale progressive cropping strategy provided by the present invention; Detailed Embodiments

[0061] The present invention aims to provide an intelligent detection method for wide-field high-resolution image targets based on the multi-scale progressive cropping strategy, as Figure 1 shown, including the following steps:

[0062] Step S1: Preset multiple downsampling scales of the image according to the image resolution. The multiple downsampling scales are set in the order of increasing resolution. Construct and train a density regression network and a detector to obtain an initial image;

[0063] Specifically, the initial image obtained in this embodiment is a wide-field high-resolution image (a ten-billion-pixel-level image, with an initial image size of 32768×32768). In step S1, 4 image downsampling scale specifications are preset. In the order of increasing resolution, they are:

[0064] V1 is the coarsest scale, and the resolution is reduced to 1 / 8 of the initial image resolution (4096×4096), which is used to detect extremely large objects (such as targets occupying more than 5% of the image area).

[0065] V2 is the medium-coarse scale, and the resolution is reduced to 1 / 4 of the initial image resolution (8192×8192), which is used to detect medium-sized objects (such as targets occupying 1% - 5% of the image area).

[0066] V3 is the medium-fine scale, and the resolution is reduced to 1 / 2 of the initial image resolution (16384×16384), which is used to detect smaller objects (such as targets occupying 0.1% - 1% of the image area).

[0067] V4 is the finest scale, retaining the initial image resolution, which is used to detect and process extremely small objects (such as targets occupying less than 0.1% of the image area).

[0068] In step S1, the method for constructing and training the density regression network includes the following steps:

[0069] Step S11: Construct a density regression network, collect a large number of images and divide them into a training set and a test set. Manually annotate the centers and categories of the objects in the training set images, input the annotated images into the density regression network to generate a true heat map, and input the test set images into the density regression network to generate a predicted heat map.

[0070] Step S12: Calculate the loss function between the predicted heat map and the true heat map;

[0071] The method for calculating the loss function between the predicted heat map and the real heat map is specifically as follows: the mean square error (MSE) loss function is used to measure the difference between the predicted heat map and the real heat map;

[0072] The calculation formula of the MSE loss function is:

[0073]

[0074] where N is the total number of pixels in the heat map, Yi is the i-th pixel value of the real heat map, is the i-th pixel value of the predicted heat map.

[0075] Step S13: Using the gradient descent method, by continuously adjusting the parameters of the density regression network, minimize the mean square error loss function, and preset the number of training rounds. When the preset number of training rounds is reached, or when the decrease amplitude of the loss function value for multiple consecutive rounds is less than the convergence threshold, it is determined that the training process ends;

[0076] Specifically, the gradient descent method is used to train the density regression network. In this embodiment, the Adam optimization algorithm is specifically used, which can adaptively adjust the learning rate of each parameter, and the initial learning rate is set to 0.001. During each round of training, according to the gradient direction of the loss function, the parameters of the density regression network are continuously adjusted to minimize the mean square error loss function value between the predicted heat map and the real heat map as much as possible.

[0077] The preset number of training rounds is 100 rounds. At the same time, to avoid unnecessary training, a loss function convergence threshold is also preset. In this embodiment, the convergence threshold is 0.0001. During the training process, if the decrease amplitude of the loss function value for 10 consecutive rounds is less than the convergence threshold of 0.0001, it indicates that the model has basically converged. At this time, even if the preset number of training rounds is not reached, it can be determined that the training process ends. If the condition for the loss function to converge is never met, then when the preset 100-round training is reached, the training process will also end.

[0078] In step S1, the method for constructing and training the detector specifically includes the following steps:

[0079] Step S11': Construct a detector, which includes a classifier and a regressor. Collect a large number of images and divide them into a training set and a test set. Label the target objects in the training set images, including the class labels of the objects and the target object coordinates and bounding box data for the regression task;

[0080] Specifically, the detector is based on a deep learning model, such as an object detection model based on a convolutional neural network (CNN), such as Faster R-CNN, YOLO series, etc. The model structure inside the detector extracts features from the input image and generates predictions for the target objects through a series of operations such as convolutional layers, pooling layers, and fully connected layers.

[0081] For the classifier, the training set includes image samples with true class labels (in one-hot encoded form) and their corresponding key regions.

[0082] For the regressor, the training set includes image samples with true bounding box data and their corresponding key regions. The true bounding box data includes: the center point coordinates (x center , y center ) of the bounding box, the width w, and the height h.

[0083] Step S12’: Input the training set and test set images into the detector in sequence. By comparing with the annotation information, calculate the classifier loss and the regressor loss. Preset the weights of the classifier loss and the regressor loss, and use the weighted sum of the classifier loss and the regressor loss as the total loss function of the detector.

[0084] Specifically, the training of the classifier uses the cross-entropy loss function to measure the difference between the predicted class probability distribution and the true class label. By minimizing the cross-entropy loss function, the training of the classifier to distinguish the target classes is completed. The expression of the cross-entropy loss function is as follows:

[0085]

[0086] where C is the total number of classes, y c is the c-th element (taking values of 0 or 1) in the one-hot encoding of the true class label, and p c is the c-th element in the predicted class probability distribution.

[0087] For the regressor, calculate the offset of the center point coordinates between the predicted bounding box and the true bounding box width offset Δw and height offset where is the center point of the predicted bounding box, (x center , y center ) is the center point coordinates of the true bounding box, is the width of the predicted bounding box, w is the width of the true bounding box, To predict the width of the bounding box, h is the width of the ground truth bounding box. These offsets are used as the regression targets, and the Smooth L1 Loss function is used to measure the prediction error of these offsets. The expression is as follows:

[0088]

[0089] where x is the offset error between the predicted value and the ground truth value.

[0090] The weighted sum of the classifier loss and the regressor loss is used as the total loss function of the detector, and the expression is:

[0091] L total = λ CE L CE + λ Smooth L1 L Smooth L1

[0092] where λ CE and λ SmoothL1 are the preset weight coefficients of the classifier loss and the regressor loss respectively, which are used to adjust the relative importance of the classification loss and the regression loss in the total loss.

[0093] Step S13’: Preset the number of training rounds. When the preset number of training rounds is reached, or when the decrease amplitude of the loss function value is less than the convergence threshold for multiple consecutive rounds, it is determined that the training process ends.

[0094] The working principle of step S13’ is the same as that of step S13, and will not be elaborated here.

[0095] Step S2: Input an image. According to the preset downsampling scale order, perform a single downsampling process on the image, output the downsampled image to the density regression network to obtain the heatmap of the center points of the target objects, and process the heatmap to obtain the key regions containing the target objects. Specifically, it includes the following steps:

[0096] Step S21: Input an image. According to the preset downsampling scale order, perform a single downsampling operation on the image.

[0097] Specifically, initially input the original image. According to the preset multiple image scale downsampling specifications, first perform the first downsampling on the original image according to the V1 scale.

[0098] Step S22: The density regression network recognizes the input image to obtain the heatmap of the center points of the target objects. The confidence level of the existence of the target center points is marked at each pixel position in the heatmap;

[0099] Specifically, in this embodiment, the density regression network is based on a convolutional neural network architecture. Components such as convolutional layers, activation function layers, and pooling layers inside it will perform feature extraction and transformation on the input image, and finally output a heatmap of the center points of the target objects. The value of each pixel position in the heatmap represents the confidence level of the center point of the target object existing at that position. The higher the confidence level, the more likely that position is the center point of the target object.

[0100] Step S23: Set a confidence threshold, and mark the pixels with a confidence level greater than the confidence threshold as key pixels;

[0101] In this embodiment, the confidence threshold can be set to 0.85, and the pixels with a confidence level greater than 0.85 are marked as key pixels.

[0102] Step S24: Perform dilation operations and connected component analysis on the key pixels to generate key regions.

[0103] In this embodiment, in order to connect adjacent key pixels to form a more complete target region, dilation operations are performed on the marked key pixels. A 3×3 all-ones matrix is preset as the dilation kernel. Starting from the upper left corner of the heatmap, each pixel is traversed in turn. Align the center of the dilation kernel with each key pixel in turn, and perform operations on the neighboring pixels covered by the kernel. Mark the neighboring pixels of this key pixel as key pixels as well, thus realizing the dilation and expansion of the key pixels.

[0104] For the dilated image, scan the pixels row by row starting from the upper left corner. When a key pixel is encountered, check whether its neighboring pixels (upper, lower, left, right, upper left, upper right, lower left, lower right, a total of 8 neighboring pixels) are also key pixels. If so, mark them as belonging to the same connected component and assign the same marking value. By continuously recursively checking and marking, finally all the mutually connected key pixels are divided into different connected components. Calculate the circumscribed rectangle or the minimum convex polygon of each connected component, and use it as the key region. The key region represents the region where the target object may exist.

[0105] Step S3: Obtain the coordinates of the key region, map them to the original resolution image, mark the key region mapped in the original resolution image as the detection region, extract the detection region image and output it to the detector, and the detector outputs the target category and the bounding box coordinates after recognition. Specifically, it includes the following steps:

[0106] Step S31: Obtain the coordinates of the key region in Step S2, map them in the original resolution image, and mark them as the detection region;

[0107] Specifically, in this embodiment, if the current image scale is V1 (the coarsest scale), its resolution is reduced to 1 / 8 of the original image, and the size is 4096×4096, that is, the width W1 = 4096 and the height H1 = 4096; the original resolution image scale is V4 (the finest scale), and the size is 32768×32768, that is, the width W2 = 32768 and the height H2 = 32768.

[0108] Obtain the coordinates of the key region at the current V1 scale, including the upper left corner coordinates (X1, Y1) and the lower right corner coordinates (X2, Y2) of the key region.

[0109] Calculate the scaling ratio in the horizontal direction (width direction) and the scaling ratio in the vertical direction (height direction)

[0110] Calculate the coordinates of the rectangle frame of the key region after mapping according to the scaling ratio.

[0111] Specifically, the calculation formula for the upper left corner coordinates (X1', Y1') of the key region after mapping is:

[0112] Abscissa: X1' = X1 × r w = X1 × 8;

[0113] Ordinate Y1' = Y1 × r h = Y1 × 8.

[0114] The calculation formula for the lower right corner coordinates (X2', Y2') of the key region after mapping is:

[0115] Abscissa: X2' = X2 × r w = X2 × 8;

[0116] Ordinate Y2' = Y2 × r h = Y2 × 8.

[0117] Through the above steps, the upper left corner coordinates (X1, Y1) and the lower right corner coordinates (X2, Y2) of the key pixel region at the current scale V1 can be mapped to the original resolution image scale V4, and the upper left corner coordinates (X1', Y1') and the lower right corner coordinates (X2', Y2') after mapping are obtained, so as to determine the position of the key region in the original resolution image. Finally, the region included in the upper left corner coordinates (X1', Y1') and the lower right corner coordinates (X2', Y2') after mapping is marked as the detection region.

[0118] Step S32: Input the detection region image into the detector, and the detector performs target recognition on the detection region and generates a bounding box and a target classification result.

[0119] Input the detection area image into the trained detector. The detector can predict the exact position of the target object in the detection area image, which is represented in the form of a bounding box and usually includes the center point coordinates of the bounding box as well as the width of the bounding box and the height And the detector can also classify the detected target object and output the probability distribution (p1, p2,..., p c ) of the target object belonging to each preset class, where c is the total number of preset classes, and p c represents the probability that the target object belongs to the c-th class. The model will determine the most likely class of the target object according to the magnitude of the probability value. Finally, the output of the detector includes the detection area image with the bounding box for target recognition and the target classification label

[0120] Step S33: Use non-maximum suppression (NMS) technology to filter out redundant detection results

[0121] Specifically, for multiple predicted bounding boxes with a high degree of overlap, only the one with the highest score (such as the class probability) is retained, and the rest are deleted. This can ensure the accuracy and uniqueness of the detection results and avoid generating multiple duplicate detection boxes for the same target

[0122] Step S4: Determine whether all the operations of downsampling at the preset downsampling scales have been completed. If so, execute Step S5; if not, obtain the coordinates of the detection area, cover the detection area by generating a binary mask, and return the processed image to execute Step S2

[0123] Specifically, define Q as the number of preset image downsampling scales. In this embodiment, a total of four downsampled images with scales V1, V2, V3, and V4 are set, so a total of four rounds of downsampling processing are required, that is, Q = 4. The specific method is to define k as the number of loops, with an initial value of k = 1. Determine whether k is equal to Q. If not, then k = k + 1, and then cover the detection area and return the image to Step S2 to be input into the density regression model; when k is equal to Q, it means that all the image processing at the preset downsampling scales has been completed, and continue to execute Step S5

[0124] The method of covering the detection area by generating a binary mask is specifically as follows: Create a blank binary mask image M with the same size as the original resolution scale image, whose width is W2 = 32768 and height is H2 = 32768, and the initial value of all pixels is set to 0, representing the background. Fill the mask area: According to the coordinates (X1', Y1') and (X2', Y2') of the detection area after mapping, set the pixel values within the corresponding detection area in the mask image M to 1, representing the target area. The specific operation is to traverse all pixels (x, y) in the mask image M whose coordinates satisfy X1′≤X≤X2′ and Y1′≤Y≤Y2′, and set the value of M(x, y) to 1. No subsequent image will operate on the area set to 1. Thus, the process of covering the coordinates of the detection area is completed.

[0125] Step S5: Obtain all the bounding box coordinates and target class labels saved in step S3, map them all again in the original resolution image, and output the image.

[0126] Specifically, in step S5, all the bounding box coordinates and target class labels saved in step S3 are mapped again in the original resolution image. The principle is the same as that of step S31 above and will not be elaborated here. Thus, the target intelligent detection of the wide-field high-resolution image can be completed, and a wide-field high-resolution image with bounding boxes and target class labels can be output.

Claims

1. An intelligent target detection method for wide-field high-resolution images based on a multi-scale progressive cropping strategy, characterized in that It includes the following steps: Step S1: Preset multiple downsampling scales according to the image resolution. The multiple downsampling scales are set in the order of increasing resolution. Construct and train a density regression network and a detector, and obtain an initial image; Step S2: Input the image. According to the preset downsampling scale order, perform a single downsampling process on the image, output the downsampled image to the density regression network to obtain a heatmap of the center points of the target objects, and process the heatmap to obtain the key regions containing the target objects; Step S3: Obtain the coordinates of the key regions, map them to the original resolution image, mark the key regions mapped in the original resolution image as detection regions, extract the detection region images and output them to the detector. After the detector identifies them, output the bounding box coordinates and target category labels, and save the information output by the detector; Step S4: Determine whether all the operations of the preset downsampling scales have been completed. If so, execute Step S5; if not, obtain the coordinates of the detection region, cover the detection region by generating a binary mask, and return the covered image to execute Step S2; Step S5: Obtain all the bounding box coordinates and target category labels saved in Step S3, map them all again in the original resolution image, and output the image.

2. The intelligent target detection method for wide - field high - resolution images based on the multi - scale progressive cropping strategy according to claim 1, characterized in that, In Step S1: The specific specifications of the preset multiple downsampling scales of the image are as follows: including four scales V1, V2, V3, and V4. Among them, V1 is the coarsest scale, and the resolution is reduced to 1 / 8 of the initial image resolution; V2 is the medium-coarse scale, and the resolution is reduced to 1 / 4 of the initial image resolution; V3 is the medium-fine scale, and the resolution is reduced to 1 / 2 of the initial image resolution; V4 is the finest scale, and the initial image resolution is retained.

3. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 1, characterized in that, In Step S1, the method for constructing and training the density regression network is: Step S11: Construct a density regression network. Collect a large number of images and divide them into a training set and a test set. Manually annotate the centers and categories of the objects in the training set images, input the annotated images into the density regression network to generate a real heatmap, and input the test set images into the density regression network to generate a predicted heatmap; Step S12: Calculate the loss function between the predicted heatmap and the real heatmap; Step S13: Adopt the gradient descent method. By continuously adjusting the parameters of the density regression network, minimize the mean square error loss function. Preset the number of training rounds. When the preset number of training rounds is reached, or the continuous decline amplitude of the loss function value in multiple rounds is less than the convergence threshold, determine that the training process ends.

4. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 1, characterized in that, In Step S1, the method for constructing and training the detector is: Step S11’: Construct a detector. The detector includes a classifier and a regressor. Collect a large number of images and divide them into a training set and a test set. Annotate the target objects in the training set images, including the category labels of the objects and the target position coordinates and bounding box data for the regression task; Step S12’: Input the training set and test set images into the detector in turn. By comparing with the annotation information, calculate the classifier loss and the regressor loss. Preset the weights of the classifier loss and the regressor loss, and use the weighted sum of the classifier loss and the regressor loss as the total loss function of the detector; Step S13’: Preset the number of training rounds. When the preset number of training rounds is reached, or when the decrease in the loss function value is less than the convergence threshold for multiple consecutive rounds, it is determined that the training process ends.

5. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 3, characterized in that, In step S12, the method for calculating the loss function between the predicted heatmap and the ground truth heatmap is specifically as follows: The mean squared error (MSE) loss function is used to measure the difference between the predicted heatmap and the ground truth heatmap. The calculation formula of the MSE loss function is: where N is the total number of pixels in the heat map, and Yi is the value of the i-th pixel in the ground truth heat map, is the value of the i-th pixel in the predicted heat map:

6. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 4, characterized in that, In step S12’, the method for calculating the classifier loss is specifically as follows: For the training of the classifier, the cross-entropy loss function is used to measure the difference between the predicted class probability distribution and the ground truth class label. By minimizing the cross-entropy loss function, the training of the classifier to distinguish the target class is completed. The expression of the cross-entropy loss function is: where C is the total number of classes, yc is the c-th element in the ground truth class label (taking values of 0 or 1), and pc is the c-th element in the predicted class probability distribution. The method for calculating the regressor loss is specifically as follows: For the regressor, the smooth L1 loss function is used to measure the difference between the predicted bounding box and the ground truth bounding box. Its expression is as follows: where x is the error between the predicted value and the ground truth value.

7. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 4, characterized in that, In step S12’, the weighted sum of the classifier loss and the regressor loss is used as the total loss function of the detector, and the expression is: L total = λ CE L CE + λ Smooth L1 L Smooth L1 Among them, λ CE and λ SmoothL1 are the weight coefficients of the preset classifier loss and regressor loss, respectively.

8. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 1, characterized in that, In step S2, the method for outputting the downsampled image to the density regression network to obtain the heatmap of the center points of the target objects and processing the heatmap to obtain the key region containing the target objects is specifically as follows: Step S21: Input the image and perform a downsampling operation on the image according to the preset downsampling scale order. Step S22: The density regression network recognizes the input image to obtain the heatmap of the center points of the target objects. Each pixel position in the heatmap marks the confidence of the existence of the target center point. Step S23: Set the confidence threshold, and mark the pixels greater than the confidence threshold as key pixels. Step S24: Perform dilation operations and connected component analysis on the key pixels to generate the key region.

9. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 1, characterized in that The method of step S3 is specifically as follows: Step S31: Obtain the coordinates of the key region in step S2, map them to the original resolution image, and mark them as the detection region. Step S32: Input the detection region image into the detector. The detector performs target recognition on the detection region and generates the bounding box and the target classification result. Step S33: Use the non-maximum suppression technique to filter out redundant detection results.

10. The intelligent target detection method for wide-field high-resolution images based on the multi-scale progressive cropping strategy according to claim 9, wherein, The method of mapping the coordinates of the key region in step S2 to the original resolution image in step S31 is specifically as follows: Specifically, step S311: Obtain the current image size, with width W1 and height H1; obtain the original resolution image size, with width W2 and height H2; obtain the coordinates of the key region at the current scale, and define the upper left corner coordinates of the key region as (X1, Y1) and the lower right corner coordinates as (X2, Y2). Step S312: Calculate the scaling ratio in the horizontal direction (width direction) and the scaling ratio in the vertical direction (height direction) Step S313: According to the scaling ratio, calculate the coordinates of the rectangle of the mapped key region, and define the upper left corner coordinates of the mapped key region as (X1', Y1'). The calculation formula is: Abscissa: X1' = X1 × r w ; Vertical coordinate: Y1' = Y1 × r h ; Define the coordinates (X2', Y2') of the lower right corner of the key area after mapping. The calculation formula is as follows: Abscissa: X2' = X2 × r w ; Ordinate: Y2' = Y2 × r h ; Thus, the coordinates of the key area after mapping are from (X1', Y1') to (X2', Y2'), thereby determining the key area in the original resolution image.

Citation Information

Patent Citations

  • Target detection method and target detection device for wide-field-of-view image

    CN117079217A