High-Resolution Image Target Object Detection Method and Device Based on Reverse Segmentation
By reducing resolution screening and restoring high-resolution images, combined with sliding window slices and detectors, the problem of insufficient small and medium-sized target recall in traditional methods is solved, and efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202510551622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Traditional target detection methods have insufficient recall when dealing with small targets, making it difficult to balance detection speed and accuracy in scenarios with multiple targets, large resolution differences or complex target types.
By performing resolution reduction processing on the original high-resolution image, the image areas of the target object category are filtered out and restored to the high-resolution image, combined with sliding window slices and detectors of different structures for object detection, the image area quality is evaluated for accurate positioning.
It improves the recall rate of small targets, improves detection efficiency and accuracy, improves detection capabilities in complex scenarios, and takes into account both speed and accuracy.
Smart Images

Figure CN120070875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for detecting target objects in high-resolution images based on reverse segmentation. Background Art
[0002] In modern society, with the rapid development of monitoring technology and image processing technology, image analysis has become increasingly widely used in fields such as security monitoring, intelligent transportation, and autonomous driving. Especially in target detection, accurately identifying and locating targets is crucial for ensuring public safety and traffic management.
[0003] Traditional target detection methods rely on a two-stage processing flow. First, region screening is performed on low-resolution images, and then precise detection is carried out on high-resolution images. However, when dealing with small targets, this method often suffers from insufficient recall. Since small targets are not easily accurately identified in low-resolution images, they are omitted in subsequent stage detections, thus affecting the overall detection effect. In addition, although the progressive magnification target detection method can gradually magnify the image through a pyramid structure, if small targets are not detected in the initial stage, subsequent magnification cannot make up for this defect. Therefore, related technologies are difficult to balance detection speed and accuracy when dealing with scenarios with multiple targets, large resolution differences, or complex target types. Summary of the Invention
[0004] The present invention provides a method and device for detecting target objects in high-resolution images based on reverse segmentation, aiming to solve the problems existing in the above background art.
[0005] To solve the above technical problems, the present invention is implemented as follows:
[0006] The present invention provides a method for detecting target objects in high-resolution images based on reverse segmentation, including:
[0007] Reducing the resolution of the original high-resolution image to obtain a low-resolution image;
[0008] Filtering out image regions with non-target object contours from the low-resolution image and retaining image regions with target object contours;
[0009] Restoring the resolution of the low-resolution image with the retained image regions having target object contours to the resolution of the original high-resolution image;
[0010] Slicing the high-resolution image with restored resolution through a sliding window to obtain a plurality of image regions;
[0011] Determining the quality of each of the plurality of image regions;
[0012] Performing object detection on image regions of different qualities through detectors of different structures to obtain the object detection results of each of the multiple image regions;
[0013] Obtaining the object detection result of the original high-resolution image according to the object detection results of each of the multiple image regions.
[0014] Optionally, filtering out image regions with contours of non-target object categories from the low-resolution image and retaining image regions with contours of target object categories, including:
[0015] Identifying the contours of each target from the low-resolution image through a pre-trained segmentation model;
[0016] Classifying the contours of each target to obtain contour classification results;
[0017] Filtering out image regions of targets with contours of non-target object categories and retaining image regions of targets with contours of target object categories.
[0018] Optionally, classifying the contours of each target to obtain contour classification results, including:
[0019] For each target among the various targets, determining the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target;
[0020] For each target among the various targets, using the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, converting each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system;
[0021] Processing the coordinates of each pixel point on the contour of each target in the polar coordinate system through a pre-trained classifier to determine whether the category of the contour of each target is the target object category.
[0022] Optionally, restoring the resolution of the low-resolution image with image regions having contours of target object categories to the resolution of the original high-resolution image, including:
[0023] Determining the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and restoring the resolution of the low-resolution image with image regions having contours of target object categories to the resolution of the original high-resolution image according to the scaling ratio; or
[0024] According to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system, map the coordinates of the pixel points on the contour of the target object on the low-resolution image to the high-resolution image.
[0025] Optionally, after restoring the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image, it further includes:
[0026] According to the filtered image area whose contour is the non-target object category, perform mask marking on the high-resolution image after resolution restoration to obtain a high-resolution image including a mask area and a remaining area, where the mask area corresponds to the filtered image area whose contour is the non-target object category;
[0027] Slice the high-resolution image after resolution restoration through a sliding window to obtain a plurality of image areas, including:
[0028] Slice the remaining area in the high-resolution image after mask marking through a sliding window to obtain a plurality of image areas.
[0029] Optionally, determining the quality of each of the plurality of image areas includes:
[0030] Determine the quality level of each of the plurality of image areas according to the clarity, noise level, and texture complexity of each of the plurality of image areas;
[0031] Perform object detection on image areas of different qualities through detectors of different structures, including:
[0032] For each image area in the plurality of image areas, when the quality level of the image area is low quality, perform object detection on the image area through a pre-trained slow object detector, and when the quality level of the image area is high quality, perform object detection on the image area through a pre-trained fast object detector;
[0033] The slow object detector and the fast object detector satisfy at least one of the following:
[0034] The inference speed of the slow object detector is less than the inference speed of the fast object detector;
[0035] The number of model parameters of the slow object detector is greater than the number of model parameters of the fast object detector;
[0036] The target object feature extraction ability of the slow object detector is greater than the target object feature extraction ability of the fast object detector;
[0037] The structural complexity of the slow object detector is greater than that of the fast object detector.
[0038] Optionally, obtaining the object detection result of the original high-resolution image according to the object detection results of the respective multiple image regions includes:
[0039] Fusing the object detection results of the respective multiple image regions;
[0040] Removing redundant object position detection frames in the fused object detection result to obtain the object detection result of the original high-resolution image.
[0041] In a second aspect, an embodiment of the present disclosure provides a high-resolution image object detection device based on reverse segmentation, including:
[0042] A low-resolution processing module for reducing the resolution of the original high-resolution image to obtain a low-resolution image;
[0043] A filtering module for filtering out image regions with non-object category contours from the low-resolution image and retaining image regions with object category contours;
[0044] A restoration module for restoring the resolution of the low-resolution image with the retained image regions having object category contours to the resolution of the original high-resolution image;
[0045] A slicing module for slicing the high-resolution image after resolution restoration through a sliding window to obtain a plurality of image regions;
[0046] A quality analysis module for determining the quality of the respective multiple image regions;
[0047] An object detection module for performing object detection on image regions of different qualities through detectors of different structures to obtain object detection results of the respective multiple image regions;
[0048] A result output module for obtaining the object detection result of the original high-resolution image according to the object detection results of the respective multiple image regions.
[0049] Optionally, the filtering module includes:
[0050] A contour recognition sub-module for identifying the contours of each target from the low-resolution image through a pre-trained segmentation model;
[0051] A classification sub-module for classifying the contours of each target to obtain a contour classification result;
[0052] A filtering sub-module, configured to filter out the image regions of the targets whose contours are of non-target object categories, and retain the image regions of the targets whose contours are of target object categories.
[0053] Optionally, the classification sub-module includes:
[0054] A coordinate determination unit, configured to, for each of the respective targets, determine the coordinates of the centroid of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target;
[0055] A coordinate conversion unit, configured to, for each of the respective targets, use the coordinates of the centroid of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, and convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system, to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system;
[0056] A coordinate processing unit, configured to process the coordinates of each pixel point on the contour of each of the respective targets in the polar coordinate system through a pre-trained classifier, to determine whether the category of the contour of each of the respective targets is the target object category.
[0057] Optionally, the restoration module includes:
[0058] A scaling sub-module, configured to determine the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and according to the scaling ratio, restore the resolution of the low-resolution image with the image regions whose contours are of target object categories retained to the resolution of the original high-resolution image; or
[0059] A mapping sub-module, configured to map the coordinates of the pixel points on the contour of the target object in the low-resolution image to the high-resolution image according to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system.
[0060] Optionally, it further includes:
[0061] A masking module, configured to perform masking marking on the high-resolution image after resolution restoration according to the image regions whose contours are of non-target object categories and are filtered out, to obtain a high-resolution image including a masked region and a remaining region, where the masked region corresponds to the image regions whose contours are of non-target object categories and are filtered out;
[0062] The slicing module includes:
[0063] A slicing sub-module for slicing the remaining area in the masked high-resolution image through a sliding window to obtain a plurality of image regions.
[0064] Optionally, the quality analysis module includes:
[0065] A quality level determination sub-module for determining the quality level of each of the plurality of image regions according to the clarity, noise level, and texture complexity of each of the plurality of image regions;
[0066] The target detection module includes:
[0067] A classification detection sub-module for performing target detection on each of the plurality of image regions through a pre-trained slow target detector when the quality level of the image region is low quality, and performing target detection on the image region through a pre-trained fast target detector when the quality level of the image region is high quality;
[0068] At least one of the following is satisfied between the slow target detector and the fast target detector:
[0069] The inference speed of the slow target detector is less than the inference speed of the fast target detector;
[0070] The number of model parameters of the slow target detector is greater than the number of model parameters of the fast target detector;
[0071] The target object feature extraction ability of the slow target detector is greater than the target object feature extraction ability of the fast target detector;
[0072] The structural complexity of the slow target detector is greater than the structural complexity of the fast target detector.
[0073] Optionally, the result output module includes:
[0074] A fusion sub-module for fusing the target object detection results of each of the plurality of image regions;
[0075] A removal sub-module for removing redundant target object position detection frames in the fused target object detection results to obtain the target object detection result of the original high-resolution image.
[0076] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and capable of running on the processor, and when the computer program is executed by the processor, the steps of the high-resolution image target object detection method based on reverse segmentation as described in the first aspect are implemented.
[0077] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the high-resolution image target object detection method based on reverse segmentation as described in the first aspect are implemented.
[0078] The technical solutions provided by the present invention at least bring the following beneficial effects:
[0079] By reducing the resolution of the original high-resolution image, the present invention can quickly process large-scale image data and improve the processing efficiency. In the low-resolution image, a screening mechanism is used to effectively filter out regions with non-target object categories in the contour and retain regions with target object categories in the contour, significantly reducing the redundant calculations in subsequent detections and concentrating resources on the detection of important targets. After restoring the retained regions to the original high resolution, the image is carefully analyzed through the sliding window slicing technique, ensuring the accurate positioning of the target object. By evaluating the quality of multiple image regions and combining detectors with different structures for targeted object detection, both the detection speed and accuracy are taken into account. When dealing with complex scenarios, the present invention can effectively handle image regions with different qualities and features, improve the recall rate of small targets, provide efficient and accurate object detection outputs, and greatly improve the detection capabilities of traditional methods in multi-target and complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0081] Figure 1 It is a schematic diagram of the steps of the high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention;
[0082] Figure 2 It is a schematic diagram of the overall process of the high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention;
[0083] Figure 3 It is a block diagram of the structure of the high-resolution image target object detection device based on reverse segmentation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0085] The object detection method for high-resolution images of target objects based on reverse segmentation proposed by the present invention aims to significantly improve the detection speed of target objects while ensuring the detection accuracy of target objects through refined processing of key regions. The core idea of reverse segmentation is to utilize the fast processing ability of low-resolution images to initially filter out regions of non-target object categories and only retain the image regions containing target object categories. In this way, the subsequent high-resolution detection process can focus on the regions more likely to contain the target, thereby avoiding redundant processing of irrelevant regions and significantly improving the detection efficiency.
[0086] It can be understood that the detection method proposed by the present invention can be effectively applied to the detection scenarios of various types of target objects, including multiple fields such as traffic monitoring, intelligent security, unmanned driving, and industrial monitoring, and can meet the requirements of target detection in different scenarios.
[0087] Figure 1 is a schematic diagram of the steps of the object detection method for high-resolution images of target objects based on reverse segmentation provided by an embodiment of the present invention. As Figure 1 shown, it includes:
[0088] Step S101, reduce the resolution of the original high-resolution image to obtain a low-resolution image.
[0089] The original high-resolution image is the original input image without resolution adjustment, and its resolution is higher than the standard display or processing requirements. For example, the resolution ≥ 1920×1080 pixels (i.e., full high definition and above). The original high-resolution image can be reduced to a low-resolution image through scaling operations, such as bilinear interpolation and area sampling methods. The low-resolution image retains the global contour features of the target object while significantly reducing the data volume.
[0090] Step S102, filter out the image regions with non-target object category contours from the low-resolution image and retain the image regions with target object category contours.
[0091] Figure 2 is a schematic diagram of the overall process of the object detection method for high-resolution images of target objects based on reverse segmentation provided by an embodiment of the present invention. Please refer to Figure 2, first, Poisson distribution sampling is used to simulate and generate sparse data from the low-resolution image to increase data diversity. This can help the subsequent segmentation model better learn the features of different scenarios and targets, improving the robustness and generalization ability of the segmentation model. Further, all the contours are identified from the low-resolution image, and the categories of the contours include the target object category and the non-target object category (such as Figure 2 the vehicles and houses in Figure 2 ). Further, the identified contours are classified to determine whether they belong to the target object category (such as
[0092] the pedestrians in
[0093] Figure 2 ). Based on the classification, the image regions with contours belonging to the non-target object category are filtered out and not processed further. Instead, the image regions with contours identified as the target object category are retained for more precise detection and processing in the subsequent steps.
[0092] In an alternative embodiment, step S102 specifically includes steps S1021 - S1023:
[0093] Step S1021, identify the contours of each target from the low-resolution image through a pre-trained segmentation model.
[0094] The pre-trained segmentation model is used to process the low-resolution image. The segmentation model can identify various contours in the image. By analyzing the pixel features in the image, the segmentation model identifies the contours of each target, including multiple categories including the target object. The segmentation model outputs the identified contour information, including the boundary coordinates and category identification of each target, providing the basic data for the subsequent classification and screening steps.
[0095] Step S1022, classify the contours of each target to obtain the contour classification result.
[0096] The identified contour information is input into a classifier to classify the contours of each target, obtaining the contour classification result. The classification result of each target represents its category ("target object" or "non-target object"), providing the necessary information for the subsequent filtering step.
[0097] In an alternative embodiment, classifying the contours of each target to obtain the contour classification result includes:
[0098] For each target among the respective targets, based on the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target, determine the coordinates of the centroid of the target in the pixel coordinate system.
[0099] In the process of classifying the contours, first determine the centroid position of each target. The centroid refers to the balance point when the mass of each part of the target is evenly distributed. For each target, its contour consists of a series of pixel points. Therefore, in this embodiment, according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target, the coordinates of the centroid of the target in the pixel coordinate system are determined. Each pixel point has a corresponding coordinate in the image. These coordinates constitute the pixel coordinate system. The coordinates of the centroid of the target can be calculated by the following formula:
[0100]
[0101] In the formula, is the total number of pixels of the contour; is the th coordinate of the pixel point.
[0102] For each target among the various targets, with the coordinates of the centroid of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, each pixel point on the contour of the target is transformed from the pixel coordinate system to the polar coordinate system, and the coordinates of each pixel point on the contour of the target in the polar coordinate system are obtained.
[0103] Transform the contour of each target from the pixel coordinate system to the polar coordinate system to facilitate subsequent classification processing. Specifically, for each target, use the centroid coordinates of the target as the origin of the polar coordinate system. Each point in the polar coordinate system is represented by the distance from the origin and the angle with the horizontal axis. For each pixel point on the contour, calculate its coordinates in the polar coordinate system according to the following formula
[0104]
[0105]
[0106] Through the above coordinate transformation, the shape features of the contour can be more easily captured by the subsequent classification model, especially for the extraction of rotation-invariant features.
[0107] Through a pre-trained classifier, process the coordinates of each pixel point on the contour of each of the various targets in the polar coordinate system to determine whether the category of the contour of each of the various targets is the target object category.
[0108] The contour of each target will be classified using a pre-trained classifier. The pixel point coordinates of each target in the polar coordinate system Input it into the classifier. Use the classifier to analyze each target polar coordinate data, extract the contour features ( Figure 2 The waveform in it represents the contour features of the target in the polar coordinate system), and determine whether the contours of each target belong to the target object category. The classifier can use technologies such as deep learning. After being trained with a large amount of data, it can effectively distinguish target objects from non-target objects. The output of the classifier will indicate the category of the contour of each target.
[0109] Step S1023, filter out the image regions of the targets whose contours are of non-target object categories, and retain the image regions of the targets whose contours are of target object categories.
[0110] Traverse all the recognized contours and check the classification results of each contour one by one. Mark the image regions corresponding to all the targets recognized as non-target object categories and determine them as the regions that do not need to be processed subsequently. Eliminate the image regions of the targets whose contours are of non-target object categories from the subsequent processing flow. The pixel data of the image regions of the targets whose contours are of non-target object categories will no longer be considered, avoiding repeated processing and calculation of irrelevant regions. In contrast to the non-target object category, retain the image regions of the targets whose contours are of target object categories. The image regions of the targets whose contours are of target object categories have a higher confidence level, and the computing resources will be concentrated on these regions to improve the efficiency and accuracy of detection. It can be understood that Figure 2 "Discarding the part of the region similar to a person" in it is equivalent to retaining the image regions of the targets whose contours are of target object categories.
[0111] Step S103, restore the resolution of the low-resolution image with the image regions whose contours are of target object categories to the resolution of the original high-resolution image.
[0112] In this step, first determine the target regions in the low-resolution image that belong to the target object category. Next, in order to perform more accurate subsequent detection and analysis, restore the resolution of these retained low-resolution image regions.
[0113] In an optional implementation manner, step S103 specifically includes:
[0114] Determine the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image. According to the scaling ratio, restore the resolution of the low-resolution image with the image regions whose contours are of target object categories to the resolution of the original high-resolution image; or according to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system, map the coordinates of the pixel points on the contour of the target object in the low-resolution image to the high-resolution image.
[0115] Before performing resolution restoration, determine the scaling ratio. The scaling ratio refers to the ratio used when converting the original high-resolution image into a low-resolution image. Assume that the resolution of the original high-resolution image is (width) x (height), and the resolution of the low-resolution image is (width) x (height), then the scaling ratio can be expressed as:
[0116]
[0117] According to the scaling ratio, enlarge the image regions in the low-resolution image whose retained contours are of the target object category. This can be achieved through an interpolation algorithm (such as bilinear interpolation or cubic interpolation) to generate high-resolution image regions.
[0118] Map the relationship between the pixel coordinates of the original high-resolution image and the pixel coordinates of the low-resolution image. The mapping relationship can be established in the following way: Set the coordinates of a certain pixel point in the original high-resolution image as ( Y ), and the corresponding coordinates in the low-resolution image are ( x y ).
[0119] Based on the scaling ratio, determine how the coordinates in the low-resolution image are mapped to the high-resolution image, that is:
[0120] According to the above coordinate mapping relationship, convert the coordinates of each pixel point of the contour of the target object in the low-resolution image to the coordinates of the original high-resolution image to obtain the restored high-resolution image, ensuring that the contour of the target object can accurately reflect its position in the low-resolution image in the high-resolution image.
[0121] Step S104, slice the high-resolution image after resolution restoration through a sliding window to obtain multiple image regions.
[0122] The sliding window technique is used to extract specific regions in the image for subsequent analysis and processing. In this step, the sliding window is applied to the restored high-resolution image to generate multiple small image regions (slices), and the slices are used for object detection and feature extraction.
[0123] In an alternative embodiment, after restoring the resolution of the low-resolution image that retains the image regions with contours of the target object class to the resolution of the original high-resolution image, the following steps are further included: Based on the filtered image regions with contours of non-target object classes, perform mask marking on the high-resolution image after resolution restoration to obtain a high-resolution image including a masked region and a remaining region, where the masked region corresponds to the filtered image regions with contours of non-target object classes.
[0124] The purpose of mask marking is to clearly distinguish in the high-resolution image the known image regions with contours of non-target object classes. This embodiment aims to effectively concentrate computing resources on important regions through mask processing, avoid repeated detection of known regions, and thus improve the overall detection efficiency.
[0125] Based on the identified image regions with contours of non-target object classes, generate mask marks for these regions in the high-resolution image. Specifically, create a binary mask image for each region. In the high-resolution image, the masked region is marked as "1" or "true", while the remaining region is marked as "0" or "false". Superimpose the generated masked region on the high-resolution image to form a high-resolution image including the masked region and the remaining region. The masked region will clearly identify the parts that do not need to be detected again in subsequent processing, which is equivalent to informing different types of detectors in the subsequent steps to eliminate the detection tasks for the masked region.
[0126] After mask marking, the high-resolution image will be divided into two parts. The first part is the masked region ( Figure 2 the white region in it). The masked region corresponds to the filtered image regions with contours of non-target object classes during the reverse segmentation process. The second part is the remaining region except the masked region ( Figure 2 the region except the white region in it), that is, the part of the high-resolution image that is not mask-marked. The remaining region contains unknown targets or targets to be detected, and in this embodiment, further analysis and detection will be performed on the remaining region in subsequent steps.
[0127] Slice the high-resolution image after resolution restoration through a sliding window to obtain multiple image regions, including: Slice the remaining region in the mask-marked high-resolution image through a sliding window to obtain multiple image regions.
[0128] Set the size of the sliding window according to the expected scale and characteristics of the target. The size of the window should be able to adapt to the size of the target to ensure that sufficient detailed information can be captured during the slicing process. The sliding window starts from the upper left corner of the high-resolution image and moves step by step to the right and down according to the set stride (the distance the window moves each time). The movement of the window can be overlapping or non-overlapping, depending on the detection requirements and the characteristics of the target. Whenever the window moves to a new position in the high-resolution image, extract the image area covered by the window to form a slice. The obtained slice is the input data for subsequent target detection, focusing on the remaining area not marked by the mask.
[0129] By slicing the masked high-resolution image, multiple small image areas containing potential targets are obtained. In this embodiment, the sliding window technique is applied to slice the remaining area, enabling the focused detection of the high-resolution image to concentrate computing resources on the remaining area, ensuring efficient and accurate identification and processing of potential targets.
[0130] Step S105, determine the quality of each of the multiple image areas.
[0131] The purpose of quality assessment is to classify image areas of different qualities, that is, to perform parallel detection using different types of detectors. A classification model based on the attention mechanism can be selected to construct a quality assessment platform. Its input is a single slice, and through multi-layer feature extraction and pooling operations, a quality score vector is output. The classification model is optimized through offline training. The training data includes slice samples labeled with "high quality" (clear, low noise, distinct texture) and "low quality" (blurry, high noise, occluded) labels to learn the discriminative features of different quality levels.
[0132] In an alternative embodiment, determining the quality of each of the multiple image areas includes: determining the quality level of each of the multiple image areas according to the clarity, noise level, and texture complexity of each of the multiple image areas.
[0133] In this embodiment, the quality quantization indicators for the quality of each of the multiple image areas can include clarity, noise level, and texture complexity. Among them, clarity can be obtained by calculating the mean value of the image gradient amplitude, the noise level can be obtained by statistical analysis of the local area variance, and the complex texture degree can be calculated based on the image entropy value. The means for determining the quality quantization indicators are not limited in this embodiment of the present invention.
[0134] After obtaining the above quality quantization metrics, normalize and weight each metric for fusion to generate a comprehensive quality score, and set a quality assessment threshold. When the comprehensive quality score is greater than or equal to the quality assessment threshold, determine that the quality level of the image region is high quality and assign it to the fast detector; when the comprehensive quality score is less than the quality assessment threshold, determine that the quality level of the image region is low quality and assign it to the slow detector.
[0135] Step S106: Perform object detection on different quality image regions through detectors with different structures to obtain the object detection results of each of the multiple image regions.
[0136] In this embodiment, the fast detector is a detector for high-quality image regions and can adopt a lightweight single-stage detection model (such as the YOLO series, SSD). Its network structure is streamlined (for example, MobileNetV3 is used as the backbone network), and high-speed inference is achieved through global feature extraction and dense anchor box prediction. The fast detector has a high recall rate for clear and unoccluded slices, and the computational cost is significantly lower than that of complex models. The slow detector is a detector for low-quality image regions and can adopt a large-scale convolutional neural network or Transformer, which can effectively process blurred, occluded, or small target slices. Please refer to Figure 2 , perform parallel object detection on different quality image regions through the slow detector and the fast detector to obtain the object detection results of each of the multiple image regions.
[0137] In an alternative implementation, performing object detection on different quality image regions through detectors with different structures includes:
[0138] For each of the multiple image regions, when the quality level of the image region is low quality, perform object detection on the image region through a pre-trained slow object detector, and when the quality level of the image region is high quality, perform object detection on the image region through a pre-trained fast object detector.
[0139] The slow object detector and the fast object detector satisfy at least one of the following: the inference speed of the slow object detector is less than the inference speed of the fast object detector; the number of model parameters of the slow object detector is greater than the number of model parameters of the fast object detector; the object feature extraction ability of the slow object detector is greater than the object feature extraction ability of the fast object detector; the structural complexity of the slow object detector is greater than the structural complexity of the fast object detector.
[0140] As mentioned above, in the field of image detection, object detection has always faced the trade-off between accuracy and speed. High-precision detectors (such as the slow detector in this embodiment) usually adopt deep neural networks and multi-stage detection frameworks. Although they can effectively process regions with blur, low contrast, or noise interference, their computational cost is high and it is difficult to meet the real-time requirements. On the other hand, lightweight detectors (such as the fast detector in this embodiment) have a fast inference speed, but their detection accuracy for complex scenes (such as regions with complex textures or blurred objects) drops significantly, and the false negative rate is relatively high.
[0141] Through a quality grading strategy, the present invention combines the advantages of the two types of detectors, dynamically allocates detection tasks, and successfully solves the long-existing "accuracy-speed" contradiction in the field of object detection. The slow object detector and the fast object detector applied in this embodiment are respectively used for high-precision detection tasks and lightweight detection tasks. Their design goals are different. Specifically, the slow object detector has a slower inference speed, while the fast object detector has a faster inference speed; the slow object detector has a larger number of model parameters, its model is more complex and can handle more features, while the fast object detector has a smaller number of model parameters and a relatively simple structure, which is suitable for fast inference; the slow object detector has a stronger ability to extract object features and can extract more complex and detailed features, while the fast object detector has a weaker ability to extract object features and is suitable for handling objects with obvious features; the slow object detector has a higher structural complexity and adopts a deeper or more complex network structure, while the fast object detector has a lower structural complexity and adopts a lightweight model design.
[0142] Step S107, obtaining the object detection result of the original high-resolution image according to the object detection results of the respective multiple image regions.
[0143] Please refer to Figure 2 , after obtaining the detection results of the slow detector and the fast detector respectively, unify and splice the recognized information in the image regions with non-object categories as the contour described above to obtain the object detection result of the high-resolution image.
[0144] In an optional implementation manner, step S107 specifically includes steps S1071 - S1072:
[0145] Step S1071, fusing the object detection results of the respective multiple image regions.
[0146] Integrate the output results from the slow detector and the fast detector to form a unified detection result set. The detection result set contains the target information of all detected target objects in the remaining area of the high-resolution image, including their positions (bounding box coordinates), category information, and corresponding confidence levels in the original high-resolution image.
[0147] Step S1072, remove the redundant target object position detection frames in the fused target object detection results to obtain the target object detection results of the original high-resolution image.
[0148] In the fused target object detection results, the overlap degree between detection frames can be compared by setting a threshold. When the overlap degree of the detection frames exceeds the set threshold, the frame with the highest confidence level will be retained, and other overlapping frames will be removed. This effectively eliminates the target objects detected repeatedly, ensuring that each target object is only recognized once. The final target object detection results will contain the target information of all unique target objects in the original high-resolution image. This information can be presented in the form of bounding boxes, marking the positions and confidence levels of each target object, providing clear and accurate detection results.
[0149] By reducing the resolution of the original high-resolution image, the present invention can quickly process large-scale image data and improve the processing efficiency. In the low-resolution image, a screening mechanism is used to effectively filter out the regions of non-target object categories and retain the contours of target object targets, significantly reducing the redundant calculations in subsequent detections and concentrating resources on the detection of important targets. Secondly, after restoring the retained target object regions to the original high resolution, the image is carefully analyzed through the sliding window slicing technology to ensure the accurate positioning of the target object targets. By evaluating the quality of multiple image regions and combining detectors with different structures for targeted object detection, both the detection speed and accuracy are taken into account. When dealing with complex scenes, the present invention can effectively handle image regions with different qualities and features, improve the recall rate of small targets, provide efficient and accurate target object detection outputs, and greatly improve the detection ability of traditional methods in multi-target and complex scenes.
[0150] Figure 3 is the structural block diagram of a high-resolution image target object detection device based on reverse segmentation provided by an embodiment of the present invention, as Figure 3 shown, including:
[0151] A low-resolution processing module 201, configured to reduce the resolution of the original high-resolution image to obtain a low-resolution image;
[0152] A filtering module 202, configured to filter out the image regions with non-target object category contours and retain the image regions with target object category contours from the low-resolution image;
[0153] A restoration module 203, configured to restore the resolution of a low-resolution image with an image region retaining a contour of a target object category to the resolution of the original high-resolution image;
[0154] A slicing module 204, configured to slice the high-resolution image with restored resolution through a sliding window to obtain a plurality of image regions;
[0155] A quality analysis module 205, configured to determine the quality of each of the plurality of image regions;
[0156] A target detection module 206, configured to perform target detection on image regions of different qualities through detectors of different structures to obtain target object detection results of each of the plurality of image regions;
[0157] A result output module 207, configured to obtain a target object detection result of the original high-resolution image according to the target object detection results of each of the plurality of image regions.
[0158] In an optional implementation manner, the filtering module includes:
[0159] A contour recognition sub-module, configured to recognize contours of each target from the low-resolution image through a pre-trained segmentation model;
[0160] A classification sub-module, configured to classify the contours of each target to obtain a contour classification result;
[0161] A filtering sub-module, configured to filter out image regions of targets with contours of non-target object categories and retain image regions of targets with contours of target object categories.
[0162] In an optional implementation manner, the classification sub-module includes:
[0163] A coordinate determination unit, configured to, for each target among the targets, determine the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target;
[0164] A coordinate conversion unit, configured to, for each target among the targets, with the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system;
[0165] A coordinate processing unit, configured to process the coordinates of each pixel point on the contour of each target in the polar coordinate system through a pre-trained classifier to determine whether the category of the contour of each target is a target object category.
[0166] In an alternative embodiment, the reduction module includes:
[0167] A scaling sub-module for determining a scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and according to the scaling ratio, restoring the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image; or
[0168] A mapping sub-module for mapping the coordinates of the pixel points on the contour of the target object in the low-resolution image to the high-resolution image according to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system.
[0169] In an alternative embodiment, it further includes:
[0170] A mask module for masking and marking the high-resolution image after resolution restoration according to the image area whose contour is the non-target object category that has been filtered, to obtain a high-resolution image including a mask area and a remaining area, where the mask area corresponds to the image area whose contour is the non-target object category that has been filtered;
[0171] The slicing module includes:
[0172] A slicing sub-module for slicing the remaining area in the high-resolution image after mask marking through a sliding window to obtain a plurality of image areas.
[0173] In an alternative embodiment, the quality analysis module includes:
[0174] A quality level determination sub-module for determining the quality level of each of the plurality of image areas according to the clarity, noise level, and texture complexity of each of the plurality of image areas;
[0175] The target detection module includes:
[0176] A classification and detection sub-module for performing target detection on each of the plurality of image areas through a pre-trained slow target detector when the quality level of the image area is low-quality, and performing target detection on the image area through a pre-trained fast target detector when the quality level of the image area is high-quality;
[0177] The following at least one is satisfied between the slow target detector and the fast target detector:
[0178] The inference speed of the slow object detector is less than that of the fast object detector;
[0179] The number of model parameters of the slow object detector is greater than that of the fast object detector;
[0180] The ability of the slow object detector to extract object features is greater than that of the fast object detector;
[0181] The structural complexity of the slow object detector is greater than that of the fast object detector.
[0182] In an alternative embodiment, the result output module includes:
[0183] A fusion sub-module for fusing the object detection results of the multiple image regions;
[0184] A removal sub-module for removing redundant object position detection frames in the fused object detection results to obtain the object detection results of the original high-resolution image.
[0185] The embodiments of the present disclosure also provide an electronic device, including a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements each process of the above-mentioned embodiments of the high-resolution image object detection method based on reverse segmentation and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0186] The embodiments of the present application also provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements each process of the above-mentioned embodiments of the high-resolution image object detection method based on reverse segmentation and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0187] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, devices, electronic devices, and storage media. Therefore, the embodiments of the present invention can take the form of completely hardware embodiments, completely software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0188] Embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods and apparatuses according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0189] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0190] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article, or terminal device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article, or terminal device including the said elements.
[0191] The above has introduced in detail the method and device for detecting target objects in high-resolution images based on reverse segmentation. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for detecting target objects in high-resolution images based on reverse segmentation, characterized in that, Including: Reducing the resolution of the original high-resolution image to obtain a low-resolution image; Filtering out the image regions with contours of non-target object categories from the low-resolution image and retaining the image regions with contours of target object categories, where all the image regions corresponding to the targets identified as non-target object categories are regions that do not need to be processed subsequently; Restoring the resolution of the low-resolution image with the image regions having contours of target object categories retained to the resolution of the original high-resolution image; Performing mask marking on the high-resolution image with the resolution restored according to the filtered image regions with contours of non-target object categories to obtain a high-resolution image including a mask region and a remaining region, where the mask region corresponds to the filtered image regions with contours of non-target object categories; Slicing the high-resolution image with the resolution restored through a sliding window to obtain a plurality of image regions, including: slicing the remaining region in the mask-marked high-resolution image through a sliding window to obtain a plurality of image regions; Determining the quality of each of the plurality of image regions, including: determining the quality level of each of the plurality of image regions according to the clarity, noise level, and texture complexity of each of the plurality of image regions; Performing target detection on the image regions of different qualities through detectors of different structures to obtain the target object detection results of each of the plurality of image regions, where performing target detection on the image regions of different qualities through detectors of different structures includes: for each image region in the plurality of image regions, when the quality level of the image region is low quality, performing target detection on the image region through a pre-trained slow target detector, and when the quality level of the image region is high quality, performing target detection on the image region through a pre-trained fast target detector; Obtaining the target object detection result of the original high-resolution image according to the target object detection results of each of the plurality of image regions.
2. The method according to claim 1, wherein Filtering out the image regions with contours of non-target object categories from the low-resolution image and retaining the image regions with contours of target object categories, including: Identifying the contours of each target from the low-resolution image through a pre-trained segmentation model; Classifying the contours of each target to obtain a contour classification result; Filtering out the image regions of the targets with contours of non-target object categories and retaining the image regions of the targets with contours of target object categories.
3. The method according to claim 2, wherein Classifying the contours of each target to obtain a contour classification result, including: For each target among the targets, determining the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target; For each target among the targets, taking the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, and converting each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system; Process the coordinates of each pixel point on the contour of each of the said targets in the polar coordinate system through a pre-trained classifier to determine whether the category of the contour of each of the said targets is the target object category.
4. The method according to claim 1, wherein Restore the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image, including: Determine the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and according to the scaling ratio, restore the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image; or According to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system, map the coordinates of the pixel points on the contour of the target object on the low-resolution image to the high-resolution image.
5. The method according to claim 1, characterized in that, At least one of the following is satisfied between the slow target detector and the fast target detector: The inference speed of the slow target detector is less than the inference speed of the fast target detector; The number of model parameters of the slow target detector is greater than the number of model parameters of the fast target detector; The target object feature extraction ability of the slow target detector is greater than the target object feature extraction ability of the fast target detector; The structural complexity of the slow target detector is greater than the structural complexity of the fast target detector.
6. The method according to any one of claims 1-5, characterized in that, Obtain the target object detection result of the original high-resolution image according to the target object detection results of each of the multiple image areas, including: Fuse the target object detection results of each of the multiple image areas; Remove the redundant target object position detection frames in the fused target object detection result to obtain the target object detection result of the original high-resolution image.
7. The high-resolution image target object detection device based on reverse segmentation, characterized in that, Including: A low-resolution processing module for reducing the resolution of the original high-resolution image to obtain a low-resolution image; A filtering module for filtering out the image areas whose contours are non-target object categories from the low-resolution image and retaining the image areas whose contours are target object categories, where all the image areas corresponding to the targets identified as non-target object categories are areas that do not need to be processed subsequently; A restoration module for restoring the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image; masking the high-resolution image after resolution restoration according to the filtered image areas whose contours are non-target object categories to obtain a high-resolution image including a masked area and a remaining area, where the masked area corresponds to the filtered image areas whose contours are non-target object categories; A slicing module for slicing the high-resolution image after resolution restoration through a sliding window to obtain multiple image areas, including: slicing the remaining area in the masked high-resolution image through a sliding window to obtain multiple image areas; A quality analysis module, configured to determine the quality of each of the multiple image regions, including: determining the quality level of each of the multiple image regions according to the sharpness, noise level, and texture complexity of each of the multiple image regions; A target detection module, configured to perform target detection on image regions of different qualities through detectors of different structures, to obtain the target object detection results of each of the multiple image regions. Among them, performing target detection on image regions of different qualities through detectors of different structures includes: for each image region among the multiple image regions, when the quality level of this image region is low quality, performing target detection on this image region through a pre-trained slow target detector, and when the quality level of this image region is high quality, performing target detection on this image region through a pre-trained fast target detector; A result output module, configured to obtain the target object detection result of the original high-resolution image according to the target object detection results of each of the multiple image regions.
8. An electronic device, characterized in that, Comprising: A processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Pedestrian detection and recognition system, method and device and computer readable storage medium
CN111274991A
Brake beam falling detection method based on deep learning
CN112634242A
Image acquisition method based on human eye attention perception mechanism
CN116664821A
One-billion-pixel-level target detection method and device, medium and electronic equipment
CN117893735A