High-resolution image target object detection method and device based on reverse segmentation
By performing reverse segmentation and sliding window slicing technology on high-resolution images, combined with detectors with different structures, the problem of insufficient recall of traditional object detection methods in small targets and complex scenarios is solved, and efficient and accurate object detection is achieved.
Patent Information
- Application Number
- CN202510551622.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-29
AI Technical Summary
When traditional object detection methods face small targets and complex scenarios, the recall rate is insufficient, making it difficult to balance detection speed and accuracy.
The high-resolution image target object detection method based on reverse segmentation is adopted, and the resolution of the original high-resolution image is reduced, the non-target area is filtered, the target area is retained, and detailed detection is performed through sliding windows and detectors of different structures in subsequent steps.
It significantly improves the recall rate of small targets, improves the ability to take into account both detection speed and accuracy, and improves the detection ability of traditional methods in multiple targets and complex scenarios.
Smart Images

Figure CN120070875A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method and device for detecting target objects in high-resolution images based on reverse segmentation. Background Art
[0002] In modern society, with the rapid development of surveillance technology and image processing technology, image analysis has become increasingly widely used in fields such as security surveillance, intelligent transportation, and autonomous driving. Especially in target detection, accurately identifying and locating targets is crucial for ensuring public safety and traffic management.
[0003] Traditional target detection methods rely on a two-stage processing flow. First, region screening is performed on low-resolution images, and then precise detection is carried out on high-resolution images. However, when dealing with small targets, this method often has the problem of insufficient recall rate. Since small targets are not easily identified accurately in low-resolution images, they are missed in subsequent detection stages, thus affecting the overall detection effect. In addition, although the progressive magnification target detection method can gradually magnify the image through a pyramid structure, if small targets are not detected in the initial stage, subsequent magnification cannot make up for this defect. Therefore, related technologies are difficult to balance detection speed and accuracy when dealing with scenarios with multiple targets, large resolution differences, or complex target types. Summary of the Invention
[0004] The present invention provides a method and device for detecting target objects in high-resolution images based on reverse segmentation, aiming to solve the problems existing in the above background art.
[0005] To solve the above technical problems, the present invention is implemented as follows: The present invention provides a method for detecting target objects in high-resolution images based on reverse segmentation, including: Reducing the resolution of the original high-resolution image to obtain a low-resolution image; Filtering out image regions with non-target object contours from the low-resolution image and retaining image regions with target object contours; Restoring the resolution of the low-resolution image with the retained image regions having target object contours to the resolution of the original high-resolution image; Slicing the high-resolution image with restored resolution through a sliding window to obtain a plurality of image regions; Determining the quality of each of the plurality of image regions; Performing target detection on image regions of different qualities through detectors with different structures to obtain target object detection results for each of the plurality of image regions; Obtain the object detection result of the original high-resolution image based on the object detection results of the respective multiple image regions.
[0006] Optionally, filtering out the image regions with non-target object categories in their contours from the low-resolution image and retaining the image regions with target object categories in their contours includes: Identify the contours of each target from the low-resolution image through a pre-trained segmentation model; Classify the contours of each target to obtain a contour classification result; Filter out the image regions of the targets with non-target object categories in their contours and retain the image regions of the targets with target object categories in their contours.
[0007] Optionally, classifying the contours of each target to obtain a contour classification result includes: For each target among the respective targets, determine the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target; For each target among the respective targets, with the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system; Through a pre-trained classifier, process the coordinates of each pixel point on the contour of each target in the polar coordinate system to determine whether the category of the contour of each target is a target object category.
[0008] Optionally, restoring the resolution of the low-resolution image with image regions having contours of target object categories to the resolution of the original high-resolution image includes: Determine the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and according to the scaling ratio, restore the resolution of the low-resolution image with image regions having contours of target object categories to the resolution of the original high-resolution image; or According to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system, map the coordinates of the pixel points on the contour of the target object in the low-resolution image to the high-resolution image.
[0009] Optionally, after restoring the resolution of the low-resolution image with image regions having contours of target object categories to the resolution of the original high-resolution image, it further includes: Mask the high-resolution image after resolution restoration according to the image regions whose filtered contours are non-target object categories, to obtain a high-resolution image including a masked region and a remaining region, where the masked region corresponds to the image regions whose filtered contours are non-target object categories; Slice the high-resolution image after resolution restoration through a sliding window to obtain a plurality of image regions, including: Slice the remaining region in the masked high-resolution image through a sliding window to obtain a plurality of image regions.
[0010] Optionally, determining the quality of each of the plurality of image regions includes: Determine the quality level of each of the plurality of image regions according to the clarity, noise level, and texture complexity of each of the plurality of image regions; Perform object detection on image regions of different qualities through detectors of different structures, including: For each image region in the plurality of image regions, when the quality level of the image region is low quality, perform object detection on the image region through a pre-trained slow object detector, and when the quality level of the image region is high quality, perform object detection on the image region through a pre-trained fast object detector; The slow object detector and the fast object detector satisfy at least one of the following: The inference speed of the slow object detector is less than the inference speed of the fast object detector; The number of model parameters of the slow object detector is greater than the number of model parameters of the fast object detector; The target object feature extraction ability of the slow object detector is greater than the target object feature extraction ability of the fast object detector; The structural complexity of the slow object detector is greater than the structural complexity of the fast object detector.
[0011] Optionally, obtaining the object detection result of the original high-resolution image according to the object detection results of each of the plurality of image regions includes: Fuse the object detection results of each of the plurality of image regions; Remove redundant object position detection frames in the fused object detection result to obtain the object detection result of the original high-resolution image.
[0012] In a second aspect, an embodiment of the present disclosure provides a high-resolution image object detection device based on reverse segmentation, including: A low-resolution processing module for reducing the resolution of the original high-resolution image to obtain a low-resolution image; A filtering module, configured to filter out image regions with contours of non-target object categories from the low-resolution image and retain image regions with contours of target object categories; A restoration module, configured to restore the resolution of the low-resolution image retaining the image regions with contours of target object categories to the resolution of the original high-resolution image; A slicing module, configured to slice the high-resolution image with restored resolution through a sliding window to obtain a plurality of image regions; A quality analysis module, configured to determine the quality of each of the plurality of image regions; A target detection module, configured to perform target detection on image regions of different qualities through detectors of different structures to obtain target object detection results for each of the plurality of image regions; A result output module, configured to obtain the target object detection result of the original high-resolution image according to the target object detection results for each of the plurality of image regions.
[0013] Optionally, the filtering module includes: A contour recognition sub-module, configured to recognize the contours of each target from the low-resolution image through a pre-trained segmentation model; A classification sub-module, configured to classify the contours of each target to obtain a contour classification result; A filtering sub-module, configured to filter out image regions of targets with contours of non-target object categories and retain image regions of targets with contours of target object categories.
[0014] Optionally, the classification sub-module includes: A coordinate determination unit, configured to, for each target among the respective targets, determine the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target; A coordinate conversion unit, configured to, for each target among the respective targets, with the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system; A coordinate processing unit, configured to process the coordinates of each pixel point on the contour of each target in the polar coordinate system through a pre-trained classifier to determine whether the category of the contour of each target is the target object category.
[0015] Optionally, the restoration module includes: A scaling sub-module, configured to determine a scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and according to the scaling ratio, restore the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image; or A mapping sub-module, configured to map the coordinates of the pixel points on the contour of the target object on the low-resolution image to the high-resolution image according to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system.
[0016] Optionally, it further includes: A mask module, configured to perform mask marking on the high-resolution image after resolution restoration according to the image area whose contour is the non-target object category that is filtered out, to obtain a high-resolution image including a mask area and a remaining area, where the mask area corresponds to the image area whose contour is the non-target object category that is filtered out; The slicing module includes: A slicing sub-module, configured to slice the remaining area in the high-resolution image after mask marking through a sliding window to obtain a plurality of image areas.
[0017] Optionally, the quality analysis module includes: A quality level determination sub-module, configured to determine the quality level of each of the plurality of image areas according to the clarity, noise level, and texture complexity of each of the plurality of image areas; The target detection module includes: A classification detection sub-module, configured to perform target detection on each of the plurality of image areas through a pre-trained slow target detector when the quality level of the image area is low quality, and perform target detection on the image area through a pre-trained fast target detector when the quality level of the image area is high quality; At least one of the following is satisfied between the slow target detector and the fast target detector: The inference speed of the slow target detector is less than the inference speed of the fast target detector; The number of model parameters of the slow target detector is greater than the number of model parameters of the fast target detector; The target object feature extraction ability of the slow target detector is greater than the target object feature extraction ability of the fast target detector; The structural complexity of the slow target detector is greater than the structural complexity of the fast target detector.
[0018] Optionally, the result output module includes: A fusion sub-module, configured to fuse the target object detection results of the respective multiple image regions; A removal sub-module, configured to remove redundant target object position detection frames in the fused target object detection results to obtain the target object detection result of the original high-resolution image.
[0019] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the high-resolution image target object detection method based on reverse segmentation as described in the first aspect are implemented.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the high-resolution image target object detection method based on reverse segmentation as described in the first aspect are implemented.
[0021] The technical solution provided by the present invention at least brings the following beneficial effects: By reducing the resolution of the original high-resolution image, the present invention can quickly process large-scale image data and improve the processing efficiency. In the low-resolution image, a screening mechanism is adopted to effectively filter out regions with non-target object categories in the contour and retain regions with target object categories in the contour, significantly reducing the redundant calculations in subsequent detections and concentrating resources on the detection of important targets. After restoring the retained regions to the original high resolution, the image is carefully analyzed through the sliding window slicing technique to ensure the accurate positioning of the target object. By evaluating the quality of multiple image regions and combining detectors with different structures for targeted object detection, both the detection speed and accuracy are taken into account. When dealing with complex scenes, the present invention can effectively handle image regions with different qualities and features, improve the recall rate of small targets, provide efficient and accurate target detection outputs, and greatly improve the detection ability of traditional methods in multi-target and complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic diagram of the steps of a high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the overall process of the high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention; Figure 3 It is a structural block diagram of the high-resolution image target object detection device based on reverse segmentation provided by an embodiment of the present invention. Specific embodiments
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] The high-resolution image target object detection method based on reverse segmentation proposed by the present invention aims to significantly improve the target object detection speed while ensuring the detection accuracy of the target object through refined processing of key regions. The core idea of reverse segmentation is to utilize the fast processing ability of low-resolution images to initially filter out non-target object category regions and only retain the image regions containing target object categories. In this way, the subsequent high-resolution detection process can focus on the regions more likely to contain the target, thereby avoiding redundant processing of irrelevant regions and significantly improving the detection efficiency.
[0026] It can be understood that the detection method proposed by the present invention can be effectively applied to the detection scenarios of various types of target objects, including multiple fields such as traffic monitoring, intelligent security, unmanned driving, and industrial monitoring, and can meet the target detection requirements in different scenarios.
[0027] Figure 1 It is a schematic diagram of the steps of the high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention. As Figure 1 shown, it includes: Step S101, reduce the resolution of the original high-resolution image to obtain a low-resolution image.
[0028] The original high-resolution image is the original input image without resolution adjustment, and its resolution is higher than the standard display or processing requirements, such as a resolution ≥ 1920 × 1080 pixels (i.e., full high-definition and above). The original high-resolution image can be reduced to a low-resolution image through scaling operations, such as bilinear interpolation and area sampling methods. The low-resolution image retains the global contour features of the target object while significantly reducing the data volume.
[0029] Step S102: Filter out the image regions with contours of non-target object categories from the low-resolution image, and retain the image regions with contours of target object categories.
[0030] Figure 2 It is a schematic diagram of the overall process of the high-resolution image target object detection method based on reverse segmentation provided by an embodiment of the present invention. Please refer to Figure 2 , first, simulate and generate sparse data from the low-resolution image through Poisson distribution sampling to increase data diversity. This can help the subsequent segmentation model better learn the features of different scenarios and targets, and improve the robustness and generalization ability of the segmentation model. Further, identify all the contours from the low-resolution image, and the categories of the contours include target object categories and non-target object categories (such as Figure 2 the vehicles and houses in Figure 2 ). Further, classify the identified contours to determine whether they belong to the target object category (such as
[0031] the pedestrians in ). Based on the classification, filter out the image regions with contours of non-target object categories and no longer perform subsequent processing. On the contrary, retain the image regions with contours identified as target object categories for more accurate detection and processing in subsequent steps.
[0032] In an alternative embodiment, step S102 specifically includes steps S1021 - S1023:
[0033] Step S1021: Identify the contours of each target from the low-resolution image through a pre-trained segmentation model.
[0034] Process the low-resolution image using a pre-trained segmentation model. The segmentation model can identify various contours in the image. By analyzing the pixel features in the image, the segmentation model identifies the contours of each target, including multiple categories including target objects. The segmentation model outputs the identified contour information, including the boundary coordinates and category identifiers of each target, providing the basic data for subsequent classification and screening steps.
[0035] In an alternative embodiment, classifying the contours of each target to obtain the contour classification result includes: For each of the respective targets, based on the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target, determine the coordinates of the center of gravity of the target in the pixel coordinate system.
[0036] During the process of classifying the contours, first determine the position of the center of gravity of each target. The center of gravity refers to the balance point when the mass of each part of the target is evenly distributed. For each target, its contour consists of a series of pixel points. Therefore, in this embodiment, based on the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target, determine the coordinates of the center of gravity of the target in the pixel coordinate system. Each pixel point has a corresponding coordinate in the image. These coordinates constitute the pixel coordinate system. The coordinates of the center of gravity of the target can be calculated by the following formula:
[0037] In the formula, is the total number of pixels of the contour; is the th coordinate of the pixel point.
[0038] For each of the respective targets, with the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system.
[0039] Convert the contour of each target from the pixel coordinate system to the polar coordinate system to facilitate subsequent classification processing. Specifically, for each target, use the center of gravity coordinates of the target as the origin of the polar coordinate system. Each point in the polar coordinate system is represented by the distance from the origin and the angle with the horizontal axis. For each pixel point on the contour, calculate its coordinates in the polar coordinate system according to the following formula
[0040]
[0041]
[0042] Through the above coordinate transformation, the shape features of the contour can be more easily captured by the subsequent classification model, especially for the extraction of rotation-invariant features.
[0043] Using a pre-trained classifier, process the coordinates of each pixel point on the contour of each of the said targets in the polar coordinate system to determine whether the category of the contour of each of the said targets is the target object category.
[0044] A pre-trained classifier will be used to classify the contours of each target. The pixel point coordinates of each target in the polar coordinate system will be input into the classifier. The classifier analyzes the polar coordinate data of each target, extracts contour features ( Figure 2 the waveform in it represents the contour features of the target in the polar coordinate system), and determines whether the contour of each target belongs to the target object category. The classifier can use technologies such as deep learning, and through training with a large amount of data, effectively distinguish target objects from non-target objects. The output of the classifier will indicate the category of the contour of each target.
[0045] Step S1023, filter out the image regions of the targets whose contours are of non-target object categories, and retain the image regions of the targets whose contours are of target object categories.
[0046] Traverse all the identified contours and check the classification results of each contour one by one. Mark the image regions corresponding to all the targets identified as non-target object categories and determine them as regions that do not need to be processed subsequently. Eliminate the image regions of the targets whose contours are of non-target object categories from the subsequent processing flow. The pixel data of the image regions of the targets whose contours are of non-target object categories will no longer be considered, avoiding repeated processing and calculation of irrelevant regions. In contrast to non-target object categories, retain the image regions of the targets whose contours are of target object categories. The image regions of the targets whose contours are of target object categories have a higher confidence level, and computing resources will be concentrated on these regions to improve the efficiency and accuracy of detection. It can be understood that Figure 2 "discarding the parts of regions similar to humans" in it is equivalent to retaining the image regions of the targets whose contours are of target object categories.
[0047] Step S103, restore the resolution of the low-resolution image with the image regions whose contours are of target object categories to the resolution of the original high-resolution image.
[0048] In this step, first determine the target regions in the low-resolution image that belong to the target object category. Next, in order to perform more precise subsequent detection and analysis, restore the resolution of these retained low-resolution image regions.
[0049] In an optional implementation manner, step S103 specifically includes: Determine the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image. According to the scaling ratio, restore the resolution of the low-resolution image with the image area whose contour is the target object category to the resolution of the original high-resolution image; or according to the coordinate mapping relationship between the pixel coordinates of the original high-resolution image in the pixel coordinate system and the pixel coordinates of the low-resolution image in the pixel coordinate system, map the coordinates of the pixel points on the contour of the target object on the low-resolution image to the high-resolution image.
[0050] Before performing resolution restoration, determine the scaling ratio. The scaling ratio refers to the ratio used when converting the original high-resolution image into a low-resolution image. Assume that the resolution of the original high-resolution image is (width) x (height), and the resolution of the low-resolution image is (width) x (height), then the scaling ratio can be expressed as:
[0051] According to the scaling ratio, perform magnification processing on the image area in the low-resolution image whose contour is the target object category. It can be achieved through an interpolation algorithm (such as bilinear interpolation or cubic interpolation) to generate a high-resolution image area.
[0052] Map the relationship between the pixel coordinates of the original high-resolution image and the pixel coordinates of the low-resolution image. The mapping relationship can be established in the following way: Set the coordinates of a certain pixel point in the original high-resolution image as ( Y ), and the corresponding coordinates in the low-resolution image are ( x y ).
[0053] Through the scaling ratio, determine how the coordinates in the low-resolution image are mapped to the high-resolution image, that is:
[0054] According to the above coordinate mapping relationship, convert the coordinates of each pixel point on the contour of the target object in the low-resolution image to the coordinates of the original high-resolution image to obtain the restored high-resolution image, ensuring that the contour of the target object in the high-resolution image can accurately reflect its position in the low-resolution image.
[0055] Step S104, slice the high-resolution image after resolution restoration through a sliding window to obtain multiple image areas.
[0056] The sliding window technique is used to extract specific regions in an image for subsequent analysis and processing. In this step, the sliding window is applied to the restored high-resolution image to generate multiple small image regions (slices), which are used for object detection and feature extraction.
[0057] In an alternative embodiment, after restoring the resolution of the low-resolution image that retains the image regions with contours of the target object class to the resolution of the original high-resolution image, it further includes: masking the restored high-resolution image according to the image regions with contours of non-target object classes that have been filtered out, to obtain a high-resolution image including a masked region and a remaining region, where the masked region corresponds to the image regions with contours of non-target object classes that have been filtered out.
[0058] The purpose of the masking is to clearly distinguish, in the high-resolution image, the known image regions with contours of non-target object classes. This embodiment aims to effectively concentrate computing resources on important regions through the masking process, avoid repeated detection of known regions, and thus improve the overall detection efficiency.
[0059] According to the identified image regions with contours of non-target object classes, generate masking marks for these regions in the high-resolution image, where a binary mask image is created for each region. In the high-resolution image, the masked region is marked as "1" or "true", while the remaining region is marked as "0" or "false". Superimpose the generated masked region on the high-resolution image to form a high-resolution image including the masked region and the remaining region. The masked region will clearly identify the parts that do not need to be detected again in subsequent processing, which is equivalent to informing the subsequent detectors of different classes to exclude the detection tasks for the masked region.
[0060] After the masking, the high-resolution image will be divided into two parts. The first part is the masked region ( Figure 2 the white region in it). The masked region corresponds to the image regions with contours of non-target object classes that have been filtered out during the reverse segmentation process. The second part is the remaining region except the masked region ( Figure 2 the region except the white region in it), that is, the part of the high-resolution image that has not been masked. The remaining region contains unknown targets or targets to be detected, and this embodiment will perform further analysis and detection on the remaining region in subsequent steps.
[0061] Slice the restored high-resolution image through the sliding window to obtain multiple image regions, including: slicing the remaining region in the masked high-resolution image through the sliding window to obtain multiple image regions.
[0062] Set the size of the sliding window according to the expected scale and characteristics of the target. The size of the window should be able to adapt to the size of the target to ensure that sufficient detailed information can be captured during the slicing process. The sliding window starts from the upper left corner of the high-resolution image and moves step by step to the right and down according to the set stride (the distance the window moves each time). The movement of the window can be overlapping or non-overlapping, depending on the detection requirements and the characteristics of the target. Whenever the window moves to a new position in the high-resolution image, extract the image area covered by the window to form a slice. The obtained slice is the input data for subsequent target detection, focusing on the remaining area not marked by the mask.
[0063] By slicing the high-resolution image after mask marking, multiple small image areas containing potential targets are obtained. In this embodiment, the sliding window technique is applied to slice the remaining area, enabling the focused detection of the high-resolution image to concentrate computing resources on the remaining area, ensuring efficient and accurate identification and processing of potential targets.
[0064] Step S105, determine the quality of each of the multiple image areas.
[0065] The purpose of quality assessment is to classify image areas of different qualities, that is, to perform parallel detection using different types of detectors. A classification model based on the attention mechanism can be selected to build a quality assessment platform. Its input is a single slice, and through multi-layer feature extraction and pooling operations, a quality score vector is output. The classification model is optimized through offline training. The training data includes slice samples labeled with "high quality" (clear, low noise, distinct texture) and "low quality" (blurry, high noise, occluded) labels to learn the discriminative features of different quality levels.
[0066] In an alternative embodiment, determining the quality of each of the multiple image areas includes: determining the quality level of each of the multiple image areas according to the clarity, noise level, and texture complexity of each of the multiple image areas.
[0067] In this embodiment, the quality quantization indicators for the quality of each of the multiple image areas may include clarity, noise level, and texture complexity. Among them, clarity can be obtained by calculating the mean value of the image gradient amplitude, the noise level can be obtained by local area variance statistics, and the complex texture degree can be calculated based on the image entropy value. The means for determining the quality quantization indicators are not limited in the embodiments of the present invention.
[0068] After obtaining the above quality quantification indicators, normalize and weighted fuse each indicator to generate a comprehensive quality score, and set a quality evaluation threshold. When the comprehensive quality score is greater than or equal to the quality evaluation threshold, determine that the quality level of the image region is high quality and assign it to the fast detector; when the comprehensive quality score is less than the quality evaluation threshold, determine that the quality level of the image region is low quality and assign it to the slow detector.
[0069] Step S106, perform object detection on different quality image regions through detectors with different structures to obtain the object detection results of each of the multiple image regions.
[0070] In this embodiment, the fast detector is a detector for high-quality image regions, and a lightweight single-stage detection model (such as the YOLO series, SSD) can be used. Its network structure is concise (for example, MobileNetV3 is used as the backbone network), and through global feature extraction and dense anchor box prediction, high-speed inference is achieved. The fast detector has a high recall rate for clear and unoccluded slices, and the computational cost is significantly lower than that of complex models. The slow detector is a detector for low-quality image regions, and a large-scale convolutional neural network or Transformer can be used, which can effectively process blurred, occluded, or small target slices. Please refer to Figure 2 , perform parallel object detection on different quality image regions through the slow detector and the fast detector to obtain the object detection results of each of the multiple image regions.
[0071] In an alternative embodiment, performing object detection on different quality image regions through detectors with different structures includes: For each of the multiple image regions, when the quality level of the image region is low quality, perform object detection on the image region through a pre-trained slow object detector, and when the quality level of the image region is high quality, perform object detection on the image region through a pre-trained fast object detector.
[0072] The slow object detector and the fast object detector satisfy at least one of the following: the inference speed of the slow object detector is less than the inference speed of the fast object detector; the number of model parameters of the slow object detector is greater than the number of model parameters of the fast object detector; the object feature extraction ability of the slow object detector is greater than the object feature extraction ability of the fast object detector; the structural complexity of the slow object detector is greater than the structural complexity of the fast object detector.
[0073] As described above, in the field of image detection, object detection always faces the trade-off between accuracy and speed. High-precision detectors (such as the slow detector in this embodiment) usually adopt deep neural networks and multi-stage detection frameworks. Although they can effectively process areas with blur, low contrast, or noise interference, the computational cost is high and it is difficult to meet the real-time requirements. On the other hand, lightweight detectors (such as the fast detector in this embodiment) have a fast inference speed, but the detection accuracy for complex scenes (such as areas with complex textures or blurred objects) drops significantly, and the false negative rate is relatively high.
[0074] Through a quality grading strategy, the present invention combines the advantages of the two types of detectors, dynamically allocates detection tasks, and successfully solves the long-existing "accuracy-speed" contradiction in the field of object detection. The slow object detector and the fast object detector applied in this embodiment are respectively used for high-precision detection tasks and lightweight detection tasks. Their design goals are different. Specifically, the slow object detector has a slower inference speed, while the fast object detector has a faster inference speed; the slow object detector has a larger number of model parameters, its model is more complex and can handle more features, while the fast object detector has a smaller number of model parameters and a relatively simple structure, which is suitable for fast inference; the slow object detector has a stronger ability to extract object features and can extract more complex and detailed features, while the fast object detector has a weaker ability to extract object features and is suitable for processing objects with obvious features; the slow object detector has a higher structural complexity and adopts a deeper or more complex network structure, while the fast object detector has a lower structural complexity and adopts a lightweight model design.
[0075] Step S107, obtaining the object detection result of the original high-resolution image according to the object detection results of the respective multiple image regions.
[0076] Please refer to Figure 2 , after obtaining the detection results of the slow detector and the fast detector respectively, unify and splice the information identified in the image regions with non-object categories as the contour described above to obtain the object detection result of the high-resolution image.
[0077] In an optional implementation manner, step S107 specifically includes steps S1071 - S1072: Step S1071, fusing the object detection results of the respective multiple image regions.
[0078] Integrate the output results from the slow detector and the fast detector to form a unified detection result set. The detection result set contains all the detected object target information in the remaining regions of the high-resolution image, including its position (bounding box coordinates), category information, and corresponding confidence in the original high-resolution image.
[0079] Step S1072: Remove redundant target object position detection frames from the fused target object detection results to obtain the target object detection results of the original high-resolution image.
[0080] In the fused target object detection results, the overlap degree between detection frames can be compared by setting a threshold. When the overlap degree of the detection frames exceeds the set threshold, the frame with the highest confidence will be retained, and other overlapping frames will be removed. This effectively eliminates the target objects detected repeatedly, ensuring that each target object is only recognized once. The final target object detection results will contain all the unique target object information in the original high-resolution image. This information can be presented in the form of bounding boxes, marking the position and confidence of each target object, providing clear and accurate detection results.
[0081] By reducing the resolution of the original high-resolution image, the present invention can quickly process large-scale image data and improve the processing efficiency. In the low-resolution image, a screening mechanism is used to effectively filter out the regions of non-target object categories and retain the outlines of target object targets, significantly reducing the redundant calculations in subsequent detections and concentrating resources on the detection of important targets. Secondly, after restoring the retained target object regions to the original high resolution, the image is carefully analyzed through the sliding window slicing technology to ensure the accurate positioning of the target object targets. By evaluating the quality of multiple image regions and combining detectors with different structures for targeted object detection, both the detection speed and accuracy are taken into account. When dealing with complex scenes, the present invention can effectively handle image regions with different qualities and features, improve the recall rate of small targets, provide efficient and accurate target object detection outputs, and greatly improve the detection ability of traditional methods in multi-target and complex scenes.
[0082] Figure 3 It is a structural block diagram of a high-resolution image target object detection device based on reverse segmentation provided by an embodiment of the present invention. As Figure 3 shown, it includes: A low-resolution processing module 201, configured to reduce the resolution of the original high-resolution image to obtain a low-resolution image; A filtering module 202, configured to filter out the image regions with outlines of non-target object categories from the low-resolution image and retain the image regions with outlines of target object categories; A restoration module 203, configured to restore the resolution of the low-resolution image with the retained image regions with outlines of target object categories to the resolution of the original high-resolution image; A slicing module 204, configured to slice the high-resolution image after resolution restoration through a sliding window to obtain a plurality of image regions; A quality analysis module 205 for determining the quality of each of the multiple image regions; A target detection module 206 for performing target detection on image regions of different qualities through detectors of different structures to obtain the target object detection results of each of the multiple image regions; A result output module 207 for obtaining the target object detection result of the original high-resolution image according to the target object detection results of each of the multiple image regions.
[0083] In an alternative embodiment, the filtering module includes: A contour recognition sub-module for recognizing the contours of each target from the low-resolution image through a pre-trained segmentation model; A classification sub-module for classifying the contours of each target to obtain contour classification results; A filtering sub-module for filtering out the image regions of the targets whose contours are of non-target object categories and retaining the image regions of the targets whose contours are of target object categories.
[0084] In an alternative embodiment, the classification sub-module includes: A coordinate determination unit for, for each target among the respective targets, determining the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target; A coordinate conversion unit for, for each target among the respective targets, using the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system to convert each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system to obtain the coordinates of each pixel point on the contour of the target in the polar coordinate system; A coordinate processing unit for processing the coordinates of each pixel point on the contour of each target in the polar coordinate system through a pre-trained classifier to determine whether the category of the contour of each target is a target object category.
[0085] In an alternative embodiment, the restoration module includes: A scaling sub-module for determining the scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and restoring the resolution of the low-resolution image with the image regions whose contours are of target object categories retained to the resolution of the original high-resolution image according to the scaling ratio; or A mapping sub-module, configured to map the coordinates of the pixels on the contour of the target object in the low-resolution image to the high-resolution image according to the coordinate mapping relationship between the coordinates of the pixels of the original high-resolution image in the pixel coordinate system and the coordinates of the pixels of the low-resolution image in the pixel coordinate system.
[0086] In an alternative embodiment, it further includes: A mask module, configured to perform mask marking on the high-resolution image after resolution restoration according to the image regions whose filtered contours are non-target object categories, to obtain a high-resolution image including a mask region and a remaining region, where the mask region corresponds to the image regions whose filtered contours are non-target object categories; The slicing module includes: A slicing sub-module, configured to slice the remaining region in the mask-marked high-resolution image through a sliding window to obtain a plurality of image regions.
[0087] In an alternative embodiment, the quality analysis module includes: A quality level determination sub-module, configured to determine the quality level of each of the plurality of image regions according to the clarity, noise level, and texture complexity of each of the plurality of image regions; The target detection module includes: A classification detection sub-module, configured to perform target detection on each of the plurality of image regions through a pre-trained slow target detector when the quality level of the image region is low-quality, and perform target detection on the image region through a pre-trained fast target detector when the quality level of the image region is high-quality; At least one of the following is satisfied between the slow target detector and the fast target detector: The inference speed of the slow target detector is less than the inference speed of the fast target detector; The number of model parameters of the slow target detector is greater than the number of model parameters of the fast target detector; The target object feature extraction ability of the slow target detector is greater than the target object feature extraction ability of the fast target detector; The structural complexity of the slow target detector is greater than the structural complexity of the fast target detector.
[0088] In an alternative embodiment, the result output module includes: A fusion sub-module, configured to fuse the target object detection results of each of the plurality of image regions; A removal sub-module, configured to remove redundant target object position detection frames in the fused target object detection result, so as to obtain the target object detection result of the original high-resolution image.
[0089] Embodiments of the present disclosure further provide an electronic device, including a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, it implements each process of the above-mentioned embodiments of the method for detecting a target object in a high-resolution image based on inverse segmentation, and can achieve the same technical effects. To avoid repetition, details are not described herein again.
[0090] Embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned embodiments of the method for detecting a target object in a high-resolution image based on inverse segmentation, and can achieve the same technical effects. To avoid repetition, details are not described herein again. Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0091] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, an electronic device, and a storage medium. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0092] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the method and apparatus according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks. These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing terminal devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the process Figure 1one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.
[0093] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0094] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0095] The above has introduced in detail the method and device for detecting a target object in a high-resolution image based on reverse segmentation provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A high-resolution image target object detection method based on reverse segmentation, characterized in that: include: The resolution of the original high-resolution image is reduced to obtain a low-resolution image; From the low-resolution image, filter out image regions whose contours are non-target object categories, and retain image regions whose contours are target object categories; Restoring the resolution of the low-resolution image, where the image region with the contour of the target object category is retained, to the resolution of the original high-resolution image; Slicing the high-resolution image after resolution restoration through a sliding window to obtain multiple image regions; determining a quality of each of the plurality of image regions; Performing target detection on image regions of different qualities by using detectors of different structures to obtain target object detection results for each of the multiple image regions; The target object detection result of the original high-resolution image is obtained according to the target object detection results of each of the multiple image regions.
2. The method according to claim 1, characterized in that From the low-resolution image, filtering out image regions whose contours are non-target object categories, and retaining image regions whose contours are target object categories, including: Using a pre-trained segmentation model, identifying the contours of each target from the low-resolution image; Classifying the contours of each target to obtain a contour classification result; The image regions whose contours belong to the targets of non-target object categories are filtered out, and the image regions whose contours belong to the targets of target object categories are retained.
3. The method according to claim 2, characterized in that Classifying the contours of the respective targets to obtain contour classification results, including: For each of the targets, determine the coordinates of the center of gravity of the target in the pixel coordinate system according to the coordinates of each pixel point on the contour of the target in the pixel coordinate system and the total number of contour pixels of the target; For each of the targets, taking the coordinates of the center of gravity of the target in the pixel coordinate system as the coordinate origin of the polar coordinate system, transforming each pixel point on the contour of the target from the pixel coordinate system to the polar coordinate system, and obtaining the coordinates of each pixel point on the contour of the target in the polar coordinate system; The coordinates of each pixel point on the contour of each target in the polar coordinate system are processed by a pre-trained classifier to determine whether the category of the contour of each target is the target object category.
4. The method according to claim 1, characterized in that: Restoring the resolution of the low-resolution image in which the image region having the contour of the target object category is retained to the resolution of the original high-resolution image, comprising: Determine a scaling ratio used when reducing the resolution of the original high-resolution image to the resolution of the low-resolution image, and restore the resolution of the low-resolution image where the image region with the outline of the target object category is retained to the resolution of the original high-resolution image according to the scaling ratio; or According to the coordinate mapping relationship between the coordinates of the pixel points of the original high-resolution image in the pixel coordinate system and the coordinates of the pixel points of the low-resolution image in the pixel coordinate system, the coordinates of the pixel points on the contour of the target object in the low-resolution image are mapped to the high-resolution image.
5. The method according to claim 1, characterized in that After restoring the resolution of the low-resolution image in which the image region with the contour of the target object category is retained to the resolution of the original high-resolution image, the method further includes: According to the image area whose contour is a non-target object category that is filtered out, mask marking is performed on the high-resolution image after the resolution is restored to obtain a high-resolution image including the mask area and the remaining area, wherein the mask area corresponds to the image area whose contour is a non-target object category that is filtered out; The high-resolution image after resolution restoration is sliced through a sliding window to obtain multiple image regions, including: The remaining area in the high-resolution image after mask marking is sliced through a sliding window to obtain multiple image regions.
6. The method according to claim 1, characterized in that Determining the quality of each of the plurality of image regions includes: Determining the quality level of each of the multiple image regions according to the clarity, noise level, and texture complexity of each of the multiple image regions; Detection of objects in image regions of different qualities is performed using detectors of different structures, including: For each image region among the multiple image regions, when the quality level of the image region is low quality, performing object detection on the image region by using a pre-trained slow object detector, and when the quality level of the image region is high quality, performing object detection on the image region by using a pre-trained fast object detector; At least one of the following conditions is satisfied between the slow target detector and the fast target detector: The inference speed of the slow object detector is less than the inference speed of the fast object detector; The model parameter amount of the slow target detector is greater than the model parameter amount of the fast target detector; The target object feature extraction capability of the slow target detector is greater than the target object feature extraction capability of the fast target detector; The structural complexity of the slow object detector is greater than the structural complexity of the fast object detector.
7. The method according to any one of claims 1 to 6, characterized in that: Obtaining the target object detection result of the original high-resolution image according to the target object detection results of each of the multiple image regions, including: fusing the target object detection results of the respective multiple image regions; The redundant target object position detection frames in the fused target object detection results are removed to obtain the target object detection results of the original high-resolution image.
8. A high-resolution image target object detection device based on reverse segmentation, characterized in that: include: A low-resolution processing module is used to reduce the resolution of the original high-resolution image to obtain a low-resolution image; A filtering module, used to filter out image regions whose contours are non-target object categories from the low-resolution image, and retain image regions whose contours are target object categories; A restoration module, used for restoring the resolution of the low-resolution image with the image region whose contour is the target object category to the resolution of the original high-resolution image; A slicing module, used for slicing the high-resolution image after resolution restoration through a sliding window to obtain multiple image regions; A quality analysis module, configured to determine the quality of each of the plurality of image regions; A target detection module, used to perform target detection on image regions of different qualities by using detectors of different structures, and obtain target object detection results of the respective multiple image regions; The result output module is used to obtain the target object detection result of the original high-resolution image according to the target object detection results of each of the multiple image regions.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
CFAR (Constant False Alarm Rate) and sparse representation-based high-resolution SAR (Synthetic Aperture Radar) image ship detection method
CN103400156A
Automatic brain region segmentation method and system for brain tissue three-dimensional image with cell resolution level
CN110675372A
Pedestrian detection and recognition system, method and device and computer readable storage medium
CN111274991A
Brake beam falling detection method based on deep learning
CN112634242A
Image acquisition method based on human eye attention perception mechanism
CN116664821A