Object Detection Method, Apparatus, Electronic Device, and Computer Readable Storage Medium

By reducing and slidingly dividing the super-large pixel image, multiple sub-maps are generated and object detection model is detected, the object detection problem in super-large pixel scenes is solved, and the accurate detection of targets in super-large pixel images is achieved.

CN114612872BActive Publication Date: 2025-06-17GUANGZHOU YAXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111562162.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-06-17
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to effectively handle pedestrian and vehicle object detection in super-large pixel scenes, especially under conditions where the scene coverage is wide, the image pixels are up to 100 million, the target scale changes greatly, and the appearance characteristics are still clear.

Method used

Multiple scale images are obtained by acquiring super-pixel images and reducing them in preset multiples. Each scale image is slided and segmented using a sliding window to generate multiple sub-maps, and input these sub-maps into the trained object detection model for detection. Then, the sub-picture detection results of each scale image are fused to obtain the detection results of the scale image, and finally the detection results of each scale image of the super-large pixel image are fused to obtain the detection results of the super-large pixel image.

Benefits of technology

Effective detection of targets in super-large pixel images is achieved, pedestrians and vehicles can be accurately identified in complex large scenes, and the accuracy and efficiency of target detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612872B_ABST
    Figure CN114612872B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a target detection method, device, electronic device and computer-readable storage medium, including: obtaining an ultra-large pixel image, and reducing the ultra-large pixel image by at least one preset multiple to obtain a corresponding scale image; using a sliding window with a first preset size to perform sliding segmentation on the scale image to obtain corresponding multiple sub-images, and inputting each sub-image into a trained target detection model to obtain the detection result of the sub-image; fusing the detection results of each sub-image to obtain the detection result corresponding to the scale image; fusing the detection results of each scale image to obtain the detection result of the ultra-large pixel image. This solution uses a target detection model to detect the sub-images obtained after reducing and segmenting the ultra-large pixel image to obtain the detection results of each sub-image, and then obtains the detection result of the ultra-large pixel image after multiple fusions based on the detection results of the sub-images, realizing the detection of targets in the ultra-large pixel image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology. Specifically, the present application relates to an object detection method, device, electronic device, and computer-readable storage medium. Background Art

[0002] Multi-object detection of pedestrians and vehicles is an object detection technology that uses computer vision technology to determine whether there are pedestrian or vehicle objects in an image and gives precise positioning. This technology is widely used in fields such as intelligent monitoring, autonomous driving, and smart cities. The detection process is as Figure 1 shown (in the figure, "Person" indicates that the object category in the corresponding object box is "pedestrian", and "Vehicle" indicates that the object category in the corresponding object box is "vehicle"). Currently, the mainstream object detection algorithms are mainly divided into two categories. One is the two-stage detection algorithm represented by Faster RCNN, and the other is the one-stage detection algorithm represented by SSD (Single Shot MultiBox Detector), YOLO (You Only Look Once) series.

[0003] Ultra-large pixel scenes generally refer to scenes with a relatively wide coverage area such as intersections, stations, large-scale event sites, and commercial plazas and their surrounding areas. The image pixels of these scenes reach hundreds of millions, and the global view covers natural scenes of square kilometers. There are up to thousands of people in the scene, and the scale change of a single object can reach up to a hundred times. The appearance features of local objects are still clearly distinguishable in the maximum magnification view. The images in ultra-large pixel scenes are as Figure 2 shown.

[0004] Currently, the object detection technology of pedestrians and vehicles in small-scale scenes and conventional images with low video image resolution is relatively mature. In order to better ensure people's safe travel and intelligent interaction and entertainment, the demand for monitoring and analyzing tasks of pedestrians and vehicle objects in large-scale scenes has increased sharply. However, the characteristics and difficulties of images in these scenes that are different from conventional images make it impossible for traditional object detection algorithms to directly perform effective image analysis on these scenes. Summary of the Invention

[0005] The purpose of the present application is to at least solve one of the above technical defects. The technical solutions provided by the embodiments of the present application are as follows:

[0006] In a first aspect, an embodiment of the present application provides an object detection method, including:

[0007] Obtain an ultra-large pixel image, and reduce the ultra-large pixel image by at least one preset multiple to obtain a corresponding scaled image;

[0008] For each scale image, use a sliding window of the first preset size to perform sliding segmentation on the scale image, obtain corresponding multiple sub-images, and input each sub-image into a trained object detection model to obtain the detection result of the sub-image. The trained object detection model is obtained by training with sub-image samples marked with detection results. The detection results include target box coordinates, target categories, and target box confidence levels;

[0009] Fuse the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image;

[0010] Fuse the detection results of each scale image corresponding to the super-large pixel image to obtain the detection result of the super-large pixel image.

[0011] In an alternative embodiment of the present application, for each scale image, using a sliding window of the first preset size to perform sliding segmentation on the scale image to obtain corresponding multiple sub-images includes:

[0012] Use the sliding window to slide on the scale image according to a preset step size. The area corresponding to the sliding window on the scale image after each slide is a sub-image. The ratio of the preset step size to the width of the sliding window is within a preset ratio range;

[0013] If the sliding window exceeds the boundary of the scale image after one slide, translate the sliding window into the scale image to obtain the corresponding sub-image.

[0014] In an alternative embodiment of the present application, the object detection model includes a backbone Backbone module, an intermediate Neck module, and an output module;

[0015] Input each sub-image into the trained object detection model to obtain the detection result of the sub-image, including:

[0016] Respectively extract the target features of the sub-image through the Transformer layer and the deformable convolution DCN layer in the Backbone layer; then fuse the target feature information through the Neck module to obtain the corresponding fused features; finally, the output module respectively outputs the corresponding initial detection results based on the fused features corresponding to at least two network layers in the Neck module, and obtain the detection result corresponding to the sub-image based on multiple initial detection results.

[0017] In an alternative embodiment of the present application, obtaining the detection result corresponding to the sub-image based on multiple initial detection results includes:

[0018] Project all the target boxes corresponding to multiple initial detection results into the sub-image, and obtain at least one set of target boxes with the first intersection over union (IOU) not less than the first preset threshold;

[0019] For each group of target bounding boxes with the first IOU not less than the first preset threshold, use the Weighted Box Fusion (WBF) algorithm to obtain the corresponding fused target bounding box based on this group of target bounding boxes with the first IOU not less than the first preset threshold;

[0020] Based on the fused target bounding boxes corresponding to each group of target bounding boxes with the first IOU not less than the first preset threshold, and other target bounding boxes except each group of target bounding boxes with the first IOU not less than the first preset threshold, obtain the detection result of the sub - image.

[0021] In an alternative embodiment of the present application, the trained object detection model is obtained through the following method:

[0022] For at least one super - large pixel image sample annotated with detection results, downscale it by at least one preset multiple to obtain the corresponding scale image sample, and for each scale image sample, use a sliding window of the first preset size to slide - split the scale image to obtain a preset number of sub - image samples;

[0023] Perform joint data augmentation on each sub - image sample to obtain the corresponding data - augmented sub - image sample, and use a preset amount of data - augmented sub - image samples to train the initial object detection model until the loss function meets the preset conditions to obtain the trained object detection model;

[0024] Among them, the loss function includes the target bounding box coordinate loss sub - function, and the Quality Focal Loss (QFL) sub - function obtained from the target classification loss sub - function and the target bounding box confidence loss sub - function.

[0025] In an alternative embodiment of the present application, performing joint data augmentation on each sub - image sample to obtain the corresponding data - augmented sub - image sample includes:

[0026] Obtain the detection results of the target bounding boxes in each sub - image sample that are not greater than the second preset size;

[0027] Copy the target bounding boxes corresponding to each detection result, and after translation and rotation by a preset angle, paste them into the non - target area of the sub - image to obtain the corresponding data - augmented sub - image sample.

[0028] In an alternative embodiment of the present application, fusing the detection results of each sub - image corresponding to each scale image to obtain the detection result corresponding to this scale image includes:

[0029] Based on the splitting method of each scale image, splice the sub - images with detection results;

[0030] For the previous sub - figure among two adjacent front - and - back sub - figures, if the target box corresponding to the detection result of the previous sub - figure is located on the left side of the mid - line of the overlapping area of the two sub - figures or intersects with the mid - line, the detection result is retained; if the target box corresponding to the detection result of the previous sub - figure is located on the right side of the mid - line, the detection result is discarded. For the subsequent sub - figure among two adjacent front - and - back sub - figures, if the target box corresponding to the detection result of the subsequent sub - figure is located on the right side of the mid - line, the detection result is retained; if the target box corresponding to the detection result of the subsequent sub - figure is located on the left side of the mid - line or intersects with the mid - line, the detection result is discarded.

[0031] In a second aspect, an embodiment of the present application provides an object detection device, including:

[0032] A scale image acquisition module, configured to acquire an ultra - large pixel image and reduce the ultra - large pixel image by at least one preset multiple to obtain a corresponding scale image;

[0033] A sub - figure acquisition and detection module, for each scale image, using a sliding window with a first preset size to perform sliding segmentation on the scale image to obtain corresponding multiple sub - figures, and inputting each sub - figure into a trained object detection model to obtain the detection result of the sub - figure. The trained object detection model is obtained by training with sub - figure samples marked with detection results, and the detection results include target box coordinates, target categories, and target box confidence levels;

[0034] A first detection result fusion module, configured to fuse the detection results of each sub - figure corresponding to each scale image to obtain the detection result corresponding to the scale image;

[0035] A second detection result fusion module, configured to fuse the detection results of each scale image corresponding to the ultra - large pixel image to obtain the detection result of the ultra - large pixel image.

[0036] In an alternative embodiment of the present application, the sub - figure acquisition and detection module is specifically configured to:

[0037] Use the sliding window to slide on the scale image at a preset step size. Each area corresponding to the sliding window on the scale image after each slide is a sub - figure, and the ratio of the preset step size to the width of the sliding window is within a preset ratio range;

[0038] If the sliding window exceeds the boundary of the scale image after one slide, translate the sliding window into the scale image to obtain the corresponding sub - figure.

[0039] In an alternative embodiment of the present application, the object detection model includes a backbone Backbone module, an intermediate Neck module, and an output module; the sub - figure acquisition and detection module is specifically configured to:

[0040] Extract the target features of the sub - graph through the Transformer layer and the deformable convolutional network (DCN) layer in the Backbone layer respectively; then fuse the target feature information through the Neck module to obtain the corresponding fused features; finally, the output module outputs the corresponding initial detection results based on the fused features corresponding to at least two network layers in the Neck module, and obtains the detection result corresponding to the sub - graph based on multiple initial detection results.

[0041] In an alternative embodiment of the present application, the sub - graph acquisition and detection module is further configured to:

[0042] Project all the target boxes corresponding to multiple initial detection results into the sub - graph, and obtain at least one group of target boxes with the intersection over union (IOU) not less than the first preset threshold;

[0043] For each group of target boxes with the first IOU not less than the first preset threshold, use the weighted box fusion (WBF) algorithm to obtain the corresponding fused target box based on this group of target boxes with the first IOU not less than the first preset threshold;

[0044] Based on the fused target boxes corresponding to each group of target boxes with the first IOU not less than the first preset threshold and the other target boxes except each group of target boxes with the first IOU not less than the first preset threshold, obtain the detection result of the sub - graph.

[0045] In an alternative embodiment of the present application, the device further includes a training module, which is used for:

[0046] For at least one super - pixel image sample labeled with detection results, reduce it by at least one preset multiple to obtain the corresponding scale image sample, and for each scale image sample, use a sliding window with a first preset size to slide - cut the scale image to obtain a preset number of sub - graph samples;

[0047] Perform joint data augmentation on each sub - graph sample to obtain the corresponding data - augmented sub - graph sample, and use the preset amount of data - augmented sub - graph samples to train the initial object detection model until the loss function meets the preset conditions to obtain the trained object detection model;

[0048] Wherein, the loss function includes the target box coordinate loss sub - function and the quality focal loss (QFL) sub - function obtained from the target classification loss sub - function and the target box confidence loss sub - function.

[0049] In an alternative embodiment of the present application, the training module is specifically used for:

[0050] Obtain the detection results of the target boxes not larger than the second preset size in each sub - graph sample;

[0051] Copy the target bounding box corresponding to each detection result, and after translation and rotation by a preset angle, paste it into the non-target area of the sub-image to obtain the sub-image sample after corresponding data augmentation.

[0052] In an alternative embodiment of the present application, the first detection result fusion module is specifically configured to:

[0053] Based on the segmentation method of each scale image, splice the sub-images with detection results;

[0054] For the previous sub-image among two adjacent front and rear sub-images, if the target bounding box corresponding to the detection result of the previous sub-image is located on the left side of the midline of the overlapping area of the two sub-images or intersects with the midline, the detection result is retained; if the target bounding box corresponding to the detection result of the previous sub-image is located on the right side of the midline, the detection result is discarded; for the subsequent sub-image among two adjacent front and rear sub-images, if the target bounding box corresponding to the detection result of the subsequent sub-image is located on the right side of the midline, the detection result is retained; if the target bounding box corresponding to the detection result of the subsequent sub-image is located on the left side of the midline or intersects with the midline, the detection result is discarded.

[0055] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor;

[0056] A computer program is stored in the memory;

[0057] The processor is configured to execute the computer program to implement the method provided in the first aspect embodiment or any alternative embodiment of the first aspect.

[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method provided in the first aspect embodiment or any alternative embodiment of the first aspect.

[0059] In a fifth aspect, an embodiment of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the method provided in the first aspect embodiment or any alternative embodiment of the first aspect when executed.

[0060] The beneficial effects brought by the technical solution provided by the present application are:

[0061] By performing downsampling on the super-large pixel image, multiple scale images are obtained. Then, a sliding window is used to slide and divide each scale image to obtain multiple sub-images. Each sub-image is input into the target detection model to obtain the detection results of each sub-image. Then, the detection results of all sub-images of each scale image are fused to obtain the detection results of the scale image. Finally, the detection results of all scale images of the super-large pixel image are fused to obtain the detection results of the super-large pixel image. This solution uses the target detection model to detect the sub-images obtained by downsampling and dividing the super-large pixel image to obtain the detection results of each sub-image, and then fuses the detection results of the sub-images multiple times to obtain the detection results of the super-large pixel image, realizing the detection of targets in the super-large pixel image. Description of the Drawings

[0062] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.

[0063] Figure 1 Schematic diagram of the target detection process in the prior art;

[0064] Figure 2 Example diagram of an image in a super-large pixel scenario;

[0065] Figure 3 Schematic flow diagram of a target detection method provided by an embodiment of the present application;

[0066] Figure 4 Schematic diagram of the Resize operation and sliding window division operation on a super-large pixel image in an example of an embodiment of the present application;

[0067] Figure 5 Schematic diagram of the structure of a target detection model in an example of an embodiment of the present application;

[0068] Figure 6 Schematic diagram of the Transformer structure in a target detection model in an example of an embodiment of the present application;

[0069] Figure 7 Schematic diagram of joint data augmentation for sub-image samples in an example of an embodiment of the present application;

[0070] Figure 8 Schematic diagram of the fusion process of the detection results of two adjacent front and rear sub-images in an example of an embodiment of the present application;

[0071] Figure 9 Schematic diagram of the overall process of target detection implementation in an example of an embodiment of the present application;

[0072] Figure 10The embodiment of the present application provides a structural block diagram of an object detection device;

[0073] Figure 11 The embodiment of the present application provides a schematic structural diagram of an electronic device. Detailed implementation manners

[0074] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be construed as a limitation to the present application.

[0075] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0076] To make the purpose, technical solution and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0077] Figure 3 The embodiment of the present application provides a schematic flowchart of an object detection method, as Figure 3 shown, the method may include:

[0078] Step S301, obtain an ultra-large pixel image, and reduce the ultra-large pixel image by at least one preset multiple to obtain a corresponding scale image.

[0079] Among them, the ultra-large pixel image is the image corresponding to the ultra-large pixel scenario, and it is necessary to detect and obtain the objects in the ultra-large pixel image, for example, pedestrians, vehicles, etc.

[0080] Specifically, after obtaining the super-large pixel image, in order to facilitate subsequent processing, its size can be reduced. In the embodiments of the present application, the super-large pixel image can be reduced by one or more preset multiples. This reduction process can also be referred to as "Resize" processing of the super-large pixel image. The preset multiples used in this "Resize" processing can be 0.2, 0.4, or 0.6. In other words, the super-large pixel image can be reduced by 0.2, 0.4, and 0.6 times respectively, thereby obtaining three corresponding scale images.

[0081] It can be understood that the preset multiples can be selected according to actual needs, and the embodiments of the present application do not make any limitations.

[0082] Step S302: For each scale image, use a sliding window of the first preset size to perform sliding segmentation on the scale image, obtaining corresponding multiple sub-images, and input each sub-image into the trained target detection model to obtain the detection result of the sub-image. The trained target detection model is obtained by training with sub-image samples marked with detection results, and the detection results include target box coordinates, target categories, and target box confidence levels.

[0083] Specifically, each scale image obtained after reducing the super-large pixel image is segmented. Specifically, a sliding window of the first preset size is used to perform sliding segmentation on each scale image, and the scale image is divided into multiple sub-images. Then, each sub-image obtained by segmenting each image is input into the trained target detection model to obtain the detection results of the sub-images, that is, the positions of the target boxes (and target box coordinates) in the sub-images, the target categories of the targets in the target boxes, and the target box confidence levels.

[0084] It can be understood that the sub-images obtained by segmenting each scale image are equivalent to images in a small scene or equivalent to conventional images, and can be directly input into the target detection model for processing, and then the detection results of each sub-image are fused to obtain the detection result of the entire scale image.

[0085] It should be noted that the first preset size can be selected according to actual needs or determined with reference to the annotation information during the training process, which will be described in detail later.

[0086] Step S303: Fuse the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image.

[0087] Specifically, in the previous step, the detection results of each sub-image of each scale image have been obtained. Then, by fusing the detection results of all sub-images corresponding to the scale image, the detection result of the scale image can be obtained.

[0088] Step S304: Fuse the detection results of each scale image corresponding to the super-large pixel image to obtain the detection result of the super-large pixel image.

[0089] Specifically, in the previous step, the detection results of each scale image corresponding to the super-large pixel image have been obtained. Then, by fusing the detection results of all scale images corresponding to the super-large pixel image, the detection result of the super-large pixel image can be obtained, realizing the object detection of the super-large pixel image.

[0090] The solution provided in this application reduces the super-large pixel image to obtain multiple scale images, then uses a sliding window to slide and split each scale image into multiple sub-images, inputs each sub-image into the object detection model to obtain the detection results of each sub-image, then fuses the detection results of all sub-images of each scale image to obtain the detection result of the scale image, and finally fuses the detection results of all scale images of the super-large pixel image to obtain the detection result of the super-large pixel image. This solution uses the object detection model to detect the sub-images after reducing and splitting the super-large pixel image to obtain the detection results of each sub-image, and then obtains the detection result of the super-large pixel image through multiple fusions based on the detection results of the sub-images, realizing the detection of the object in the super-large pixel image.

[0091] In an alternative embodiment of the present application, for each scale image, using a sliding window with a first preset size to slide and split the scale image to obtain corresponding multiple sub-images, including:

[0092] Use the sliding window to slide on the scale image at a preset step length. Each time it slides, the area corresponding to the sliding window on the scale image is a sub-image, and the ratio of the preset step length to the width of the sliding window is within a preset ratio range;

[0093] If the sliding window exceeds the boundary of the scale image after one slide, translate the sliding window into the scale image to obtain the corresponding sub-image.

[0094] Specifically, when sliding and splitting each scale image, first determine the size of the sliding window, and then use this sliding window to slide and split each scale image respectively. Specifically, slide the sliding window on each scale image at a preset step length. Each time it slides, the area on the scale image covered by the sliding window is used as a sub-image. Among them, there may be an overlapping area (overlap) between the sub-images generated by two consecutive slides during the sliding process, and the size of the overlapping area is controlled by the preset step length.

[0095] Such as Figure 4As shown, first, a Resize operation is performed on the super-large pixel image to reduce the super-large pixel image to 0.2, 0.4, and 0.6 times its original size (i.e., three preset multiples) to obtain three corresponding scale images. Then, the preset step size is used to make the overlap ratio range between the two adjacent sub-images be between 0.1 and 0.5. Finally, the sub-images of the third batch corresponding to the three scale images are obtained.

[0096] Furthermore, the sliding window can slide on the scale image in the order from left to right and then from top to bottom. If the sliding window exceeds the scale image, it is generally exceeded to the right or downwards. Then, the sliding window can be moved upwards or leftwards into the scale image for processing, so as to obtain sub-images with complete sizes.

[0097] In an alternative embodiment of the present application, the object detection model includes a backbone Backbone module, an intermediate Neck module, and an output module;

[0098] Each sub-image is input into the trained object detection model to obtain the detection result of the sub-image, including:

[0099] The object features of the sub-image are respectively extracted through the Transformer layer and the deformable convolutional DCN layer in the Backbone layer; then, the object feature information is fused through the Neck module to obtain the corresponding fused features; finally, the output module respectively outputs the corresponding initial detection results based on the fused features corresponding to at least two network layers in the Neck module, and based on the multiple initial detection results, the detection result corresponding to the sub-image is obtained.

[0100] Among them, the object detection model includes a Backbone module, an intermediate Neck module, and an output module. After each sub-image is input into the object detection model, after the feature extraction of the Backbone module and the Neck module, the detection results of each sub-image are output through the output module.

[0101] For example, such as Figure 5As shown in the figure, the function of the Backbone module (layers 1-12 in the figure) is to extract the features of pedestrian and vehicle targets in the input sub-graph (assuming the sub-graph size is img_size*img_size), mainly including Focus, C3, and SPP structures. In this example, the Backbone of YOLOv5x is improved as follows: a. Optimize the 3 C3 structures in layer 12 into Transformer structures to extract the global information of the image; b. Replace the traditional convolutions with fixed shapes in the C3 structures of layers 3, 5, 7, and 9 with deformable convolutions (Deformable Convolution Net, DCN). DCN adds learnable position offset parameters to the convolution action area with a fixed shape, so that the sampling points of the convolution kernel spread into a non-grid shape to better extract the features of targets with large appearance shape differences and serious occlusions.

[0102] The function of the Neck module (layers 13-33 in the figure) is to fuse the target features extracted by each convolutional layer in the Backbone module. In this example, four feature layers output from layers 5, 7, 9, and 12 in the Backbone (feature map sizes are img_size / 8, img_size / 16, img_size / 32, img_size / 64 respectively) are selected for feature fusion according to the PANet network structure.

[0103] The function of the output module is to predict the category and coordinates of the targets in the feature map. In this example, the category and coordinates of the targets in the feature maps output from layers 24, 27, 30, and 33 of the Neck module are predicted to obtain multiple target boxes (corresponding to multiple initial detection results), and then the WBF (Weighted-Boxes-Fusion) method is used to screen the target boxes to obtain more accurate target categories and target box coordinates (corresponding to the final detection results of the sub-graph).

[0104] Furthermore, as Figure 6As shown in the figure, Flatten in the Transformer structure flattens the feature map. Linear projection and MLP (Multilayer Perceptron) are implemented through fully connected layers. MultiHead Attention is a multi-head attention function implemented in the PyTorch deep learning framework (a Python machine learning library based on the open-source Torch). ⊕ represents the operation of adding feature maps. The input of the Transformer is the feature map output by the 11th layer of the Backbone module. After Flatten and linear projection, the input values q, k, and v of the MultiHead Attention function are obtained. The output of MultiHead Attention is the weighted feature map, which is then summed with the vector of Flatten. Finally, after performing a reshape operation on the output, a feature map with the same depth and width as the original input feature map is obtained.

[0105] In an alternative embodiment of the present application, obtaining the detection result corresponding to the sub-graph based on the multiple initial detection results includes:

[0106] Project all the target bounding boxes corresponding to the multiple initial detection results into the sub-graph, and obtain at least one group of target bounding boxes with the first Intersection over Union (IOU) not less than the first preset threshold;

[0107] For each group of target bounding boxes with the first IOU not less than the first preset threshold, use the Weighted Boxes Fusion (WBF) algorithm to obtain the corresponding fused target bounding box based on this group of target bounding boxes with the first IOU not less than the first preset threshold;

[0108] Based on the fused target bounding boxes corresponding to each group of target bounding boxes with the first IOU not less than the first preset threshold, and other target bounding boxes other than each group of target bounding boxes with the first IOU not less than the first preset threshold, obtain the detection result of the sub-graph.

[0109] Specifically, first project all the target bounding boxes corresponding to the multiple initial detection results into the same sub-graph area, then calculate the first IOU of each group of overlapping target bounding boxes, and find each group of target bounding boxes with the first IOU not less than the first preset threshold (the first preset threshold can be set to 0.7). Then, use the WBF algorithm to process each group of target bounding boxes with the first IOU not less than the first preset threshold to obtain the corresponding fused target bounding boxes respectively. Finally, the multiple fused target bounding boxes and those target bounding boxes that have not been processed by the WBF algorithm are the final target bounding boxes of the sub-graph. Then, the target bounding box coordinates, target bounding box confidence, and target bounding box category corresponding to these final target bounding boxes are the detection results of the sub-graph.

[0110] Among them, the WBF method weights and fuses the confidence of each group of target boxes and the target box coordinates with the first IOU greater than the first preset threshold (the first preset threshold can be set to 0.7) to obtain the final target box confidence and target box coordinates. The fusion weight is determined by the target box confidence in the initial detection result. Compared with the traditional target box screening method, the WBF method can obtain more accurate target box confidence and target box coordinates.

[0111] In an alternative embodiment of the present application, the trained object detection model is obtained through the following method:

[0112] For at least one super-large pixel image sample annotated with detection results, corresponding scale image samples are obtained by shrinking it by at least one preset multiple. For each scale image sample, the scale image is slid and sliced using a sliding window with the first preset size to obtain a preset number of sub-image samples;

[0113] Perform joint data augmentation on each sub-image sample to obtain the corresponding data-augmented sub-image sample, and use the preset amount of data-augmented sub-image samples to train the initial object detection model until the loss function meets the preset conditions to obtain the trained object detection model;

[0114] Among them, the loss function includes a target box coordinate loss sub-function, and a quality focal loss QFL sub-function obtained from a target classification loss sub-function and a target box confidence loss sub-function.

[0115] Specifically, the initial object detection model is trained with sub-image samples to obtain the trained object detection model. Among them, the process of obtaining sub-image samples is as follows: First, annotate the super-large pixel image sample, that is, annotate the target category, target box coordinates, and target box confidence. Then, shrink the super-large pixel image sample by multiple preset multiples to obtain multiple scale image samples. Finally, perform sliding slicing on each scale image sample using a sliding window to obtain multiple batches of sub-image samples.

[0116] It can be understood that the shrinking or sliding slicing process of the super-large pixel image sample in the process of obtaining sub-image samples is the same as the shrinking or sliding slicing process of the super-large pixel image to be detected in the model application process, and the preset multiples and the size of the sliding window used are also the same.

[0117] It should be noted that if the target box of the sub - image sample obtained in the above - mentioned manner is incomplete, then calculate the second IOU between the incomplete target box on the sub - image sample and the original target box in the corresponding scale image sample, and retain the target boxes whose second IOU is not less than the second preset threshold (the first preset threshold can be set to 0.5). Mask the target boxes whose second IOU is less than the second preset threshold. In other words, use a unified gray value (such as [0, 0, 0], etc.) to replace the image gray value within the target box area. After such processing, it is equivalent to erasing the target from the sub - image sample.

[0118] It can be understood that the masking process is only required during the model training stage and is not required during the model application stage.

[0119] In addition, count the sizes of the target boxes annotated in the super - large pixel image sample and obtain the maximum size (max_w, max_h) of the target boxes among them. Then, the range of the first preset size of the sliding window can be set to (0.2 - 0.6)*(max_w, max_h), that is, the first preset size can be selected from this range.

[0120] Furthermore, after obtaining multiple batches of sub - image samples, randomly split a preset number of sub - image samples into training data and test data according to a ratio of 3:1; perform joint data augmentation on the training data; use the officially open - sourced YOLOv5x model and model training hyperparameters as the pre - trained model and hyperparameters, and adopt the strategy of random multi - scale input to train the pedestrian and vehicle target detection model (i.e., the initial target detection model). The model training aims to minimize the calculated value of the loss function. After the model iterates a specific number of rounds, it is tested with the test set; select the model with the largest mAP (Mean Average Precision) value on the test set as the optimal model.

[0121] Specifically, the total loss Loss during model training consists of the target box coordinate loss L box 、the target box confidence loss L obj and the target classification loss L cls : Loss = L box + L obj + L cls , where L box is calculated using the CIOU loss. In this application, L obj and L cls are optimized to Quality Focal Loss (QFL), and the specific calculation method is as follows:

[0122] L obj (σ), L cls (σ)= -|y - σ| β((1-y)log(1-σ)+ylog(σ))

[0123] In the formula, y is the IOU value of the predicted target box and the true target box (ranging from 0 to 1), σ is the predicted target box confidence or target classification probability, and the β value here can be 2. QFL supports the joint representation of target box coordinate quality and target classification probability or target box confidence while having the characteristics of Focal Loss balancing positive and negative, difficult and easy samples; it avoids the situation where a true negative sample with a low classification probability is better than a true sample with a low classification probability but a low position score in candidate box screening due to predicting an unreliable extremely high position score.

[0124] In an optional embodiment of the present application, performing joint data enhancement on each sub-image sample to obtain a corresponding sub-image sample after data enhancement includes:

[0125] Obtaining a detection result in which the target frame in each sub-image sample is not larger than a second preset size;

[0126] The target frame corresponding to each detection result is copied and, after translation and rotation by a preset angle, pasted into the non-target area of ​​the sub-image to obtain the corresponding data-enhanced sub-image sample.

[0127] Specifically, for each sub-image sample, the detection results corresponding to the target frames that are no larger than the second preset size are obtained, the target frames corresponding to these detection results are translated and rotated by a preset angle, and then pasted to the non-target area of ​​the sub-image to obtain the corresponding data-enhanced sub-image samples. The target information in the processed sub-image samples is richer.

[0128] It should be noted that the joint data enhancement of the sub-images, in addition to the above operations, may also include image perturbations of the sub-image samples, changes in brightness, contrast, saturation, hue, noise addition, random scaling, random cropping, flipping, rotation, random erasing, and other traditional data enhancement methods that come with the open source yolov5 algorithm. In actual processing, the sub-image samples may be subjected to the above traditional data enhancement methods first, and then the detection results of the target box in each sub-image sample that is not larger than the second preset size may be translated and rotated by a preset angle, and then the enhanced result may be pasted to the non-target area of ​​the sub-image.

[0129] For example, if Figure 7 As shown in FIG. 1 , the images and their annotation information in the sub-image samples whose pedestrian and vehicle target frame sizes are less than 15*15 pixels (i.e., the second preset size) are jointly enhanced. Specifically, Figure 7 The vehicle and pedestrian targets in the left image are copied, translated, rotated by a certain angle, and then pasted into the non-target area of ​​the original image. Figure 7Enhanced sub - graph sample of the right figure.

[0130] In an alternative embodiment of the present application, the detection results of each sub - graph corresponding to each scale image are fused to obtain the detection result corresponding to the scale image, including:

[0131] Based on the segmentation method of each scale image, the sub - graphs with detection results are stitched together;

[0132] For the front sub - graph among two adjacent front - and - back sub - graphs, if the target box corresponding to the detection result of the front sub - graph is on the left side of or intersects with the mid - line of the overlapping area of the two sub - graphs, the detection result is retained; if the target box corresponding to the detection result of the front sub - graph is on the right side of the mid - line, the detection result is discarded. For the back sub - graph among two adjacent front - and - back sub - graphs, if the target box corresponding to the detection result of the back sub - graph is on the right side of the mid - line, the detection result is retained; if the target box corresponding to the detection result of the back sub - graph is on the left side of or intersects with the mid - line, the detection result is discarded.

[0133] Specifically, for each scale image, after obtaining the detection results of all sub - graphs corresponding to the scale image, the sub - graphs with detection results are stitched together, and during the stitching process, the target boxes in two adjacent front - and - back sub - graphs are fused.

[0134] Among them, the order and position of sub - graph stitching correspond to those during the sliding segmentation of the scale image.

[0135] Among them, for the front sub - graph among two adjacent front - and - back sub - graphs, if the target box corresponding to the detection result of the front sub - graph is on the left side of or intersects with the mid - line of the overlapping area of the two sub - graphs, the detection result is retained; if the target box corresponding to the detection result of the front sub - graph is on the right side of the mid - line, the detection result is discarded. For the back sub - graph among two adjacent front - and - back sub - graphs, if the target box corresponding to the detection result of the back sub - graph is on the right side of the mid - line, the detection result is retained; if the target box corresponding to the detection result of the back sub - graph is on the left side of or intersects with the mid - line, the detection result is discarded. For example, as Figure 8 shown, when sub - graph A and sub - graph B are stitched together (assuming that both sub - graphs respectively contain target boxes ①, ②, and ③), with the mid - line of the two overlapping areas (i.e., the dotted line in the figure) as a reference, during the result fusion of sub - graph A and sub - graph B, sub - graph A retains the detection results of target box ① and target box ②, and discards the detection result of target box ③; sub - graph B retains the detection result of target box ③, and discards the detection results of target box ① and target box ②. The fusion method for other adjacent front - and - back sub - graphs is the same.

[0136] It should be noted that after obtaining the detection results of the images at each scale of the super-large pixel image, each scale image is enlarged proportionally to the size of the super-large pixel image and the target boxes therein, and then all the target boxes in each scale image are projected onto the same super-large pixel image area, and the third IOU corresponding to multiple groups of target boxes with overlaps is obtained. For each group of target boxes with a third IOU greater than the third preset threshold (the third preset threshold can be 0.7), the target box with the highest confidence in the target boxes is retained, and other target boxes are deleted. In addition, all target boxes other than those with a third IOU greater than the third preset threshold in each group are retained. The target box coordinates, target box confidence, and target types of these retained target boxes are the detection results of the super-large pixel image.

[0137] The following further illustrates the solution of the embodiment of the present application through an example. As Figure 9 shown, the implementation process of this solution can be divided into an S1 data preprocessing module, an S2 model training module, and an S3 model inference module. The execution processes of each module can include:

[0138] S1 data preprocessing module:

[0139] S1.1. Obtain large-scene image data that has been labeled in different scenarios, that is, super-large pixel images with annotation information;

[0140] S1.2. Conduct statistical analysis on the annotation information in the data to obtain the scale information of pedestrian and vehicle targets;

[0141] S1.3. Use the scale information of the target as a reference to perform parallel multi-scale sliding window slicing on the super-large pixel image data to obtain multi-scale sub-images.

[0142] S2 model training module:

[0143] S2.1. Randomly split the multi-scale sub-images in a ratio of 3:1 to obtain training data and test data;

[0144] S2.2. Perform joint data augmentation on the training data;

[0145] S2.3. Use the officially open-sourced YOLOv5x model and model training hyperparameters as the pre-trained model and hyperparameters, and adopt the strategy of randomly multi-scale input to train the pedestrian and vehicle target detection model. The model training aims to minimize the calculated value of the loss function, and the model is tested with the test set after each specific number of iterations;

[0146] S2.4. Select the model with the largest mAP (Mean Average Precision) value on the test set as the optimal model, that is, obtain the trained target detection model;

[0147] S3 Model Inference Module:

[0148] S3.1. Obtain an image in an ultra-large pixel scene, that is, obtain an ultra-large pixel image to be detected.

[0149] S3.1. Perform online multi-scale sliding window slicing on the ultra-large pixel image in the manner of S1.3 to obtain multi-scale sub-images.

[0150] S3.2. Use the optimal model obtained in S2.4 to perform parallel inference on the above multi-scale sub-images to obtain candidate bounding boxes (i.e., target boxes) of the target in each sub-image.

[0151] S3.3. Use the WBF method to filter the candidate bounding boxes to obtain the detection results of each sub-image.

[0152] S3.4. Perform boundary fusion on the detection results of each sub-image to obtain the detection result of the ultra-large pixel image.

[0153] Figure 10 This application provides a structural block diagram of an object detection device, as Figure 10 shown. The device 1000 may include: a scale image acquisition module 1001, a sub-image acquisition and detection module 1002, a first detection result fusion module 1003, and a second detection result fusion module 1004, where:

[0154] The scale image acquisition module 1001 is used to acquire an ultra-large pixel image and reduce the ultra-large pixel image by at least one preset multiple to obtain a corresponding scale image.

[0155] The sub-image acquisition and detection module 1002 is used to, for each scale image, perform sliding segmentation on the scale image using a sliding window of a first preset size to obtain corresponding multiple sub-images, and input each sub-image into a trained object detection model to obtain the detection result of the sub-image. The trained object detection model is obtained by training with sub-image samples marked with detection results, and the detection results include target box coordinates, target categories, and target box confidence levels.

[0156] The first detection result fusion module 1003 is used to fuse the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image.

[0157] The second detection result fusion module 1004 is used to fuse the detection results of each scale image corresponding to the ultra-large pixel image to obtain the detection result of the ultra-large pixel image.

[0158] The solution provided by this application obtains multiple scale images by shrinking the super-large pixel image, then uses a sliding window to slide and split each scale image to obtain multiple sub-images, and inputs each sub-image into the target detection model to obtain the detection results of each sub-image. Then, the detection results of all sub-images of each scale image are fused to obtain the detection result of the scale image. Finally, the detection results of all scale images of the super-large pixel image are fused to obtain the detection result of the super-large pixel image. This solution uses the target detection model to detect the sub-images obtained by shrinking and splitting the super-large pixel image to obtain the detection results of each sub-image, and then fuses the detection results of the sub-images multiple times to obtain the detection result of the super-large pixel image, realizing the detection of targets in the super-large pixel image.

[0159] In an alternative embodiment of this application, the sub-image acquisition and detection module is specifically configured to:

[0160] Use a sliding window to slide on the scale image at a preset step size. Each time it slides, the area corresponding to the sliding window on the scale image is a sub-image, and the ratio of the preset step size to the width of the sliding window is within a preset ratio range;

[0161] If the sliding window exceeds the boundary of the scale image after one slide, translate the sliding window into the scale image to obtain the corresponding sub-image.

[0162] In an alternative embodiment of this application, the target detection model includes a backbone Backbone module, an intermediate Neck module, and an output module; the sub-image acquisition and detection module is specifically configured to:

[0163] Extract the target features of the sub-image through the Transformer layer and the deformable convolution DCN layer in the Backbone layer respectively; then fuse the target feature information through the Neck module to obtain the corresponding fused features; finally, the output module outputs the corresponding initial detection results based on the fused features corresponding to at least two network layers in the Neck module respectively, and obtains the detection result corresponding to the sub-image based on multiple initial detection results.

[0164] In an alternative embodiment of this application, the sub-image acquisition and detection module is further configured to:

[0165] Project all the target boxes corresponding to multiple initial detection results into the sub-image, and obtain at least one set of target boxes with the first intersection over union (IOU) not less than the first preset threshold;

[0166] For each set of target boxes with the first IOU not less than the first preset threshold, use the weighted box fusion (WBF) algorithm to obtain the corresponding fused target box based on this set of target boxes with the first IOU not less than the first preset threshold;

[0167] Based on the fused target bounding boxes corresponding to the target bounding boxes in each group with the first IOU not less than the first preset threshold, and other target bounding boxes except those in each group with the first IOU not less than the first preset threshold, obtain the detection result of the sub - graph.

[0168] In an alternative embodiment of the present application, the apparatus further includes a training module for:

[0169] For at least one super - pixel image sample annotated with detection results, downscale it by at least one preset multiple to obtain the corresponding scaled image sample, and for each scaled image sample, use a sliding window of the first preset size to slide - cut the scaled image to obtain a preset number of sub - graph samples;

[0170] Perform joint data augmentation on each sub - graph sample to obtain the corresponding data - augmented sub - graph sample, and use the preset amount of data - augmented sub - graph samples to train the initial object detection model until the loss function meets the preset conditions to obtain the trained object detection model;

[0171] Wherein, the loss function includes a target bounding box coordinate loss sub - function, and a quality focal loss QFL sub - function obtained from a target classification loss sub - function and a target bounding box confidence loss sub - function.

[0172] In an alternative embodiment of the present application, the training module is specifically used for:

[0173] Obtain the detection results in each sub - graph sample where the target bounding box is not greater than the second preset size;

[0174] Copy the target bounding box corresponding to each detection result, and after translation and rotation by a preset angle, paste it into the non - target area of the sub - graph to obtain the corresponding data - augmented sub - graph sample.

[0175] In an alternative embodiment of the present application, the first detection result fusion module is specifically used for:

[0176] Based on the segmentation method of each scaled image, splice the sub - graphs with detection results;

[0177] For the front sub - graph among two adjacent front - and - back sub - graphs, if the target bounding box corresponding to the detection result of the front sub - graph is on the left side of or intersects with the mid - line of the overlapping area of the two sub - graphs, retain the detection result; if the target bounding box corresponding to the detection result of the front sub - graph is on the right side of the mid - line, discard the detection result; for the back sub - graph among two adjacent front - and - back sub - graphs, if the target bounding box corresponding to the detection result of the back sub - graph is on the right side of the mid - line, retain the detection result; if the target bounding box corresponding to the detection result of the back sub - graph is on the left side of or intersects with the mid - line, discard the detection result.

[0178] Next, refer to Figure 11, which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application (for example, a terminal device or a server that executes the Figure 3 method shown) 1100. The electronic devices in the embodiments of the present application may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0179] The electronic device includes: a memory and a processor. The memory is used to store programs for executing the methods described in the above various method embodiments; the processor is configured to execute the programs stored in the memory. Here, the processor may be referred to as the processing device 1101 described below, and the memory may include at least one of the read-only memory (ROM) 1102, random access memory (RAM) 1103, and storage device 1108 described below, as specifically shown below:

[0180] As Figure 11 shown, the electronic device 1100 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1101, which may perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 1102 or the programs loaded from the storage device 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device 1100 are also stored. The processing device 1101, ROM 1102, and RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.

[0181] Generally, the following devices may be connected to the I / O interface 1105: an input device 1106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1108 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1109. The communication device 1109 may allow the electronic device 1100 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0182] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1109, or installed from the storage device 1108, or installed from the ROM 1102. When the computer program is executed by the processing device 1101, the above functions defined in the method of the embodiment of the present application are performed.

[0183] It should be noted that the above computer-readable storage medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0184] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0185] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0186] The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to:

[0187] Obtain an ultra-large pixel image, and reduce the ultra-large pixel image by at least one preset multiple to obtain a corresponding scaled image; for each scaled image, use a sliding window of a first preset size to perform sliding segmentation on the scaled image to obtain a corresponding plurality of sub-images, and input each sub-image into a trained object detection model to obtain the detection result of the sub-image. The trained object detection model is obtained by training with sub-image samples marked with detection results, and the detection results include target box coordinates, target categories, and target box confidence levels; fuse the detection results of each sub-image corresponding to each scaled image to obtain the detection result corresponding to the scaled image; fuse the detection results of each scaled image corresponding to the ultra-large pixel image to obtain the detection result of the ultra-large pixel image.

[0188] Computer program code for performing the operations of the present application can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0189] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0190] The modules or units described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module or unit does not constitute a limitation on the unit itself in some cases. For example, the proxy link acquisition module can also be described as "the module for acquiring proxy links".

[0191] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0192] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or Flash memory), an optical fiber, a portable Compact Disc Read-Only Memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0193] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific methods implemented when the above-described computer-readable medium is executed by an electronic device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0194] Embodiments of the present application provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that when the computer device is executed, the following situations are realized:

[0195] Obtain an ultra-large pixel image, and reduce the ultra-large pixel image by at least one preset multiple to obtain a corresponding scaled image; for each scaled image, use a sliding window with a first preset size to perform sliding segmentation on the scaled image to obtain corresponding multiple sub-images, and input each sub-image into a trained target detection model to obtain the detection result of the sub-image. The trained target detection model is obtained by training with sub-image samples marked with detection results, and the detection results include target box coordinates, target categories, and target box confidence levels; fuse the detection results of each sub-image corresponding to each scaled image to obtain the detection result corresponding to the scaled image; fuse the detection results of each scaled image corresponding to the ultra-large pixel image to obtain the detection result of the ultra-large pixel image.

[0196] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0197] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A target detection method, characterized in that, Including: Obtain an ultra-large pixel image, and downscale the ultra-large pixel image by multiple preset multiples to obtain corresponding scale images; For each scale image, use a sliding window of a first preset size to perform sliding segmentation on the scale image to obtain corresponding multiple sub-images, and input each sub-image into a trained object detection model to obtain the detection result of the sub-image. The trained object detection model is obtained by training with sub-image samples marked with detection results. The detection results include target box coordinates, target categories, and target box confidence levels. The object detection model includes a backbone Backbone module, an intermediate Neck module, and an output module; Fuse the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image; Fuse the detection results of each scale image corresponding to the ultra-large pixel image to obtain the detection result of the ultra-large pixel image; Among them, the step of inputting each sub-image into the trained object detection model to obtain the detection result of the sub-image includes: Respectively extract the target features of the sub-image through the Transformer layer and the deformable convolution DCN layer in the Backbone layer; then fuse the target feature information through the Neck module to obtain corresponding fused features; finally, the output module respectively outputs corresponding initial detection results based on the fused features corresponding to at least two network layers in the Neck module, and based on multiple initial detection results, obtain the detection result corresponding to the sub-image.

2. The method according to claim 1, characterized in that, The step of using a sliding window of a first preset size to perform sliding segmentation on each scale image to obtain corresponding multiple sub-images includes: Use the sliding window to slide on the scale image at a preset step length. After each slide, the area corresponding to the sliding window on the scale image is a sub-image. The ratio of the preset step length to the width of the sliding window is within a preset ratio range; If the sliding window exceeds the boundary of the scale image after one slide, translate the sliding window into the scale image to obtain the corresponding sub-image.

3. The method according to claim 1, characterized in that, The step of obtaining the detection result corresponding to the sub-image based on the multiple initial detection results includes: Project all target boxes corresponding to the multiple initial detection results into the sub-image, and obtain at least one group of target boxes with a first intersection over union (IOU) not less than a first preset threshold; For each group of target boxes with a first IOU not less than the first preset threshold, use the weighted box fusion (WBF) algorithm to obtain a corresponding fused target box based on the group of target boxes with a first IOU not less than the first preset threshold; Based on the fused target boxes corresponding to each group of target boxes with a first IOU not less than the first preset threshold, and other target boxes except each group of target boxes with a first IOU not less than the first preset threshold, obtain the detection result of the sub-image.

4. The method according to claim 1, characterized in that, The trained object detection model is trained in the following manner: For at least one super-pixel image sample marked with detection results, scale down the super-pixel image sample by at least one preset multiple to obtain corresponding scale image samples. For each scale image sample, use a sliding window of a first preset size to slide and divide the scale image to obtain a preset number of sub-image samples; Perform joint data augmentation on each sub-image sample to obtain corresponding data-augmented sub-image samples. Use the preset number of data-augmented sub-image samples to train an initial object detection model until the loss function meets the preset conditions to obtain the trained object detection model; Wherein, the loss function includes an object box coordinate loss sub-function, and a Quality Focal Loss (QFL) sub-function obtained from an object classification loss sub-function and an object box confidence loss sub-function.

5. The method according to claim 4, characterized in that, Performing joint data augmentation on each sub-image sample to obtain corresponding data-augmented sub-image samples includes: Obtain the detection results in each sub-image sample where the object box is not larger than a second preset size; Copy the object box corresponding to each detection result, and after translation and rotation by a preset angle, paste it into the non-object area of the sub-image to obtain the corresponding data-augmented sub-image sample.

6. The method according to claim 1, characterized in that, The fusing the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image includes: Based on the segmentation method of each scale image, splice the sub-images with detection results; For the previous sub-image among two adjacent front and back sub-images, if the object box corresponding to the detection result of the previous sub-image is located on the left side of the midline of the overlapping area of the two sub-images or intersects with the midline, retain the detection result. If the object box corresponding to the detection result of the previous sub-image is located on the right side of the midline, discard the detection result. For the subsequent sub-image among two adjacent front and back sub-images, if the object box corresponding to the detection result of the subsequent sub-image is located on the right side of the midline, retain the detection result. If the object box corresponding to the detection result of the subsequent sub-image is located on the left side of the midline or intersects with the midline, discard the detection result.

7. A target detection device, characterized in that, including: A scale image acquisition module, configured to acquire a super-pixel image and scale down the super-pixel image by multiple preset multiples to obtain corresponding scale images; A sub-image acquisition and detection module, configured to, for each scale image, use a sliding window of a first preset size to slide and divide the scale image to obtain corresponding multiple sub-images, and input each sub-image into the trained object detection model to obtain the detection result of the sub-image. The trained object detection model is obtained by training with sub-image samples marked with detection results. The detection results include object box coordinates, object categories, and object box confidences. The object detection model includes a backbone module, an intermediate Neck module, and an output module; A first detection result fusion module, configured to fuse the detection results of each sub-image corresponding to each scale image to obtain the detection result corresponding to the scale image; The second detection result fusion module is used to fuse the detection results of each scale image corresponding to the super-large pixel image to obtain the detection result of the super-large pixel image; The sub-image acquisition and detection module is specifically used to extract the target features of the sub-image through the Transformer layer and the deformable convolution DCN layer in the Backbone layer respectively; then fuse the target feature information through the Neck module to obtain the corresponding fused features; finally, based on the fused features corresponding to at least two network layers in the Neck module, the output module outputs the corresponding initial detection results respectively, and based on multiple initial detection results, obtains the detection result corresponding to the sub-image.

8. An electronic device, characterized in that, It includes a memory and a processor; A computer program is stored in the memory; The processor is used to execute the computer program to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.