Perception-driven multi-object detection method, system and device
By using a perception-driven multi-object detection method, feature map pooling and edge filling are performed using the perceived size of objects. This solves the problem of classifying objects with similar features but different perceived sizes in object detection, and improves detection accuracy and robustness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2022-07-25
- Publication Date
- 2026-05-01
AI Technical Summary
Existing computer vision object detection methods struggle to accurately distinguish between different categories of objects that have similar features but different perceived sizes.
A perception-driven multi-object detection method is proposed. By extracting object images and point cloud maps, the perceptual size of the visible part of the object is calculated. Based on the perceptual size, feature map pooling and edge filling are performed to generate a fixed-size feature map for object detection.
It enables accurate detection of objects with similar features but different perceived sizes, improving the accuracy and generalization ability of object detection.
Smart Images

Figure CN115546495B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision object detection technology, and in particular to perception-driven multi-object detection methods, systems and devices. Background Technology
[0002] The mechanism by which brain visual neural networks perceive objects in the physical world serves as an important model and source of inspiration for building computer vision object detection models. Research on brain visual object perception has revealed that the perceived size of an object is determined by the combination of retinal imaging size and scene depth information in the early stages of the human visual system. Furthermore, the larger the perceived object size, the greater its proportion in the visual field, and the larger the activated area in the primary visual cortex. Inspired by these mechanisms of human visual object perception, how to achieve accurate object detection remains a gap in current computer vision object detection technology. Summary of the Invention
[0003] The present invention aims to at least partially solve one of the technical problems in the related art.
[0004] Therefore, the purpose of this invention is to propose a perception-driven multi-object detection method, system, and apparatus. This method pools the feature map of the visible portion of an object based on its area in the physical world coordinate system (i.e., the perceived size of the object). Then, the pooled feature map is filled with edges to generate a fixed-size feature map for object detection. The object detection neural network model based on perceived size proposed in this invention can eliminate interference between different categories of objects with similar image features but different perceived sizes, thus achieving accurate object detection.
[0005] To achieve the above objectives, this invention proposes a perception-driven multi-object detection method, comprising:
[0006] Extract the first feature map and the corresponding point cloud map of the solid image;
[0007] Classify the second feature maps generated on the first feature map within each anchor frame and regress the anchor frames to obtain the anchor frame classification and regression results;
[0008] Based on the anchor frame classification and regression results and the point cloud map within the anchor frame after regression, the perceptual size of the visible part of the object in the third feature map within the anchor frame after regression is calculated, and a fourth feature map is generated based on the perceptual size and the third feature map.
[0009] The fourth feature map is used for object classification and bounding box regression, and the distance of objects within the bounding box is determined based on the point cloud map corresponding to the third feature map.
[0010] The perception-driven multi-object detection method according to embodiments of the present invention may also have the following additional technical features:
[0011] Furthermore, in one embodiment of the present invention, the extraction of the first feature map and the corresponding point cloud map of the object image includes: acquiring object images of different pixels and corresponding point cloud maps, and converting the object images of different pixels and corresponding point cloud maps into object images of a first preset size and corresponding point cloud maps; inputting the object images of the first preset size and the corresponding point cloud maps into a residual network and a feature map pyramid network to extract the first feature map and the corresponding point cloud map.
[0012] Furthermore, in one embodiment of the present invention, classifying the second feature maps within each anchor box generated on the first feature map and regressing the anchor boxes to obtain anchor box classification and regression results includes: generating a first preset number of anchor boxes on the first feature map using a region proposal network; classifying the second feature maps within the first preset number of anchor boxes as foreground and background and regressing the anchor boxes to obtain the anchor box classification and regression results.
[0013] Further, in one embodiment of the present invention, the step of calculating the perceptual size of the visible part of the object in the third feature map within the regressed anchor frame based on the anchor frame classification and regression results and the point cloud map within the regressed anchor frame includes: selecting a second preset number of foreground anchor frames as suggestion frames of a second preset size according to the classification probability from the regressed anchor frames classified as foreground, and mapping the suggestion frames of the second preset size onto the first feature map and the corresponding point cloud map; calculating the perceptual size of the visible part of the object in the third feature map within the suggestion frame based on the second preset size of the suggestion frame and the average pixel value of the point cloud map within the suggestion frame.
[0014] Furthermore, in one embodiment of the present invention, generating a fourth feature map based on the perceived size and the third feature map includes: pooling the third feature map within the suggestion box according to the perceived size to generate a feature map of proportional size; and generating a fourth feature map of a third preset size by edge filling the feature map of proportional size.
[0015] Furthermore, in one embodiment of the present invention, the step of performing object classification and bounding box regression on the fourth feature map, and determining the distance of the object within the bounding box based on the point cloud map corresponding to the third feature map, includes: performing object classification and bounding box regression on the fourth feature map of the third preset size through a fully connected layer, and using the average pixel value of the point cloud map of the third feature map within the proposed box as the distance of the object within the corresponding bounding box.
[0016] To achieve the above objectives, another aspect of the present invention provides a perception-driven multi-object detection device, comprising:
[0017] The feature extraction module is used to extract the first feature map and the corresponding point cloud map of the object image;
[0018] The region suggestion module is used to classify the second feature maps within each anchor box generated on the first feature map and regress the anchor boxes to obtain the anchor box classification and regression results.
[0019] The perception calculation module is used to calculate the perceived size of the visible part of the object in the third feature map of the third feature map after regression based on the anchor frame classification and regression results and the point cloud map in the anchor frame after regression, and to generate a fourth feature map based on the perceived size and the third feature map.
[0020] The classification and regression module is used to classify objects and regress bounding boxes on the fourth feature map, and determine the distance of objects within the bounding box based on the point cloud map corresponding to the third feature map.
[0021] To achieve the above objectives, the present invention further proposes a perception-driven multi-object detection system, characterized in that it includes:
[0022] The feature acquisition module, gateway module, local database, and object detection module are connected in sequence.
[0023] The object detection module is used for:
[0024] Extract the first feature map and the corresponding point cloud map of the solid image from the local database;
[0025] Classify the second feature maps generated on the first feature map within each anchor frame and regress the anchor frames to obtain the anchor frame classification and regression results;
[0026] Based on the anchor frame classification and regression results and the point cloud map within the anchor frame after regression, the perceptual size of the visible part of the object in the third feature map within the anchor frame after regression is calculated, and a fourth feature map is generated based on the perceptual size and the third feature map.
[0027] The fourth feature map is used for object classification and bounding box regression, and the distance of objects within the bounding box is determined based on the point cloud map corresponding to the third feature map.
[0028] The perception-driven multi-object detection method, system, and apparatus of this invention incorporate the important feature of perceived object size into the object detection model, achieving accurate object detection. This solves the problem of incorrect detection of different categories of objects with similar features but different perceived sizes.
[0029] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0030] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0031] Figure 1 This is a flowchart of a perception-driven multi-object detection method according to an embodiment of the present invention;
[0032] Figure 2 This is a schematic diagram of the neural network model structure for object detection based on object size perception according to an embodiment of the present invention;
[0033] Figure 3 A schematic diagram of ground truth visualization of the KITTI Mask dataset according to an embodiment of the present invention;
[0034] Figure 4 This is a schematic diagram of point cloud data visualization based on an embodiment of the KITTI dataset according to the present invention;
[0035] Figure 5 This is a schematic diagram illustrating the visualization of detection results using this model according to an embodiment of the present invention;
[0036] Figure 6 This is a schematic diagram illustrating the visualization of detection results using Faster RCNN according to an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of the structure of a perception-driven multi-object detection device according to an embodiment of the present invention;
[0038] Figure 8 This is a schematic diagram of the structure of a perception-driven multi-object detection system according to an embodiment of the present invention. Detailed Implementation
[0039] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0040] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0041] The perception-driven multi-object detection method, apparatus, and system according to embodiments of the present invention are described below with reference to the accompanying drawings.
[0042] The model in this embodiment of the invention takes an object image and its corresponding point cloud (generated by mapping point cloud data to the image coordinate system) as input, and outputs the object's class, 2D bounding box, and distance (i.e., the distance of the object from the camera), such as... Figure 2 As shown.
[0043] Figure 1 This is a flowchart of a perception-driven multi-object detection method according to an embodiment of the present invention.
[0044] like Figure 1 As shown, the method includes, but is not limited to, the following steps:
[0045] S1, extracts the first feature map and the corresponding point cloud map of the solid image.
[0046] Understandably, the object images and corresponding point clouds of each different pixel are first converted into object images and point clouds of a uniform size (e.g., 1024*1024, unit: pixels). Then, the object images are input into the Residual Network (ResNet) and the Feature Map Pyramid Network (FPN) to extract the feature maps of the object images.
[0047] The feature map of the object image extracted by the network in step 1 is the first feature map and the corresponding point cloud map of this embodiment of the invention.
[0048] S2, classify the second feature maps generated on the first feature map within each anchor box and regress the anchor boxes to obtain the anchor box classification and regression results.
[0049] Specifically, a certain number of anchor boxes are generated on the feature maps using a Region Proposal Network (RPN), and the feature maps within the anchor boxes are classified as foreground and background, and the anchor boxes are regressed.
[0050] Furthermore, a certain number of foreground anchor boxes are selected from the regressed anchor boxes classified as foreground according to the classification probability from large to small as proposal boxes, and these proposal boxes are mapped onto feature maps and point cloud maps.
[0051] It is understood that the feature map within the suggestion box obtained after mapping the suggestion box to the feature map is the third feature map of this embodiment of the invention.
[0052] S3. Calculate the perceptual size of the visible part of the object in the third feature map of the anchor box after regression based on the anchor box classification, regression results, and point cloud map of the anchor box after regression. Generate the fourth feature map based on the perceptual size and the third feature map.
[0053] Specifically, based on the size of each proposal box (e.g., after coordinate normalization, they are (0.05*0.05, 0.09*0.13, 0.2*0.2, ...)) and the average pixel values of the point cloud map within each proposal box (i.e., the distance of the object from the camera, e.g., 50.0, 40.0, 20.0, ..., in meters), the perceptual size (i.e., the area in the physical world coordinate system, e.g., 0.125, 0.468, 0.8, ...) of the visible part of the object within each proposal box is calculated. Then, based on the perceptual size (e.g., 0.125, 0.468, 0.8, ...), the feature maps within each proposal box are pooled using Region of Interest (RoI) pooling. Pooling generates feature maps of proportionally large sizes (e.g., 6*6, 9*9, 12*12, ...). The proportional pooling principle is: when the perceptual sizes are 0-0.1, 0.1-0.2, 0.2-0.3, 0.3-0.4, 0.4-0.5, 0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, and 0.9-1.0, the pooled sizes are 5*5, 6*6, 7*7, 8*8, 9*9, 10*10, 11*11, 12*12, 13*13, and 14*14, respectively.
[0054] Furthermore, the pooled feature maps can be padded with edge padding to generate feature maps of fixed sizes (e.g., 14*14, 14*14, 14*14, ...).
[0055] It is understood that the feature map obtained by edge filling is the fourth feature map in this embodiment of the invention.
[0056] S4. Perform object classification and bounding box regression on the fourth feature map, and determine the distance of objects within the bounding box based on the point cloud map corresponding to the third feature map.
[0057] Specifically, feature maps of fixed sizes are shared to fully connected layers (FC layers) for object classification and 2D box regression, and the average pixel values of the point cloud maps within the previously proposed boxes (e.g., 50.0, 40.0, 20.0, ..., in meters) are used as the distance to the object within the corresponding 2D box.
[0058] This invention presents a perception-driven multi-object detection method, proposing a neural network model for object detection based on the perceived size of objects. The method pools the feature maps of the visible portion of an object in a proportional manner according to the area of that visible portion in the physical world coordinate system (i.e., the perceived size of the object). Then, the pooled feature maps are filled with edges to generate a fixed-size feature map for object detection. This proposed neural network model for object detection can eliminate interference between different categories of objects with similar image features but different perceived sizes, achieving accurate object detection.
[0059] The perception-driven multi-object detection method of the present invention will be further described below with reference to the accompanying drawings.
[0060] This invention establishes an experimental platform for training and testing an object detection neural network model based on object perception size.
[0061] Preferably, in this embodiment of the invention, 7481 training images from the KITTI Mask dataset and 7481 sets of point cloud data corresponding to the training set in the KITTI dataset were used to train, validate, and test the object detection neural network model based on object perception size. Figure 3 The image shows a visualization of the ground truth of the KITTI Mask dataset. The object categories include three categories: Car, Van, and Truck. The 2D bounding boxes of the visible parts of the objects are labeled. Figure 4The image shows the corresponding visualized point cloud data from the KITTI dataset. The point cloud data was obtained by LiDAR scanning with a horizontal field of view of 90° and a detection distance of approximately 70 meters. In the experiment, 7481 training set images / point cloud data were divided into 4488 training set images / point cloud data, 1495 validation set images / point cloud data, and 1498 test set images / point cloud data. The object 2D bounding box evaluation metrics included three levels: easy, medium, and hard. The 2D bounding box intersection-over-union (IoU) threshold was set to 0.7. The three difficulty levels (easy, medium, and hard) were determined based on the minimum height (40, 25, and 25 pixels, in pixels) of the ground truth and the detected object's 2D bounding box, the maximum occlusion level (0, 1, and 2) of the ground truth, and the maximum truncation level (0.15, 0.3, and 0.5), respectively. In the training experiments, we pre-trained the Residual Network (ResNet), Feature Map Pyramid Network (FPN), and Region Proposal Network (RPN) in this model using the MS-COCO dataset, and fine-tuned them using the KITTIMask dataset. We trained them in 50 batches at a constant learning rate of 0.0001 using a 1080TI GPU and a 1.7GHz CPU.
[0062] Furthermore, the experimental results of the object detection neural network model based on perceived size (PerceivedSize RCNN) proposed in this invention on the test set are shown in Table 1. AP2D represents the average accuracy of 2D bounding boxes for the three categories of cars, vans, and trucks. It can be seen that, compared with the current state-of-the-art Faster RCNN object detection algorithm, the object detection neural network model based on perceived size proposed in this invention achieves higher average accuracy of 2D bounding boxes for cars, vans, and trucks across three difficulty levels: Easy (E), Medium (M), and Hard (H). In particular, the average accuracy of 2D bounding boxes for vans and trucks is significantly higher than that of Faster RCNN. Faster RCNN is the main network in this model without incorporating perceived size. The results show that Faster RCNN struggles to distinguish between cars, vans, and trucks with similar object features. This model can effectively distinguish between cars, vans, and trucks with similar object features, exhibiting higher object detection accuracy, better generalization ability, and better robustness.
[0063] Table 1
[0064]
[0065] To better demonstrate the superior object detection performance of the proposed object detection neural network model based on object size perception, and to analyze the reasons for its better detection results, this invention provides visualized detection results of some images detected using this model and Faster R-CNN on a test set, as shown below. Figure 5 and Figure 6 As shown in the figure. The detection results of this model are as follows: Figure 5 As shown, this model can accurately detect cars and large black vans located near the center of the image; the detection results of Faster R-CNN are as follows. Figure 6 As shown, Faster R-CNN incorrectly detects the large black van near the center of the image as a car. The reason Faster R-CNN cannot accurately detect the car and van in the image is that the visible parts of the car and van have very similar image features, making accurate detection difficult based solely on image features. Faster R-CNN is an algorithm that relies solely on object image features for object detection, thus failing to accurately detect the car and van in the situation shown, resulting in a lower object detection accuracy. Our model can accurately detect the car and van in the image because it detects objects not only based on image features but also on the perceived size of the object calculated from the 2D bounding box of the visible part of the image and the corresponding distance. Therefore, although the image features of the car and van are very similar in the situation shown, making accurate detection difficult, their perceived sizes are different. Thus, our model can accurately distinguish them based on their perceived sizes.
[0066] In summary, the object detection neural network model based on object perceived size proposed in this invention can accurately detect objects based on the perceived size of the visible part of the object and image features.
[0067] According to the perception-driven multi-object detection method of the present invention, the mutual interference between different categories of objects with similar image features but different perceived sizes can be eliminated based on the perceived size of the objects, thereby achieving accurate object detection.
[0068] To achieve the above embodiments, such as Figure 7 As shown, this embodiment also provides a perception-driven multi-object detection device 10, which includes: a feature extraction module 100, a region proposal module 200, a perception computing module 300, and a classification and regression module 400.
[0069] Feature extraction module 100 is used to extract the first feature map and the corresponding point cloud map of the object image;
[0070] The region proposal module 200 is used to classify the second feature maps within each anchor box generated on the first feature map and to regress the anchor boxes to obtain the anchor box classification and regression results.
[0071] The perception computing module 300 is used to calculate the perceived size of the visible part of the object in the third feature map of the anchor box after regression based on the anchor box classification and regression results and the point cloud map in the anchor box after regression, and to generate a fourth feature map based on the perceived size and the third feature map.
[0072] The classification and regression module 400 is used to classify objects and regress bounding boxes on the fourth feature map, and to determine the distance of objects within the bounding box based on the point cloud map corresponding to the third feature map.
[0073] Furthermore, the aforementioned feature extraction module 100 is also used for:
[0074] Acquire object images of different pixels and their corresponding point cloud maps, and convert the object images of different pixels and their corresponding point cloud maps into object images of a first preset size and their corresponding point cloud maps.
[0075] The object image of the first preset size and the corresponding point cloud map are input into the residual network and the feature map pyramid network to extract the first feature map and the corresponding point cloud map.
[0076] Furthermore, the aforementioned region suggestion module 200 is also used for:
[0077] A first preset number of anchor boxes are generated on the first feature map using a region proposal network;
[0078] The second feature map within a first preset number of anchor frames is classified into foreground and background, and the anchor frames are regressed to obtain the anchor frame classification results.
[0079] Furthermore, the aforementioned perception computing module 300 is also used for:
[0080] From the anchor boxes that have been regressed and are classified as foreground, select a second preset number of foreground anchor boxes according to the classification probability as suggestion boxes of a second preset size, and map the suggestion boxes of the second preset size onto the first feature map and the corresponding point cloud map;
[0081] Based on the second preset size of the suggestion box and the average pixel value of the point cloud map within the suggestion box, calculate the perceptual size of the visible part of the object in the third feature map within the suggestion box.
[0082] Based on the perceived size, the third feature map within the suggestion box is pooled through the region of interest to generate a feature map of proportional size. The feature map of proportional size is then used to generate a fourth feature map of a third preset size through edge padding.
[0083] Furthermore, the aforementioned classification and regression module 400 is also used for:
[0084] The fourth feature map of the third preset size is passed through a fully connected layer for object classification and bounding box regression, and the average pixel value of the point cloud map of the proposed box is used as the distance of the object in the corresponding bounding box.
[0085] According to the perception-driven multi-object detection device of the present invention, the mutual interference between different categories of objects with similar image features but different perceived sizes can be eliminated based on the perceived size of the objects, thereby achieving accurate object detection.
[0086] To achieve the above embodiments, such as Figure 8 As shown, this embodiment also provides a perception-driven multi-object detection system 20, including:
[0087] The feature acquisition module 201, gateway module 202, local database 203, and object detection module 204 are connected in sequence.
[0088] Object detection module 204 is used for:
[0089] Extract the first feature map and the corresponding point cloud map of the object image from the local database 203;
[0090] Classify the second feature maps generated on the first feature map and regress the anchor boxes to obtain the anchor box classification and regression results;
[0091] The perceptual size of the visible part of the object in the third feature map of the anchor frame after regression is calculated based on the anchor frame classification and regression results and the point cloud map in the anchor frame after regression. The fourth feature map is generated based on the perceptual size and the third feature map.
[0092] Object classification and bounding box regression are performed on the fourth feature map, and the distance of objects within the bounding box is determined based on the point cloud map corresponding to the third feature map.
[0093] According to the perception-driven multi-object detection system of the present invention, the mutual interference between different categories of objects with similar image features but different perceived sizes can be eliminated based on the perceived size of the objects, thereby achieving accurate object detection.
[0094] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0095] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0096] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A perception-driven multi-object detection method, characterized in that, Includes the following steps: Extract the first feature map and the corresponding point cloud map of the solid image; A first preset number of anchor boxes are generated on the first feature map using a region proposal network; Classify the foreground and background of the second feature map within the first preset number of anchor frames and perform regression on the anchor frames to obtain the anchor frame classification and regression results; From the anchor boxes that have been regressed and are classified as foreground, a second preset number of foreground anchor boxes are selected according to the classification probability as suggestion boxes of a second preset size, and the suggestion boxes of the second preset size are mapped onto the first feature map and the corresponding point cloud map; Based on the second preset size of the suggestion box and the average pixel value of the point cloud map within the suggestion box, the perceptual size of the visible part of the object in the third feature map within the suggestion box is calculated, and a fourth feature map is generated based on the perceptual size and the third feature map. The fourth feature map is used for object classification and bounding box regression, and the distance of objects within the bounding box is determined based on the point cloud map corresponding to the third feature map.
2. The method according to claim 1, characterized in that, The extracted first feature map and corresponding point cloud map of the body image include: Obtain object images of different pixels and their corresponding point cloud maps, and convert the object images of different pixels and their corresponding point cloud maps into object images of a first preset size and their corresponding point cloud maps. The object image of the first preset size and the corresponding point cloud map are input into the residual network and the feature map pyramid network to extract the first feature map and the corresponding point cloud map.
3. The method according to claim 2, characterized in that, The step of generating a fourth feature map based on the perceived size and the third feature map includes: Based on the perceived size, the third feature map within the suggestion box is pooled through the region of interest to generate a feature map of proportional size. The feature map of the proportional size is used to generate a fourth feature map of a third preset size through edge filling operation.
4. The method according to claim 3, characterized in that, The step of performing object classification and bounding box regression on the fourth feature map, and determining the distance of objects within the bounding box based on the point cloud map corresponding to the third feature map, includes: The fourth feature map of the third preset size is used for object classification and bounding box regression through a fully connected layer, and the average pixel value of the point cloud map of the third feature map within the proposed box is used as the distance of the object within the corresponding bounding box.
5. A perception-driven multi-object detection device, characterized in that, include: The feature extraction module is used to extract the first feature map and the corresponding point cloud map of the object image; The region suggestion module is used to classify the second feature maps within each anchor box generated on the first feature map and regress the anchor boxes to obtain the anchor box classification and regression results. The perception calculation module is used to calculate the perceived size of the visible part of the object in the third feature map of the third feature map after regression based on the anchor frame classification and regression results and the point cloud map in the anchor frame after regression, and to generate a fourth feature map based on the perceived size and the third feature map. The classification and regression module is used to classify objects and regress bounding boxes on the fourth feature map, and determine the distance of objects within the bounding box based on the point cloud map corresponding to the third feature map.
6. The apparatus according to claim 5, characterized in that, The feature extraction module is also used for: Obtain object images of different pixels and their corresponding point cloud maps, and convert the object images of different pixels and their corresponding point cloud maps into object images of a first preset size and their corresponding point cloud maps. The object image of the first preset size and the corresponding point cloud map are input into the residual network and the feature map pyramid network to extract the first feature map and the corresponding point cloud map.
7. The apparatus according to claim 6, characterized in that, The perception computing module is also used for: Based on the perceived size, the third feature map within the suggestion box is pooled through the region of interest to generate a feature map of proportional size. The feature map of proportional size is then filled with edges to generate a fourth feature map of a third preset size.
8. The apparatus according to claim 7, characterized in that, The classification and regression module is also used for: The fourth feature map of the third preset size is used for object classification and bounding box regression through a fully connected layer, and the average pixel value of the point cloud map of the third feature map within the proposed box is used as the distance of the object within the corresponding bounding box.
9. A perception-driven multi-object detection system, characterized in that, include: The feature acquisition module, gateway module, local database, and object detection module are connected in sequence. The object detection module is used for: Extract the first feature map and the corresponding point cloud map of the solid image from the local database; A first preset number of anchor boxes are generated on the first feature map using a region proposal network; Classify the foreground and background of the second feature map within the first preset number of anchor frames and perform regression on the anchor frames to obtain the anchor frame classification and regression results; From the anchor boxes that have been regressed and are classified as foreground, a second preset number of foreground anchor boxes are selected according to the classification probability as suggestion boxes of a second preset size, and the suggestion boxes of the second preset size are mapped onto the first feature map and the corresponding point cloud map; Based on the second preset size of the suggestion box and the average pixel value of the point cloud map within the suggestion box, the perceptual size of the visible part of the object in the third feature map within the suggestion box is calculated, and a fourth feature map is generated based on the perceptual size and the third feature map. The fourth feature map is used for object classification and bounding box regression, and the distance of objects within the bounding box is determined based on the point cloud map corresponding to the third feature map.
Citation Information
Patent Citations
Object recognition model training method and device and object recognition method and system
CN112287860A
Pluggable aerial image target positioning detection method, system and device
CN112364843A