A target detection method, device and computer device

By projecting point clouds and images on the graphics processor, feature extraction and fusion, the problem of large transmission delay and high computational volume between devices is solved, and end-to-end real-time object detection is realized, improving the accuracy and speed of detection.

CN114140758BActive Publication Date: 2025-07-25YAOYAO
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111450576.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-30
Publication Date
2025-07-25
Estimated Expiration
2041-11-30

AI Technical Summary

Technical Problem

In the prior art, the fusion algorithm of image and lidar cannot be implemented simultaneously on the graphics processor, resulting in large delays in data transmission between devices, high calculation amounts and serious time-consuming, and real-time object detection cannot be achieved.

Method used

Projection, feature extraction and fusion of point clouds and images are carried out on the graphics processor, downsampling and mapping of point clouds are realized using CUDA, and end-to-end real-time object detection is completed in combination with neural networks. Category prediction and bounding box prediction of target areas are completed on the GPU through projection and fusion.

Benefits of technology

The end-to-end real-time operation of the image and lidar fusion algorithm on the graphics processor is realized, which improves the accuracy and speed of detection and reduces the computational complexity and memory footprint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114140758B_ABST
    Figure CN114140758B_ABST
Patent Text Reader

Abstract

The present application provides a target detection method, apparatus and computer device. The method is applied to a graphics processing unit and includes: obtaining an original point cloud and an original image corresponding to a target area; projecting each sparse point in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and a pixel point; extracting the point cloud features of the original point cloud and the image features of the original image; fusing each sparse point feature with the corresponding pixel point feature according to the corresponding relationship between each sparse point and the pixel point to obtain a target fusion feature corresponding to the target area; performing class prediction and bounding box prediction on the target area based on the target fusion feature to obtain a detection target. In the present application, the entire target detection process including projection and fusion is completed on the graphics processing unit, which can achieve end-to-end real-time operation, and the target fusion feature includes high-level semantic information, increasing the accuracy of detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition, and in particular, to an object detection method, apparatus, and computer device. Background Art

[0002] In an autonomous driving perception system, cameras and lidars are essential sensors for autonomous vehicles. Cameras can collect RGB color information and texture information of the surrounding environment, simulating human visual perception imaging. The advantage of cameras is that they can accurately describe the texture information of objects, but lack the depth information of objects. Lidars rely on continuous scanning of laser beams to complete the scene reproduction of the surrounding environment. Laser beams can generate laser points on the surface of objects. It can collect the accurate XYZ coordinates and reflectivity of the surrounding environment in the radar coordinate system. The advantage of lidars is that they can obtain the depth information of objects, but lack the texture information of objects. The texture information of objects and the lack of depth information can be retained through the fusion algorithm of images and lidars.

[0003] However, there are two major problems in the existing 3D object detection based on the fusion of images and laser point clouds. One is that the overall design of the fusion algorithm of images and lidars is relatively complex. From the input of point clouds and images to the output of detection results, the entire model cannot be implemented on a Graphics Processing Unit (GPU) at the same time, resulting in multiple transmissions of data between devices during the processing process, causing extremely large delays and unable to meet practical applications. The other is that the neural network part and the non-neural network data processing part of the existing algorithms are complicated, time-consuming, and occupy a large amount of memory, resulting in a high complexity and large computational amount of the algorithm model. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides an object detection method, apparatus, and computer device. The specific solutions are as follows:

[0005] In a first aspect, an embodiment of the present application provides an object detection method, the method including:

[0006] Obtain the original point cloud and original image corresponding to the target area, where the original point cloud includes a plurality of sparse points, and the original image includes a plurality of pixel points;

[0007] Project each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel points;

[0008] Extract the point cloud features of the original point cloud and extract the image features of the original image, where the point cloud features include a plurality of sparse point features, and the image features include a plurality of pixel point features;

[0009] According to the corresponding relationship between each sparse point and the pixel point, fuse the sparse point features with the corresponding pixel point features to obtain the target fusion feature corresponding to the target area;

[0010] Based on the target fusion feature, perform class prediction and bounding box prediction on the target area to obtain the detection target.

[0011] According to a specific implementation manner disclosed in the present application, the step of projecting each sparse point in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel point includes:

[0012] Based on the formula , establish the corresponding relationship between the sparse point and the pixel point, where is the two-dimensional coordinate value of the pixel point in the image coordinate system, is the projection matrix from the camera coordinate system to the image coordinate system, with a size of 3*4, is the rotation matrix of the camera, with a size of 4*4, is the projection matrix from the radar to the camera, with a size of 4*4, is the three-dimensional coordinate value of the sparse point in the point cloud coordinate system.

[0013] According to a specific implementation manner disclosed in the present application, the point cloud data corresponding to each sparse point includes three-dimensional coordinate values and reflectivity. The steps of extracting the point cloud features of the original point cloud include;

[0014] Perform parallel downsampling on the original point cloud;

[0015] Based on the point cloud data corresponding to each sparse point in the downsampled original point cloud, extract the sparse point features and neighborhood features corresponding to each sparse point. Among them, taking any one of the sparse points as the key point, the sparse points within the preset radius range are the neighboring points corresponding to the key point, and the neighborhood feature is composed of the concatenation of the point cloud data of the neighboring points corresponding to the sparse point;

[0016] Fuse each sparse point feature and the corresponding neighborhood feature into the point cloud feature.

[0017] According to a specific implementation manner disclosed in the present application, the steps of determining the neighboring points of each sparse point include:

[0018] Judge whether the number N of sparse points within the preset radius range centered on the key point is greater than or equal to the preset number M, where N is a positive integer;

[0019] If N M, arrange the N sparse points in ascending order according to the distance between each sparse point and the key point, and determine the first M sparse points corresponding to the order as the neighboring points of the key point;

[0020] If N < M, arrange the N sparse points in ascending order according to the distances between the sparse points and the key points, and determine the first M - N sparse points as supplementary points. Duplicate the supplementary points, and determine the N sparse points and the M - N supplementary points as the neighboring points of the key points.

[0021] According to a specific implementation manner disclosed in the present application, the step of fusing each sparse point feature with the corresponding pixel point feature according to the correspondence between the sparse points and the pixel points to obtain the fusion feature corresponding to the target region includes:

[0022] Based on the correspondence between each sparse point and the pixel point, fuse each sparse point feature with the corresponding pixel point feature to obtain a first fusion feature;

[0023] Interpolate the first fusion feature to obtain a second fusion feature;

[0024] Extract the high-level semantic feature in the second fusion feature through two Linear-BN-ReLU layers as the target fusion feature corresponding to the target region.

[0025] According to a specific implementation manner disclosed in the present application, the step of interpolating the first fusion feature to obtain a second fusion feature includes:

[0026] Select any one of the sparse points corresponding to the first fusion feature as the original point;

[0027] Arrange all the sparse points corresponding to the first fusion feature in ascending order according to the distance values between the sparse points and the original point to obtain a first sequence;

[0028] Select the first K sparse points in the first sequence as the associated points of the original point, where K is a positive integer;

[0029] Normalize the distances from each associated point to the original point to obtain the weights of each associated point;

[0030] Multiply the weight corresponding to each associated point by the sparse point feature corresponding to each associated point to obtain the upsampling feature of each original point;

[0031] Combine the upsampling features of each sparse point corresponding to the first fusion feature into the second fusion feature.

[0032] According to a specific implementation manner disclosed in the present application, the step of performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain the detection target includes:

[0033] Perform class prediction and bounding box prediction on the target region based on the target fusion feature to obtain a first bounding box with multiple different class scores corresponding to the target class;

[0034] Sort the multiple first bounding boxes according to the class scores to obtain a second sequence;

[0035] Repeat the step of selecting a target bounding box from the second sequence until all target bounding boxes are found;

[0036] Determine the target object corresponding to each target bounding box as the detection target;

[0037] Among them, the step of repeating the step of selecting a target bounding box from the second sequence until all target bounding boxes are found includes:

[0038] Select the first bounding box with the largest class score in the second sequence as the target bounding box according to a preset rule;

[0039] Retain the first bounding boxes with an overlap degree less than or equal to a preset threshold as the second bounding boxes, where the overlap degree is the ratio of the area of the intersection part of the first bounding box and the target bounding box to the area of the union part;

[0040] Sort each of the second bounding boxes according to the class scores into a third sequence, and use the third sequence as the new second sequence.

[0041] In a second aspect, an embodiment of the present application provides a target detection device applied to a graphics processor. The device includes:

[0042] An acquisition module, configured to acquire the original point cloud and the original image corresponding to the target region, where the original point cloud includes multiple sparse points, and the original image includes multiple pixel points;

[0043] A projection module, configured to project each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel point;

[0044] An extraction module, configured to extract the point cloud feature of the original point cloud and extract the image feature of the original image, where the point cloud feature includes multiple sparse point features, and the image feature includes multiple pixel point features;

[0045] A fusion module, configured to fuse each sparse point feature and the corresponding pixel point feature according to the corresponding relationship between each sparse point and the pixel point to obtain the target fusion feature corresponding to the target region;

[0046] The detection module, or the method described in any of the embodiments in the second aspect, is used to perform class prediction and bounding box prediction on the target region based on the target fusion feature to obtain a detection target.

[0047] In a third aspect, an embodiment of the present application provides a computer device, which includes a graphics processor and a memory. The memory stores a computer program, and when the computer program is executed on the graphics processor, it implements the target detection method described in any of the embodiments in the first aspect.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed on a processor, it implements the target detection method described in any of the embodiments in the first aspect.

[0049] Compared with the prior art, the present application has the following beneficial effects:

[0050] The target detection method provided by the present application is applied to a graphics processor and includes: obtaining the original point cloud and the original image corresponding to the target region; projecting each sparse point in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel point; extracting the point cloud feature of the original point cloud and the image feature of the original image; fusing each sparse point feature with the corresponding pixel point feature according to the corresponding relationship between each sparse point and the pixel point to obtain the target fusion feature corresponding to the target region; performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain a detection target. In the present application, the entire target detection process including projection and fusion is completed on the graphics processor, which can achieve end-to-end real-time operation, and the target fusion feature includes high-level semantic information, increasing the accuracy of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the protection scope of the present invention. In each drawing, similar components are numbered similarly.

[0052] Figure 1 FIG. 1 is a schematic flowchart of a target detection method provided by an embodiment of the present application;

[0053] Figure 2 FIG. 2 is a schematic flowchart of a target detection method provided by an embodiment of the present application;

[0054] Figure 3 FIG. 3 is a block diagram of a target detection device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0056] The components of the embodiments of the present invention generally described and illustrated in the figures herein may be arranged and designed in a variety of different configurations. Thus, the detailed description of the embodiments of the present invention provided in the figures is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.

[0057] In the following text, the terms "including", "having" and their cognates used in various embodiments of the present invention are only intended to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or precluding the possibility of adding one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0058] In addition, the terms "first", "second", "third", etc. are only used for descriptive distinction and should not be construed as indicating or implying relative importance.

[0059] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present invention belong. The terms (such as those defined in a general use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or being overly formal, unless clearly defined in the various embodiments of the present invention.

[0060] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments may be combined with each other.

[0061] Currently, the fusion algorithms based on images and lidar are mainly divided into three different types: pre-fusion, depth fusion, and post-fusion. Among them, depth fusion includes three steps: pre-processing, intermediate processing, and post-processing. Pre-processing includes the projection of point clouds to images and the downsampling of point clouds, which is implemented on the Central Processing Unit (CPU) and takes a long time; the intermediate processing is implemented on the Graphics Processing Unit (GPU) and includes a point cloud feature extraction branch, an image feature extraction branch, and a fusion branch; and the post-processing includes non-maximum suppression and bounding box decoding, which takes a long time to implement on the CPU side.

[0062] Generally speaking, the data processing of the entire fusion network needs to be transmitted multiple times on different hardware devices, and each process takes a relatively long time, resulting in the current fusion algorithm being unable to perform real-time inference and application.

[0063] See Figure 1 and Figure 2 , Figure 1 is one of the schematic flowcharts of an object detection method provided by an embodiment of the present application. Figure 2 is the second schematic flowchart of an object detection method provided by an embodiment of the present application. The object detection method is applied to a graphics processing unit. As Figure 1 shown, the method mainly includes:

[0064] Step S101, obtain the original point cloud and the original image corresponding to the target area, where the original point cloud includes a plurality of sparse points, and the original image includes a plurality of pixel points.

[0065] When performing object detection, the graphics processing unit can, according to the actual needs of the user, select any spatial area as the target area, and respectively obtain the original point cloud and the original image corresponding to the target area, that is, Figure 2 the input point cloud and the input image shown in. Among them, the original point cloud is composed of sparse points, and the original image is composed of pixel points. The original point cloud can be collected by lidar, and the original image can be collected by a camera. Among them, the point data corresponding to each sparse point includes its three-dimensional coordinate values (X, Y, Z) in the point cloud coordinate system, and the reflectivity of each sparse point. The pixel point data corresponding to each pixel point is composed of RGB three-channel data.

[0066] Step S102, project each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel points.

[0067] After obtaining the original point cloud and the original image corresponding to the target area, according to the pre-set calibration parameters, through coordinate system conversion, each sparse point in the original point cloud in the radar coordinate system can be converted to the image coordinate system, that is, the original point cloud is projected onto the original image, and the Figure 2 point cloud mapping image shown in is obtained, so as to realize the one-to-one correspondence between the sparse points in the original point cloud and the pixel points in the original image, and establish the corresponding relationship between each sparse point and the pixel point. Among them, the calibration parameters include the rigid body transformation matrix from the radar coordinate system to the camera coordinate system, and the projection matrix from the camera coordinate system to the image coordinate system.

[0068] Specifically, the corresponding relationship between each sparse point and the pixel point can be calculated by the following formula:

[0069] ,

[0070] where is the two-dimensional coordinate value of the pixel point in the image coordinate system, is the projection matrix from the camera coordinate system to the image coordinate system, with a size of 3*4, is the rotation matrix of the camera, with a size of 4*4, is the projection matrix from the radar to the camera, with a size of 4*4, is the three-dimensional coordinate value of the sparse point in the point cloud coordinate system.

[0071] Step S103, extract the point cloud features of the original point cloud and extract the image features of the original image, where the point cloud features include multiple sparse point features, and the image features include multiple pixel point features.

[0072] After obtaining the original point cloud and the original image corresponding to the target area, extract the point cloud features of the original point cloud and the image features of the original image respectively. Specifically, the point clouds in each frame of the original point cloud can be sampled at a fixed value first, that is, the point cloud collected in each frame is downsampled or upsampled to the same value, which is convenient for subsequent alignment and sampling processing. The principle of downsampling in fixed-value sampling is formulated according to the radar scanning principle. The point cloud scanned by the radar is usually dense near and sparse far away. Therefore, sampling is only carried out in the perceptible region of interest. In the field of image processing, the region of interest is the region where the user performs target detection or analysis. A large number of distant points are retained during the sampling process, so as to ensure that the features represented by the distant points will not be lost, while the nearby points are randomly sampled to ensure that the algorithm maintains a certain robustness during the training iteration process.

[0073] The process of extracting the point cloud features of the original point cloud and the process of extracting the image features of the original image are introduced separately below.

[0074] Regarding the extraction process of point cloud features, that is, the point cloud data corresponding to each of the above-mentioned sparse points includes three-dimensional coordinate values and reflectivity, the steps of extracting the point cloud features of the original point cloud include:

[0075] Performing parallel downsampling on the original point cloud;

[0076] Based on the point cloud data corresponding to each of the sparse points in the downsampled original point cloud, extracting the sparse point features and neighborhood features corresponding to each sparse point. Among them, taking any one of the sparse points as the key point, the sparse points within a preset radius are the neighboring points corresponding to the key point, and the neighborhood feature is composed of the concatenation of the point cloud data of the neighboring points corresponding to the sparse point;

[0077] Fusing each of the sparse point features and the corresponding neighborhood features into the point cloud features.

[0078] In specific implementation, the three-dimensional space corresponding to the original point cloud is divided into several voxels through grid-based downsampling, and some points are taken inside each voxel, ultimately achieving the purpose of reducing the point cloud resolution. The above downsampling process can also be replaced by farthest point sampling or random sampling. During the downsampling process, grid parallel sampling can be implemented based on the Compute Unified Device Architecture (CUDA) to ensure the sampling speed and greatly reduce the time consumption during the sampling process. CUDA is a parallel computing platform and programming model, through which the GPU can be conveniently used for general computing.

[0079] After sampling, the original point cloud is more evenly distributed in three-dimensional space. To extract the semantic information of each sparse point, the sparse point features and neighborhood features corresponding to each sparse point are required. Among them, semantic information is divided into three different types: visual layer, object layer, and concept layer. The visual layer includes color, texture, and shape, etc., and these features are all called low-level features or low-level semantic information; the object layer is also called the intermediate layer and contains attribute features, which are used to describe the state of an object at a certain moment; while the concept layer is the high-level layer, which is used to express things closest to human understanding. For example, in a certain visual area, there is sand, blue sky, and sea water. The visual layer is to distinguish them piece by piece, the object layer is sand, blue sky, and sea water, and the concept layer is the beach. The sparse point features and neighborhood features corresponding to each sparse point can be extracted based on the point cloud data corresponding to each sparse point in the downsampled original point cloud. Among them, taking any one of the sparse points as the key point, the sparse points within a preset radius are the neighboring points corresponding to the key point, and the neighborhood feature is composed of the concatenation of the point cloud data of the neighboring points corresponding to the sparse point.

[0080] In the above process of extracting point cloud features, the steps of determining the neighboring points of each sparse point include:

[0081] Determine whether the number N of sparse points within the preset radius centered on the key point is greater than or equal to the preset number M, where N is a positive integer;

[0082] If N ≥M, sort the N sparse points in ascending order according to the distance between each sparse point and the key point, and determine the first M sparse points in the sorted order as the neighboring points of the key point;

[0083] If N < M, sort the N sparse points in ascending order according to the distance between each sparse point and the key point, determine the first M - N sparse points as supplementary points, copy the supplementary points, and determine the N sparse points and the M - N supplementary points as the neighboring points of the key point.

[0084] In specific implementation, the neighboring points corresponding to the sparse points within the preset radius of each sparse point and with the distance value ranking among the top 16 or 32 can be found by means of traversal. Further, in order to reduce the time consumption caused by traversal, a CUDA-based "neighboring point query and grouping" can be used to quickly query neighboring points. By querying the points closer in distance among the 27 grids closest to the sampling point or the 125 grids of the second closest neighbors in the downsampled original point cloud. If the number of neighboring points is less than 16 or 32, the closest sparse points are copied to reach the preset number M required in the calculation, that is, 16 or 32 mentioned above. In specific implementation, the specific value of M can be customized according to actual usage requirements and application scenarios, and no specific limitation is made here.

[0085] In specific implementation, in order to extract deeper or more features, the downsampling process can be divided into multiple stages, such as four stages: successively performing downsampling, query grouping, and feature aggregation at 1 / 4, 1 / 16, 1 / 64, and 1 / 256, for example Figure 2 as shown in "1 / 4 downsampling + grouped aggregation features" to increase the receptive field of each aggregation. Since the number of features obtained by the preset receptive field is a fixed value, if the degree of downsampling is greater, then the radius of the receptive field needs to be expanded so that the receptive field can obtain local features in a larger range.

[0086] For the image features in the original image, they can be extracted through a neural network model. See Figure 2 . The extraction of image features mainly consists of neural network layers, including conv convolution, bn normalization, and relu activation functions. The role of these neural network layers is to extract local features in the image. Similarly, in order to extract deeper-level features, we set four feature extraction layers to extract the deep features of the image. The dimension of each feature extraction layer is consistent with the feature dimension of the original point cloud after downsampling at each stage, facilitating the next fusion.

[0087] Step S104: According to the correspondence between each sparse point and the pixel point, fuse each sparse point feature with the corresponding pixel point feature to obtain the target fusion feature corresponding to the target area.

[0088] After obtaining the point cloud feature and the image feature in step S103, based on the correspondence between each sparse point and the pixel point, fuse each sparse point feature with the corresponding pixel point feature to obtain the first fusion feature. Interpolate the first fusion feature to obtain the second fusion feature. This first fusion feature is the fusion feature corresponding to the last layer in the downsampling process in step S103. Therefore, feature interpolation is required to interpolate the first fusion feature to restore the low-resolution dimensional feature to the high-resolution original-sized point cloud. Then, through two Linear-BN-ReLU layers, extract the high-level semantic feature in the second fusion feature as the target fusion feature corresponding to the target area.

[0089] In specific implementation, the correspondence between each sparse point and the pixel point obtained in step S102 can be used to match the sparse points in the original point cloud after four times of downsampling with the pixel points respectively, so as to ensure that the sparse point feature and the pixel point feature are fused at the same position. When fusing features, feature extraction can also be performed on the neighborhood features of each sparse point again to make the distribution of features relatively consistent, that is, normalize the features to ensure that the sparse point feature and the pixel point feature are of the same order of magnitude, and then fuse the features layer by layer. The above implementation steps are Figure 2 the process of "feature fusion and alignment" shown.

[0090] The step of interpolating the first fusion feature to obtain the second fusion feature includes:

[0091] Select any one of the sparse points corresponding to the first fusion feature as the original point;

[0092] Arrange all the sparse points corresponding to the first fusion feature in ascending order according to the distance value between the sparse point and the original point to obtain the first sequence;

[0093] Select the first K sparse points in the first sequence as the associated points of the original point, where K is a positive integer;

[0094] Normalize the distances from each associated point to the original point to obtain the weights of each associated point;

[0095] Multiply the weight corresponding to each associated point by the sparse point feature corresponding to each associated point to obtain the upsampled feature of each original point;

[0096] Combine the upsampled features of each sparse point corresponding to the first fusion feature into the second fusion feature.

[0097] Specifically, any sparse point corresponding to the first fusion feature can be selected as the original point, and CUDA is used for parallel processing to query the K sparse points with relatively close distances in the downsampled original point cloud for the original point. For example, K can be set to 3. The sparse point features corresponding to these 3 sparse points are taken as the features of the original point. The distances from these 3 sparse points to the original point are normalized to obtain the weights of these 3 sparse points, and then the 3 weights are respectively multiplied by the sparse point feature values corresponding to the 3 sparse points as the upsampled features of the original point.

[0098] In specific implementation, the Linear-BN-ReLU layer in the neural network processing layer can be replaced by Conv2d-BN-Relu / LeakyRelu / Relu, which can also achieve the purpose of feature extraction. The embodiments of the present application do not make specific limitations.

[0099] Step S105: Based on the target fusion feature, perform class prediction and bounding box prediction on the target region to obtain the detection target.

[0100] After obtaining the fusion feature corresponding to the target region, perform class prediction and bounding box prediction based on the fusion feature. Class prediction is used to determine the classes of each target to be detected in the target region, such as cats, dogs, cars, etc. Class prediction will output the scores of several classes corresponding to the target to be detected. After passing through the sigmoid function, the index with the highest score is taken as its corresponding class. The bounding box prediction is used to obtain the relative length, width, and height, the absolute values of the three-dimensional coordinates, and the orientation angle of the target to be detected.

[0101] After performing class prediction and bounding box prediction on the target region based on the target fusion feature, perform bounding box decoding and non-maximum suppression on the output detection result. Bounding box decoding is to calculate the true offset, length, width, and height, and orientation angle of the target by resolving the relative offset of the bounding box prediction. Using non-maximum suppression, the bounding boxes with overlapping between the same target classes can be filtered, and the bounding box with the highest score corresponding to the target class is retained as the final target bounding box.

[0102] The step of performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain the detection target includes:

[0103] Performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain a first bounding box with multiple different class scores corresponding to the target class;

[0104] Sorting the multiple first bounding boxes according to the class scores to obtain a second sequence;

[0105] Repeat the step of selecting the target bounding box from the second sequence until all the target bounding boxes are found;

[0106] Determine the target objects corresponding to each of the target bounding boxes as the detection targets;

[0107] Among them, the step of repeating the step of selecting the target bounding box from the second sequence until all the target bounding boxes are found includes:

[0108] Select the first bounding box with the largest class score in the second sequence as the target bounding box according to the preset rule;

[0109] Keep the first bounding boxes with an overlap degree less than or equal to the preset threshold as the second bounding boxes, where the overlap degree is the ratio of the area of the intersection part of the first bounding box and the target bounding box to the area of the union part;

[0110] Sort each of the second bounding boxes according to the class score to form a third sequence, and use the third sequence as the new second sequence.

[0111] The following explains the above steps S1051 - S1056 through a specific example. If it is necessary to detect the detection targets with the target class of vehicle from a picture, and multiple detection targets may be recognized for the same class, and there may be multiple highly overlapping bounding boxes for each detection target. For example, assume that two cars, a and b, are recognized in the image by using class prediction and bounding box prediction. There are 5 first bounding boxes for a and 5 first bounding boxes for b. Select the first bounding box with the largest class score as the target bounding box. Suppose this target bounding box is a bounding box corresponding to a. Then, keep the first bounding boxes with an overlap degree less than or equal to the preset threshold, which are the first bounding boxes belonging to b, and delete the other first bounding boxes of a because their overlap degree with the target bounding box is too large. Then, select again based on the remaining first bounding boxes and compare the overlap degrees, and all the target bounding boxes can be selected.

[0112] Since each box is independent, fast parallel bounding box decoding can be implemented based on CUDA. In non - maximum suppression, the filtering of the bounding boxes can also be completed on the GPU side, improving the calculation speed and data processing efficiency.

[0113] The object detection method provided by this application uses CUDA to implement point cloud downsampling and point cloud - to - image mapping on the GPU side. The non - neural network part uses CUDA to implement the deep fusion of point cloud features and image features on the GPU side, and the neural network part is directly completed on the GPU. The entire object detection process can maintain a very high processing speed, solving the problem that end - to - end real - time operation cannot be achieved in the deep fusion of the original image and the original point cloud, and also reducing the misdetection problem caused by insufficient semantic information contained in the fused features, improving the detection accuracy.

[0114] Corresponding to the above method embodiments, refer to Figure 3 , this application also provides a target detection device 300, which is applied to a graphics processing unit. The target detection device 300 includes:

[0115] An acquisition module 301, configured to acquire the original point cloud and the original image corresponding to the target area. Among them, the original point cloud includes a plurality of sparse points, and the original image includes a plurality of pixel points;

[0116] A projection module 302, configured to project each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel point;

[0117] An extraction module 303, configured to extract the point cloud features of the original point cloud and extract the image features of the original image. Among them, the point cloud features include a plurality of sparse point features, and the image features include a plurality of pixel point features;

[0118] A fusion module 304, configured to fuse each sparse point feature with the corresponding pixel point feature according to the corresponding relationship between each sparse point and the pixel point to obtain the target fusion feature corresponding to the target area;

[0119] A detection module 305, configured to perform class prediction and bounding box prediction on the target area based on the target fusion feature to obtain a detection target.

[0120] The target detection device provided by this application uses CUDA to implement point cloud downsampling and point cloud to image mapping on the GPU side. The non-neural network part uses CUDA to implement deep fusion of point cloud features and image features on the GPU side, while the neural network part is directly completed on the GPU. The entire target detection process can maintain a very high processing speed, solves the problem that end-to-end real-time operation cannot be achieved in the deep fusion based on the original image and the original point cloud, and also reduces the misdetection problem caused by insufficient semantic information contained in the fusion features, improving the detection accuracy.

[0121] For the specific implementation process of the provided target detection device, computer device, and computer-readable storage medium, reference may be made to the specific implementation process of the target detection method provided in the above embodiments, which will not be elaborated here one by one.

[0122] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0123] In addition, each functional module or unit in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0124] If the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0125] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A target detection method, characterized in that, Applied to a graphics processing unit, the method includes: Obtain the original point cloud and the original image corresponding to the target area, where the original point cloud includes a plurality of sparse points, and the original image includes a plurality of pixel points; Project each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel points; Extract the point cloud features of the original point cloud and the image features of the original image, where the point cloud features include a plurality of sparse point features, and the image features include a plurality of pixel point features; According to the corresponding relationship between each sparse point and the pixel points, fuse each sparse point feature with the corresponding pixel point feature to obtain the target fusion feature corresponding to the target area; Based on the target fusion feature, perform class prediction and bounding box prediction on the target area to obtain the detection target; where The point cloud data corresponding to each sparse point includes three-dimensional coordinate values and reflectivity. The step of extracting the point cloud features of the original point cloud includes: Perform parallel downsampling on the original point cloud; Based on the point cloud data corresponding to each sparse point in the downsampled original point cloud, extract the sparse point feature and the neighborhood feature corresponding to each sparse point. Where, taking any one of the sparse points as the key point, the sparse points within the preset radius range are the neighboring points corresponding to the key point, and the neighborhood feature is composed of the concatenation of the point cloud data of the neighboring points corresponding to the sparse point; Fuse each of the sparse point features and the corresponding neighborhood features into the point cloud features.

2. The method according to claim 1, wherein The step of projecting each sparse point in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel points includes: Based on the formula , establish the correspondence between sparse points and pixel points, where is the two-dimensional coordinate value of the pixel point in the image coordinate system, is the projection matrix from the camera coordinate system to the image coordinate system, with a size of 3*4, is the rotation matrix of the camera, with a size of 4*4, is the projection matrix from the radar to the camera, with a size of 4*4, is the three-dimensional coordinate value of the sparse point in the point cloud coordinate system.

3. The method according to claim 1, characterized in that, The step of determining the neighboring points of each sparse point includes: Judge whether the number N of sparse points within the preset radius range centered on the key point is greater than or equal to the preset number M, where N is a positive integer; If N is less than M, arrange the N sparse points in ascending order according to the distances between the sparse points and the key points, and determine the sparse points corresponding to the first M orders as the adjacent points of the key points; If N < M, sort the N sparse points in ascending order according to the distance between each sparse point and the key point, and determine the first M - N sparse points as supplementary points, copy the supplementary points, and determine the N sparse points and the M - N supplementary points as the neighboring points of the key point.

4. The method according to claim 1, wherein The step of fusing each sparse point feature with the corresponding pixel point feature according to the corresponding relationship between the sparse point and the pixel points to obtain the fusion feature corresponding to the target area includes: Based on the corresponding relationship between each sparse point and the pixel points, fuse each sparse point feature with the corresponding pixel point feature to obtain the first fusion feature; Interpolate the first fusion feature to obtain the second fusion feature; Extract the high-level semantic features in the second fusion feature through two Linear-BN-ReLU layers as the target fusion feature corresponding to the target area.

5. The method according to claim 4, characterized in that The step of interpolating the first fusion feature to obtain the second fusion feature includes: Select any one of the sparse points corresponding to the first fusion feature as the original point; Sort all the sparse points corresponding to the first fusion feature in ascending order according to the distance value between the sparse point and the original point to obtain the first sequence; Select the first K sparse points in the first sequence as the associated points of the original point, where K is a positive integer; Normalize the distances from each of the associated points to the original points to obtain the weights of each of the associated points; Multiply the weight corresponding to each associated point by the sparse point feature corresponding to each associated point to obtain the upsampled feature of each original point; Combine the upsampled features of the sparse points corresponding to the first fusion feature into the second fusion feature.

6. The method according to claim 1, wherein The steps of performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain the detection target include: Performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain a first bounding box with multiple different class scores corresponding to the target class; Sort the multiple first bounding boxes according to the class scores to obtain a second sequence; Repeatedly execute the step of selecting a target bounding box from the second sequence until all target bounding boxes are found; Determine the target object corresponding to each of the target bounding boxes as the detection target; Among them, the step of repeatedly executing the step of selecting a target bounding box from the second sequence until all target bounding boxes are found includes: Select the first bounding box with the largest class score in the second sequence as the target bounding box according to a preset rule; Retain the first bounding boxes with an overlap degree less than or equal to a preset threshold as the second bounding boxes, where the overlap degree is the ratio of the area of the intersection part of the first bounding box and the target bounding box to the area of the union part; Sort each of the second bounding boxes according to the class scores into a third sequence, and use the third sequence as the new second sequence.

7. A target detection device, characterized in that, Applied to a graphics processing unit, the device includes: An acquisition module for acquiring the original point cloud and the original image corresponding to the target region, where the original point cloud includes a plurality of sparse points, and the original image includes a plurality of pixel points; A projection module for projecting each of the sparse points in the original point cloud onto the original image to obtain the corresponding relationship between each sparse point and the pixel points; An extraction module for extracting the point cloud feature of the original point cloud and extracting the image feature of the original image, where the point cloud feature includes a plurality of sparse point features, and the image feature includes a plurality of pixel point features; A fusion module for fusing each sparse point feature and the corresponding pixel point feature according to the corresponding relationship between each sparse point and the pixel points to obtain the target fusion feature corresponding to the target region; A detection module for performing class prediction and bounding box prediction on the target region based on the target fusion feature to obtain the detection target, where The point cloud data corresponding to each sparse point includes three-dimensional coordinate values and reflectivity. The steps of extracting the point cloud feature of the original point cloud include; Perform parallel downsampling on the original point cloud; Based on the point cloud data corresponding to each of the sparse points in the downsampled original point cloud, extract the sparse point feature and the neighborhood feature corresponding to each sparse point. Among them, taking any one of the sparse points as the key point, the sparse points within a preset radius range are the neighboring points corresponding to the key point, and the neighborhood feature is composed of the concatenation of the point cloud data of the neighboring points corresponding to the sparse point; Fuse each of the sparse point features and the corresponding neighborhood features into the point cloud feature.

8. A computer device, characterized in that, The computer includes a graphics processor and a memory, and the memory stores a computer program. When the computer program is executed on the graphics processor, it implements the object detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program is executed on a processor, it implements the object detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal data fusion drivable area detection method based on point cloud up-sampling

    CN112731436A