Target recognition method, device and equipment
By acquiring and fusion point cloud data in the target environment, extracting and fusion point features and multi-view features, the problem of difficult to take into account the accuracy and speed of target recognition in the prior art is solved, efficient target recognition is achieved, and the safety of autonomous driving vehicles is improved.
Patent Information
- Application Number
- CN202111445551.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-30
AI Technical Summary
The prior art is difficult to take into account the high accuracy of target recognition and the fast target recognition speed, and cannot meet the requirements of the real-time and accuracy of target detection of autonomous driving vehicle systems, especially in complex road conditions, which is difficult to ensure the safety of autonomous driving.
By obtaining the current frame point cloud data and multi-frame historical point cloud data in the target environment, the point features and multi-view features of each point are extracted, and the point dimension description features are fused into point dimension description features for target recognition. The method includes extraction and fusion of multi-view features, and utilizing multi-task learning for object recognition.
It achieves the ability to ensure laser point cloud perception performance on the basis of taking into account high accuracy and fast identification speed, thereby improving the safety of autonomous vehicles under complex road conditions.
Smart Images

Figure CN113902043B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and specifically to target recognition methods, devices and equipment. Background Art
[0002] In the fields of autonomous driving and robotics, machine perception is an important component, and perception sensors include laser radar, cameras, ultrasound, millimeter-wave radar, etc. Compared with sensors such as cameras, ultrasound, and millimeter-wave radar, the laser point cloud signal of multi-line laser radar contains accurate target position information and target geometry information, so it plays an important role in the perception of autonomous driving and robotics.
[0003] At present, commonly used methods for realizing target recognition using laser point cloud include: detection methods based on traditional segmentation, deep learning methods based on laser point cloud projection, 3D laser point cloud detection methods based on voxelization, point cloud 3D detection methods based on point dimension features, etc. However, in the process of realizing the present invention, the inventors found that the above methods all have the problem of not being able to take into account both high target recognition accuracy and fast target recognition speed, and it is difficult to meet the requirements of the autonomous driving vehicle system for real-time and accuracy of target detection, and it is impossible to ensure the safety of autonomous driving under complex road conditions. Summary of the invention
[0004] The present application provides a target recognition method to solve the problem that the prior art cannot achieve both high target recognition accuracy and high target recognition speed. The present application also provides a target recognition device and equipment, and a vehicle.
[0005] The present application provides a target recognition method, comprising:
[0006] Obtain the current frame point cloud data and multiple frames of historical point cloud data in the target environment;
[0007] Obtaining point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data, and obtaining multi-view features based on the current frame point cloud data and the multiple frames of historical point cloud data;
[0008] Obtaining point dimension description features in the target environment according to the point features of each point and the multi-view features;
[0009] The target in the target environment is identified according to the point dimension description features in the target environment.
[0010] Optionally, obtaining multi-view features according to the current frame point cloud data and the multiple frames of historical point cloud data includes:
[0011] Aligning the multiple frames of historical point cloud data to the coordinate system of the current frame point cloud data by positioning;
[0012] For each frame of point cloud data in the multiple frames of historical point cloud data, extract features of a top view;
[0013] The point cloud data of the current frame is projected into the perspective of the front view to extract the features of the front view.
[0014] Optionally, extracting features of the top view from each frame of point cloud data in the multiple frames of historical point cloud data includes:
[0015] voxelize each frame of point cloud data in the multiple frames of historical point cloud data;
[0016] Extract features within non-empty voxels in each frame of point cloud data;
[0017] The features in all non-empty voxels are spliced together to obtain the features of the top view accumulated in multiple frames.
[0018] Optionally, also include:
[0019] For a same laser point in the historical point cloud data and the current frame point cloud data, obtaining features of a top view and features of a front view corresponding to the same laser point;
[0020] The features of the top view and the features of the front view corresponding to the same laser point are spliced to obtain multi-view features corresponding to the same laser point.
[0021] Optionally, obtaining the point dimension description feature in the target environment according to the point feature of each point and the multi-view feature includes:
[0022] The point feature of the same laser point is spliced with the multi-view features corresponding to the same laser point to obtain the point dimension description features corresponding to the same laser point.
[0023] Optionally, the identifying the target in the target environment according to the point dimension description feature in the target environment includes:
[0024] Multi-task learning is performed on the point dimension description features of each laser point, and the multi-task learning includes but is not limited to center point, size and direction supervision to achieve the target recognition task in the target environment.
[0025] Optionally, the multi-task learning further includes point cloud segmentation, and the method further includes:
[0026] Point cloud segmentation is performed on the point dimension description features of each laser point to achieve the point cloud segmentation task in the target environment.
[0027] Optionally, the identifying the target in the target environment according to the point dimension description feature in the target environment includes:
[0028] Predicting a target frame for the target environment according to point dimension description features in the target environment;
[0029] The method further comprises:
[0030] According to the predicted offset value of the foreground point, the foreground point is offset to be used as the foreground offset point;
[0031] Select multiple target key points from the foreground offset points;
[0032] Get the foreground offset point set corresponding to each target key point;
[0033] Determining a target prediction frame according to the foreground offset point set;
[0034] The target prediction boxes are filtered to eliminate redundant target prediction boxes.
[0035] The present application also provides a target recognition device, comprising:
[0036] A multi-frame point cloud acquisition unit is used to acquire the current frame point cloud data and multi-frame historical point cloud data in the target environment;
[0037] A point feature extraction unit, used to obtain the point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data;
[0038] A multi-view feature extraction unit, used to obtain multi-view features according to the current frame point cloud data and the multiple frames of historical point cloud data;
[0039] A feature fusion unit, used for obtaining a point dimension description feature in the target environment according to the point feature of each point and the multi-view feature;
[0040] The target recognition unit is used to recognize the target in the target environment according to the point dimension description features in the target environment.
[0041] The present application also provides an electronic device, including:
[0042] Processor; and
[0043] The memory is used to store a program for implementing the target identification method. The device is powered on and runs the program of the method through the processor.
[0044] The present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the computer-readable storage medium is run on a computer, the computer executes the above-mentioned various methods.
[0045] The present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the above-mentioned various methods.
[0046] Compared with the prior art, this application has the following advantages:
[0047] The target recognition method provided by the embodiment of the present application obtains the current frame point cloud data and multi-frame historical point cloud data in the target environment; obtains the point features of each point in the current frame point cloud data and the multi-frame historical point cloud data, and obtains multi-view features based on the current frame point cloud data and the multi-frame historical point cloud data; obtains the point dimension description features in the target environment based on the point features of each point and the multi-view features; and identifies the target in the target environment based on the point dimension description features in the target environment. This processing method allows multi-view point dimension description features to be obtained based on multi-frame point clouds, thereby ensuring the laser point cloud perception performance, and therefore can effectively take into account both high target recognition accuracy and fast recognition speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flow chart of an embodiment of a target recognition method provided by the present application;
[0049] Figure 2 is a schematic diagram of a network structure of an embodiment of a target recognition method provided by the present application;
[0050] Figure 3 It is a schematic diagram of target frame prediction of an embodiment of the target recognition method provided in the present application. DETAILED DESCRIPTION
[0051] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application, so the present application is not limited by the specific implementation disclosed below.
[0052] In the present application, a target recognition method, device, equipment, and vehicle are provided. In the following embodiments, various schemes are described in detail one by one.
[0053] First embodiment
[0054] Please refer to Figure 1, which is a flow chart of an embodiment of a target recognition method provided by the present application. The execution subject of the method includes but is not limited to unmanned vehicles, such as intelligent logistics vehicles, etc., and the identifiable targets include pedestrians, vehicles, buildings, trees, curbs, traffic lights, zebra crossings, etc. A target recognition method provided by the present application includes:
[0055] Step S101: Acquire the current frame point cloud data and multiple frames of historical point cloud data in the target environment.
[0056] The method provided in the embodiment of the present application can obtain the spatial coordinates of each sampling point on the surface of the environmental space object on the road where the vehicle is traveling through a three-dimensional space scanning device installed on the vehicle during the vehicle's driving process, and obtain a set of points. This massive point data is called road environment point cloud (Point Cloud) data. Through the road environment point cloud data, the scanned object surface is recorded in the form of points, and each point contains three-dimensional coordinates, some of which may contain color information (RGB) or reflection intensity information (Intensity), etc. With point cloud data, the target space can be expressed in the same spatial reference system.
[0057] The three-dimensional space scanning device may be a laser radar (Light Detection And Ranging, Lidar), which performs laser detection and measurement by laser scanning to obtain information about obstacles in the surrounding environment, such as buildings, trees, people, vehicles, etc. The measured data is a discrete point representation of a digital surface model (DSM). In specific implementation, a multi-line laser radar such as 16-line, 32-line, or 64-line can be used. The frame rate (Frame Rate) of point cloud data collected by radars with different numbers of laser beams is different. For example, 16 and 32 lines generally collect 10 frames of point cloud data per second. The three-dimensional space scanning device may also be a three-dimensional laser scanner or a photographic scanner.
[0058] The method provided in the embodiment of the present application can perform target recognition on the current road environment based on the point cloud data of the current frame and the point cloud data of multiple historical frames before the current frame (multi-frame historical point cloud data) after collecting point cloud data (current frame point cloud data) at a certain moment. The multiple historical frames can be multiple adjacent historical frames before the current frame, or one or more non-adjacent historical frames before the current frame.
[0059] Step S103: obtaining point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data, and obtaining multi-view features based on the current frame point cloud data and the multiple frames of historical point cloud data.
[0060] By combining the current frame point cloud data with the multi-frame historical point cloud data, more abundant feature information of each target in the target environment can be obtained. In this embodiment, two types of features can be obtained: point features of each point in the current frame point cloud data and the multi-frame historical point cloud data, and multi-view features obtained based on the current frame point cloud data and the multi-frame historical point cloud data.
[0061] The current frame point cloud data and the multi-frame historical point cloud data may include point cloud data of the same point and point cloud data of different points on the same target surface collected at different positions, and these point cloud data can more accurately represent the target space. Accordingly, more abundant point features of the same target can be extracted from these point cloud data, that is, the point features of each point in the current frame point cloud data and the multi-frame historical point cloud data.
[0062] In one example, the point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data can be obtained by a multi-layer perceptron. In specific implementation, the current frame point cloud data and the multiple frames of historical point cloud data can be used as input data of the multi-layer perceptron, and the output data of the multi-layer perceptron is the point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data. The multi-layer perceptron can adopt a more mature perceptron in the prior art, such as a perceptron based on a neural network.
[0063] The multi-view features may include features of the target environment observed from different perspectives, such as front view features, top view features, rear view features, top view features, left view features, and right view features, and may also be specified observation angles, such as a 45-degree top view, etc. The multi-view features may be obtained based on the current frame point cloud data and the multiple frames of historical point cloud data.
[0064] In one example, the multi-view feature can be obtained by the following steps:
[0065] Step S201: aligning the multiple frames of historical point cloud data to the coordinate system of the current frame of point cloud data through positioning.
[0066] The multi-frame historical point cloud data and the current frame point cloud data are point cloud data of the target space collected at different positions. The point cloud data of different historical frames may include point cloud data of the same target. The point cloud data of different frames of the same target are located in different coordinate systems. The multi-frame historical point cloud data must be converted to the coordinate system of the current frame point cloud data so that the point cloud data of different frames of the same target can correspond to each other. During specific implementation, the position information corresponding to each historical frame can be obtained, and based on the position information, the multi-frame historical point cloud data is aligned to the coordinate system of the current frame point cloud data. Aligning the multi-frame historical point cloud data to the coordinate system of the current frame point cloud data through positioning belongs to the scope of the prior art and will not be repeated here.
[0067] Step S203: extracting features of the top view for each frame of point cloud data in the multiple frames of historical point cloud data.
[0068] After aligning the multi-frame historical point cloud data to the coordinate system of the current frame point cloud data, the features of the top view can be extracted based on the aligned multi-frame historical point cloud data, so that rich top view features can be extracted.
[0069] In this embodiment, the features of the top view can be extracted for each frame of point cloud data in the multiple frames of historical point cloud data after the coordinate system conversion. In specific implementation, step S203 may include the following sub-steps: 1) voxelizing each frame of point cloud data in the multiple frames of historical point cloud data; 2) extracting the features in the non-empty voxels in each frame of point cloud data; 3) splicing the features in all non-empty voxels to obtain the features of the top view accumulated in multiple frames. In this processing method, the point cloud can be voxelized first and the features in each non-empty voxel can be extracted. Then these features are concatenated to form multi-frame accumulated top view features.
[0070] In another example, step S203 may include the following sub-steps: 1) superimposing the current frame point cloud data and the aligned multi-frame historical point cloud data; 2) voxelizing the superimposed multi-frame point cloud data; 3) projecting the point cloud data of non-empty voxels to the top view plane; 4) determining the characteristics of the top view based on the projected data.
[0071] Step S205: Project the current frame point cloud data into the perspective of the front view to extract features of the front view.
[0072] The current frame point cloud data usually includes a relatively comprehensive point cloud data of the side of the target (such as buildings around the road, vehicles on the road, etc.) facing the execution subject (such as the current unmanned vehicle). Therefore, based on the current frame point cloud data, relatively rich front view features of the target space can be extracted. In specific implementation, step S205 may include the following sub-steps: 1) projecting the current frame point cloud data to the front view plane; 2) determining the front view features based on the projected data.
[0073] In another example, the process of obtaining the multi-view features may further include at least one of the following steps: extracting the features of the right view for each frame of point cloud data in the multiple frames of historical point cloud data; extracting the features of the left view for each frame of point cloud data in the multiple frames of historical point cloud data; extracting the features of the upward view for each frame of point cloud data in the multiple frames of historical point cloud data; and projecting the current frame of point cloud data into the perspective of the rear view to extract the features of the rear perspective. In this way, features of more perspectives can be extracted, thereby obtaining richer feature information of the target space.
[0074] In specific implementation, the multi-view features can also be obtained by the following steps: aligning the multi-frame historical point cloud data to the coordinate system of the current frame point cloud data through positioning; extracting the features of the front view for each frame of the multi-frame historical point cloud data; projecting the current frame point cloud data to the perspective of the top view to extract the features of the top view. In this way, more abundant front view features can be extracted based on the multi-frame point cloud data.
[0075] In addition, the multi-view features can also be obtained by the following steps: aligning the multi-frame historical point cloud data to the coordinate system of the current frame point cloud data by positioning; extracting the front view features and the top view features for each frame of the multi-frame historical point cloud data and the current frame point cloud data. In this way, more abundant front view features and top view features can be extracted from the multi-frame point cloud data.
[0076] In one example, after extracting the features of the top view and the front view of the target environment, the method may further include the following steps:
[0077] Step S401: for the same laser point in the historical point cloud data and the current frame point cloud data, obtain features of the top view and features of the front view corresponding to the same laser point.
[0078] In this embodiment, the multi-view features obtained in step S103 include the overall top view features and front view features in the target environment, and the top view features and front view features corresponding to the same laser point in the historical point cloud data and the current frame point cloud data can be obtained according to the position information in the overall environment. That is, the overall top view features and front view features in the target environment are converted into top view features and front view features of each point in the target environment.
[0079] Step S403: splicing the features of the top view and the features of the front view corresponding to the same laser point to obtain multi-view features corresponding to the same laser point.
[0080] After obtaining the top view features and front view features of each point in the target environment, the top view features and front view features of each point can be spliced to obtain the multi-view features of each point.
[0081] Step S105: Obtain point dimension description features in the target environment according to the point features of each point and the multi-view features.
[0082] The target environment includes multiple laser points. In this embodiment, the point dimension description feature of each point includes not only the point feature of the point obtained based on the original point cloud data, but also the multi-view feature of each point. It can be seen that the point dimension feature data can more accurately retain the measurement information of the original target, and can represent richer feature information of each laser point, thereby improving the sensor perception performance.
[0083] In this embodiment, step S105 can be implemented in the following manner: splicing the point feature of the same laser point with the multi-view feature corresponding to the same laser point to obtain the point dimension description feature corresponding to the same laser point. In specific implementation, the point feature and the multi-view feature of a laser point can be spliced by connecting the dimensions of the point feature with the dimensions of the multi-view feature. For example, if the point feature is 10-dimensional data and the multi-view feature is 15-dimensional data, the connected feature is 25-dimensional data. In addition, the dimensions of the point feature and the dimensions of the multi-view feature can be weighted and summed. For example, if the point feature is 10-dimensional data and the point feature weight is 0.6, and the multi-view feature is 10-dimensional data and the multi-view weight is 0.4, the spliced feature is still 10-dimensional data.
[0084] Step S107: Identify the target in the target environment according to the point dimension description features in the target environment.
[0085] The point dimension description feature has richer feature information, and target recognition based on the point dimension description feature can obtain more accurate recognition results. The recognition of targets in the target environment can be recognition of target categories (such as buildings, trees, people, vehicles, etc.), recognition of target frames (such as the bounding box of the target), recognition of target point clouds (segmenting the point cloud of the overall environment into point clouds corresponding to each target), etc.
[0086] In this embodiment, step S107 can be implemented in the following manner: multi-task learning is performed on the point dimension description features of each laser point, and the multi-task learning includes but is not limited to center point, size and direction supervision to achieve the target recognition task in the target environment. The center point can be the center point of a dynamic target. The size can be the size of a dynamic target. The direction can be the orientation of a dynamic target, such as the driving direction of a vehicle. Using this multi-task learning method can effectively improve the efficiency of target recognition.
[0087] The multi-task learning may also include target point cloud segmentation. Accordingly, the method may also include the following steps: performing point cloud segmentation on the point dimension description features of each laser point to achieve the point cloud segmentation task in the target environment, such as performing point cloud segmentation on static targets (such as trees and buildings) in the target environment. In this way, point cloud data of each target in the target environment can be obtained.
[0088] The multi-task learning may also include target category recognition, and accordingly, the method may also include the following steps: identifying the target category in the target environment according to the point dimension description features in the target environment. In this way, the categories of each target in the target environment, such as buildings, trees, people, vehicles, etc., can be obtained.
[0089] In this embodiment, the specific process of target recognition is: 1) using a laser radar to obtain laser point cloud data of the current frame and multiple frames of history in the target environment; 2) based on the multi-frame original point cloud data, obtaining the point features of each laser point in the target environment, and, based on the multi-frame original point cloud data, obtaining the multi-perspective features of each laser point, such as the features of the top view and the features of the front view; 3) using a convolutional neural network to extract features from the original multi-perspective features; 4) fusing the point features of each point with the multi-perspective features processed by the convolutional neural network to obtain point dimension description features; 5) performing target recognition processing based on the point dimension description features.
[0090] In one example, target recognition processing is performed by a target recognition model based on a neural network. The target recognition model includes: a point feature extraction network, a multi-view feature processing network, a feature fusion network, and a multi-task learning network. The point feature extraction network is used to extract the point features of each laser point in the target environment based on multi-frame point cloud data. The point feature extraction network can adopt a multi-layer perceptron structure or other network structures. The multi-view feature processing network is used to perform feature transformation on the multi-view features. The multi-view feature processing network can adopt a convolutional neural network. The feature fusion network is used to obtain point dimension description features based on the point features and the transformed multi-view features. The multi-task learning network, also known as a multi-task decision network, is used to perform multi-task target recognition based on the point dimension description features. By adopting this processing method, an end-to-end cloud panoramic segmentation network based on multi-frame point clouds can be realized, which can effectively take into account high target recognition accuracy and fast target recognition speed, and is an important part of ensuring the performance of laser point cloud perception.
[0091] like Figure 2 As shown, the point feature extraction network is used to extract the point features of each laser point in the target environment according to the current frame point cloud data and the multi-frame historical point cloud data. The point feature extraction network can use a multi-layer perceptron, which can be a multi-layer fully connected network. In this embodiment, there is no receptive field between the point features of different laser points, so the points are isolated from each other. The multi-view feature processing network includes a top view feature processing network and a front view feature processing network. The top view feature processing network is used to perform feature transformation on the top view features accumulated in multiple frames. The top view feature processing network can use a U-type network, including multiple convolutional layers and multiple deconvolutional layers, so that an output feature map with the same resolution as the input feature map can be obtained, and then the transformed top view features of each laser point, that is, the point-level features of the top view, can be obtained. Similarly, the front view feature processing network can also use a U-type network to obtain the transformed front view features of each laser point, that is, the point-level features of the front view. In this embodiment, the point-level transformed multi-view features have rich structured semantic information and stronger representation capabilities. The input data of the feature fusion network include: point features of each laser point, point-level transformed top view features and point-level transformed front view features. The feature fusion method can be the splicing of different features of the same laser point (point features, transformed top view features, transformed front view features), or the weighted sum of different features. The attention mechanism can also be used for feature fusion processing. The input data of the feature fusion network include: point dimension description features of each laser point. The point dimension description features can have accurate target location information and rich target semantic information. The multi-task learning network can include a point cloud segmentation network, a target category recognition network, and a target bounding box regression network.
[0092] The target bounding box regression network can realize the three-dimensional target box prediction for each laser point. Since there will be multiple laser points (foreground points) for each foreground target, there may be a large number of redundant target prediction boxes in the target prediction box output by the target bounding box regression network. In order to eliminate a large number of redundant target prediction boxes, the target bounding box regression network can adopt the non-maximum suppression method (NMS) and use the result with the highest prediction score (confidence of the foreground target) as the final output of the target prediction box. However, the NMS method has two defects. One is that the amount of calculation is large, which is not conducive to running on low-power devices. The other problem is that a single detection box cannot achieve high-precision target box prediction due to the presence of noise.
[0093] In one example, the target bounding box regression network may predict the target box using the following steps:
[0094] Step S501: performing an offset process on the foreground point according to the offset prediction value of the foreground point to obtain a foreground offset point.
[0095] In the target environment, some targets are foreground targets, such as cars and pedestrians, and some targets are background targets, such as buildings, trees, curbs, traffic lights, etc. The laser points of foreground targets are called foreground points, and the laser points of background targets are called background points.
[0096] In this embodiment, the target category recognition network can be used to predict the target category corresponding to each laser point, and the foreground point can be determined according to the target category, such as Figure 3 The foreground point shown in (a) in FIG. 4 is a foreground point. The offset of each laser point relative to the target center point can be predicted through the target frame regression network. In this step, the foreground point is offset to the corresponding target center point according to the offset prediction value of each foreground point. The offset foreground point is referred to as the foreground offset point. Figure 3 (b) shows the foreground offset points clustered together.
[0097] Step S503: Select multiple target key points from the foreground offset points.
[0098] The multiple target key points may include points of different foreground targets in the target environment. In this embodiment, each foreground point after the offset is used as a vote, such as Figure 3 As shown in (b) in the figure, among the ballots offset to the corresponding target center point, multiple laser points can be selected from them by using the farthest point sampling method as the main ballot (i.e., the target key point), as shown in Figure 3 The key points shown in (c).
[0099] Step S505: Obtain a foreground offset point set corresponding to each target key point.
[0100] For each target key point, the foreground offset points around the point can be obtained by using the "ball query" method. For example, the ball radius can be set, the range of the ball can be determined according to the ball radius, and the original foreground points corresponding to the foreground offset points within the ball range can be used to form a target prediction frame.
[0101] Step S507: Determine a target prediction box according to the foreground offset point set.
[0102] In specific implementation, the size of the target box can be predicted by an estimator, such as averaging the center point, length, width, height, direction and score (target category confidence of each foreground point) in a single original laser point set corresponding to the main ballot through a mean estimator to obtain a target prediction box, such as Figure 3 For example, through the target category recognition network, the target category confidence of each laser point in the target environment can be predicted, such as the confidence that a certain laser point is a car is 90%, and the confidence that a certain laser point is a pedestrian is 86%.
[0103] Step S509: Filter the target prediction box.
[0104] Since the number of primary votes K may be greater than the actual number of targets, it is also necessary to filter the target prediction frames directly determined based on the primary votes to remove redundant target prediction frames. For example, if there are 5 foreground targets in the target environment and the number of primary votes k is set to 256, multiple target key points (256 key points) will be sampled for each of the 5 foreground targets, thus forming 256 target prediction frames. Obviously, there are a large number of redundant target frames that need to be filtered out.
[0105] In specific implementation, the non-maximum suppression method (NMS) can be used to use the result with the highest prediction score (confidence of the foreground target) as the final output of the target prediction box. In addition, clustering methods (such as dbscan, kmean algorithm) can also be used to aggregate the target boxes corresponding to multiple target key points sampled from each target. In this way, the box estimation is achieved at the point granularity using each foreground point.
[0106] In this embodiment, through the above steps S501-509, a more accurate target box prediction is achieved by using a voting method.
[0107] It can be seen from the above embodiments that the target recognition method provided by the embodiments of the present application obtains the current frame point cloud data and multi-frame historical point cloud data in the target environment; obtains the point features of each point in the current frame point cloud data and the multi-frame historical point cloud data, and obtains multi-view features based on the current frame point cloud data and the multi-frame historical point cloud data; obtains the point dimension description features in the target environment based on the point features of each point and the multi-view features; and identifies the target in the target environment based on the point dimension description features in the target environment. This processing method allows multi-view point dimension description features to be obtained based on multi-frame point clouds, thereby ensuring the laser point cloud perception performance, and therefore can effectively take into account both high target recognition accuracy and fast recognition speed.
[0108] Second embodiment
[0109] In the above-mentioned embodiment, a target recognition method is provided, and correspondingly, the present application also provides a target recognition device. The device corresponds to the embodiment of the above-mentioned method. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0110] A target recognition device of this embodiment includes: a multi-frame point cloud acquisition unit, a point feature extraction unit, a multi-view feature extraction unit, and a feature fusion unit.
[0111] A multi-frame point cloud acquisition unit is used to acquire the current frame point cloud data and multi-frame historical point cloud data in the target environment; a point feature extraction unit is used to obtain the point features of each point in the current frame point cloud data and the multi-frame historical point cloud data; a multi-view feature extraction unit is used to obtain multi-view features based on the current frame point cloud data and the multi-frame historical point cloud data; a feature fusion unit is used to obtain the point dimension description features in the target environment based on the point features of each point and the multi-view features; a target recognition unit is used to recognize the target in the target environment based on the point dimension description features in the target environment.
[0112] Optionally, the multi-view feature extraction unit includes:
[0113] A coordinate conversion subunit, used for aligning the multiple frames of historical point cloud data to the coordinate system of the current frame point cloud data through positioning;
[0114] A top view feature extraction subunit, used to extract the features of the top view for each frame of point cloud data in the multiple frames of historical point cloud data;
[0115] The front view feature extraction subunit is used to project the current frame point cloud data into the perspective of the front view to extract the features of the front view.
[0116] Optionally, the top view feature extraction subunit includes:
[0117] A voxelization subunit, used for voxelizing each frame of point cloud data in the multiple frames of historical point cloud data;
[0118] A voxel feature extraction subunit is used to extract features within non-empty voxels in each frame of point cloud data;
[0119] The voxel feature stitching subunit is used to stitch the features in all non-empty voxels to obtain the features of the top view accumulated in multiple frames.
[0120] Optionally, the device further comprises:
[0121] A point-granular multi-view feature acquisition unit, for obtaining, for a same laser point in the historical point cloud data and the current frame point cloud data, features of a top view and features of a front view corresponding to the same laser point;
[0122] The point-granular multi-view feature stitching unit is used to stitch the features of the top view and the features of the front view corresponding to the same laser point to obtain the multi-view features corresponding to the same laser point.
[0123] Optionally, the feature fusion unit is specifically used to splice the point feature of the same laser point with the multi-view features corresponding to the same laser point to obtain the point dimension description features corresponding to the same laser point.
[0124] Optionally, the target recognition unit is specifically used to perform multi-task learning on the point dimension description features of each laser point, and the multi-task learning includes but is not limited to center point, size and direction supervision to achieve the target recognition task in the target environment.
[0125] Optionally, the multi-task learning also includes point cloud segmentation, which performs point cloud segmentation on the point dimension description features of each laser point to achieve the point cloud segmentation task in the target environment.
[0126] Optionally, the multi-task learning also includes: predicting a target frame for the target environment based on point dimension description features in the target environment; a target frame prediction unit, used to offset the foreground point according to the offset prediction value of the foreground point, as a foreground offset point; selecting multiple target key points from the foreground offset points; obtaining a foreground offset point set corresponding to each target key point; determining a target prediction frame based on the foreground offset point set; and screening the target prediction frames to eliminate redundant target prediction frames.
[0127] Third embodiment
[0128] In the above-mentioned embodiment, a target recognition method is provided, and correspondingly, the present application also provides an electronic device. The device corresponds to the above-mentioned method embodiment. Since the device embodiment is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described below is only illustrative.
[0129] The present application provides an electronic device, comprising: a processor and a memory; wherein the memory is used to store a program for implementing the above-mentioned target recognition method, and the device is powered on and runs the program of the method through the processor.
[0130] The electronic device may be a server or an unmanned vehicle.
[0131] In one example, the electronic device is a server, which can store multi-frame historical point cloud data of the unmanned vehicle in a target environment, receive current frame point cloud data of the target environment uploaded by the unmanned vehicle, obtain the current frame point cloud data and the point features of each point in the multi-frame historical point cloud data, and obtain multi-perspective features based on the current frame point cloud data and the multi-frame historical point cloud data; obtain point dimension description features in the target environment based on the point features of each point and the multi-perspective features; and identify the target in the target environment based on the point dimension description features in the target environment.
[0132] In another example, the electronic device is an unmanned vehicle, which collects current frame point cloud data in a target environment, obtains multi-frame historical point cloud data of the unmanned vehicle in the target environment, obtains point features of each point in the current frame point cloud data and the multi-frame historical point cloud data, and obtains multi-perspective features based on the current frame point cloud data and the multi-frame historical point cloud data; obtains point dimension description features in the target environment based on the point features of each point and the multi-perspective features; and identifies targets in the target environment based on the point dimension description features in the target environment.
[0133] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0134] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0135] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0136] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0137] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
Claims
1. A target recognition method, characterized in that: include: Obtain the current frame point cloud data and multiple frames of historical point cloud data in the target environment; Obtaining point features of each point in the current frame point cloud data and the multi-frame historical point cloud data, where there is no receptive field between point features of different laser points and the points are isolated from each other; and obtaining multi-view point-level features based on the current frame point cloud data and the multi-frame historical point cloud data; According to the point feature of each point and the point-level features of the multiple perspectives, a point dimension description feature of each laser point in the target environment is obtained, where the point dimension description feature has target position information and target semantic information; According to the point dimension description features in the target environment, the targets in the target environment are identified to obtain point cloud data of each target in the target environment.
2. The method according to claim 1, characterized in that: The step of obtaining multi-view point-level features according to the current frame point cloud data and the multiple frames of historical point cloud data includes: Aligning the multiple frames of historical point cloud data to the coordinate system of the current frame point cloud data by positioning; For each frame of point cloud data in the multiple frames of historical point cloud data, extracting point-level features of a bird's-eye view; The point cloud data of the current frame is projected into the perspective of the front view to extract point-level features of the front view.
3. The method according to claim 2, characterized in that The step of extracting point-level features of the top view from each frame of point cloud data in the multiple frames of historical point cloud data includes: voxelize each frame of point cloud data in the multiple frames of historical point cloud data; Extract features within non-empty voxels in each frame of point cloud data; The features in all non-empty voxels are spliced to obtain the point-level features of the top view accumulated in multiple frames.
4. The method according to claim 2, characterized in that: Also includes: For a same laser point in the historical point cloud data and the current frame point cloud data, obtaining features of a top view and features of a front view corresponding to the same laser point; The features of the top view and the features of the front view corresponding to the same laser point are spliced to obtain multi-view features corresponding to the same laser point.
5. The method according to claim 4, characterized in that The step of obtaining point dimension description features of each laser point in the target environment according to the point features of each point and the point-level features of the multiple perspectives includes: The point feature of the same laser point is spliced with the multi-view features corresponding to the same laser point to obtain the point dimension description features corresponding to the same laser point.
6. The method according to claim 5, characterized in that The identifying of the target in the target environment according to the point dimension description features of each laser point in the target environment includes: Multi-task learning is performed on the point dimension description features of each laser point, and the multi-task learning includes center point, size and direction supervision to achieve the target recognition task in the target environment.
7. The method according to claim 6, characterized in that The multi-task learning also includes point cloud segmentation, and the method further includes: Point cloud segmentation is performed on the point dimension description features of each laser point to achieve the point cloud segmentation task in the target environment.
8. The method according to any one of claims 1 to 7, characterized in that The identifying of the target in the target environment according to the point dimension description features of each laser point in the target environment includes: According to the point dimension description features of each laser point in the target environment, a target frame prediction is performed on the target environment; The method further comprises: According to the predicted offset value of the foreground point, the foreground point is offset to be used as the foreground offset point; Select multiple target key points from the foreground offset points; Get the foreground offset point set corresponding to each target key point; Determining a target prediction frame according to the foreground offset point set; The target prediction boxes are filtered to eliminate redundant target prediction boxes.
9. A target recognition device, characterized in that: include: A multi-frame point cloud acquisition unit is used to acquire the current frame point cloud data and multi-frame historical point cloud data in the target environment; A point feature extraction unit, used to obtain the point features of each point in the current frame point cloud data and the multiple frames of historical point cloud data, where there is no receptive field between the point features of different laser points, and the points are isolated from each other; A multi-view feature extraction unit, used to obtain multi-view point-level features according to the current frame point cloud data and the multiple frames of historical point cloud data; A feature fusion unit, used to obtain point dimension description features of each laser point in the target environment according to the point feature of each point and the point-level features of the multiple perspectives, wherein the point dimension description features have target position information and target semantic information; The target recognition unit is used to recognize the targets in the target environment according to the point dimension description features in the target environment, so as to obtain point cloud data of each target in the target environment.
10. An electronic device, characterized in that: include: processor; as well as A memory for storing a program for implementing the target recognition method according to any one of claims 1 to 8, wherein the device is powered on and runs the program of the method through the processor.
Citation Information
Patent Citations
Vehicle detection method, apparatus, computer device, and storage medium
CN109271880A
Point cloud and multi-view fused vehicle-mounted laser point cloud multi-target identification method
CN112257637A
Object recognition model training method and device and object recognition method and system
CN112287860A