Target object detection method, device, electronic device and storage medium
By acquiring point cloud data in the surrounding environment of the target vehicle and determining and fusion processing of feature types, the problem of poor detection accuracy of target objects in the prior art is solved, and the safety of vehicle autonomous driving is improved.
Patent Information
- Application Number
- CN202310897182.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-07-20
AI Technical Summary
In the prior art, the lidar detection algorithm has poor accuracy in detecting target objects, and cannot detect speed limit signs, prohibited right turn signs, etc. in a timely manner, resulting in poor safety of vehicle autonomous driving.
By obtaining the target point cloud data, service information and computing power values of the surrounding environment of the target vehicle body, determining the feature type based on the computing power value, and extracting the features of the target point cloud data under each feature type, performing fusion processing, obtaining the fusion features, and finally processing according to the network model corresponding to the service type to detect the target object.
It improves the accuracy of target objects detection, enhances the safety of vehicle autonomous driving, and can promptly detect target objects such as speed limit signs and prohibited right turn signs.
Smart Images

Figure CN116844128B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a method, device, electronic device and storage medium for detecting a target object. Background Art
[0002] LiDAR is widely used in high-level autonomous driving environment perception due to its high ranging accuracy and high stability. In recent years, with the decline in the cost of LiDAR sensors and the launch of automotive-grade products, it has also been gradually applied to consumer-grade mass-produced models.
[0003] At present, the LiDAR detection algorithms mainly include network algorithms based on single point, voxel, bird's-eye view (BEV) and range-view (RV) features. Each algorithm has its own defects. Among them, the algorithms based on Point and Voxel features often require high computing resources due to the use of 3D convolutional networks. To solve this problem, it is a common method to project the three-dimensional point cloud data into two-dimensional perspectives such as BEV or RV and use 2D convolutional networks. Although BEV features are more in line with the movement scenes of vehicles on the road plane, their sparsity limits the detection of distant targets. Although RV features are dense and can quickly query neighborhood relationships, they face the problem of different scales.
[0004] Therefore, the existing technology has poor detection accuracy of target objects and is unable to detect speed limit signs, no right turn signs, etc. in a timely manner, causing the vehicle to continue speeding or turning right, making it easy for the target vehicle to collide with other vehicles, thereby leading to poor safety of vehicle autonomous driving. Summary of the invention
[0005] The present application provides a target object detection method, device, electronic device and storage medium to solve the problem of poor vehicle autonomous driving safety caused by poor target object detection accuracy.
[0006] According to a first aspect of the present application, a method for detecting a target object is provided, comprising:
[0007] Obtaining target point cloud data, business information and computing power value of the target vehicle's surrounding environment; wherein the business information includes business type;
[0008] Determine the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, extract the features of the target point cloud data under each feature type respectively; wherein different feature types correspond to different viewing angles;
[0009] Performing fusion processing on the features of the target point cloud data under all feature types to obtain fusion features;
[0010] The fused features are processed according to a network model corresponding to the business type to detect a target object in the surrounding environment of the vehicle body and obtain a detection result.
[0011] Optionally, determining a corresponding feature type according to the computing power value of the target vehicle includes any one of the following:
[0012] When the computing power value of the target vehicle is not less than a first preset threshold, determining the corresponding feature type as a single point feature type, a bird's-eye view BEV feature type, and a depth map RV feature type;
[0013] When the computing power value of the target vehicle is greater than a second preset threshold value and less than a first preset threshold value, determining that the corresponding feature type is at least two of a single point feature type, a BEV feature type, and an RV feature type; wherein the first preset threshold value is higher than the second preset threshold value;
[0014] When the computing power value of the target vehicle is not greater than the second preset threshold, the corresponding feature type is determined to be at least one of a single-point feature type, a BEV feature type, and a RV feature type.
[0015] Optionally, the extracting the features of the target point cloud data under each feature type respectively includes at least one of the following:
[0016] Using a preset single-point feature extraction network to extract point features from the target point cloud data to obtain single-point features;
[0017] Performing a BEV perspective projection process on the target point cloud data to obtain a first BEV feature of the target point cloud data under the BEV perspective, and inputting the first BEV feature into a preset BEV feature extraction network to obtain a second BEV feature of the target point cloud data under the BEV feature type;
[0018] The target point cloud data is projected from the RV perspective to obtain a first RV feature of the target point cloud data under the RV perspective, and the first RV feature is input into a preset RV feature extraction network to obtain a second RV feature of the target point cloud data under the RV feature type.
[0019] Optionally, the projecting the target point cloud data from a BEV perspective includes:
[0020] The target point cloud data is projected at a BEV viewing angle using a first target projection method; wherein the first target projection method includes at least one of the following: a grid projection method and a Pillar projection method.
[0021] Optionally, the projecting the target point cloud data from an RV perspective includes:
[0022] The target point cloud data is projected under the RV viewing angle using a second target projection method; wherein the second target projection method includes at least one of the following: a cylindrical view projection method, a spherical view projection method, and an annular azimuthal projection method.
[0023] Optionally, the fusing of features of the target point cloud data under all feature types to obtain fused features includes:
[0024] Performing back-projection processing on the second BEV feature and / or the second RV feature to obtain a BEV point cloud feature and / or a RV point cloud feature;
[0025] The BEV point cloud features, the RV point cloud features and / or the single point features are spliced to obtain fused features.
[0026] Optionally, the step of acquiring target point cloud data of the surrounding environment of the target vehicle body includes:
[0027] Acquire original point cloud data of the environment around the target vehicle body; wherein the original point cloud data is obtained by scanning the environment around the target vehicle body by a laser radar device disposed on the target vehicle body;
[0028] The original point cloud data is preprocessed to obtain the target point cloud data; wherein the target point cloud data includes coordinate data and reflection intensity values of each target point; wherein the target point is a reflection point of multiple laser beams emitted by a laser radar device contained in the original point cloud data, and each of the reflection points has a respective reflection intensity value.
[0029] According to a second aspect of the present application, a device for detecting a target object is provided, comprising:
[0030] An acquisition module, used to acquire target point cloud data of the surrounding environment of the target vehicle body, business information and computing power value of the target vehicle; wherein the business information includes business type;
[0031] A determination extraction module is used to determine the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, respectively extract the features of the target point cloud data under each feature type; wherein different feature types correspond to different viewing angles;
[0032] A fusion processing module is used to fuse the features of the target point cloud data under all feature types to obtain fusion features;
[0033] The detection module is used to process the fusion feature according to the network model corresponding to the business type to detect the target object in the environment around the vehicle body and obtain a detection result.
[0034] According to a third aspect of the present application, there is provided an electronic device, comprising: at least one processor and a memory;
[0035] The memory stores computer-executable instructions;
[0036] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the target object detection method as described in the first aspect above.
[0037] According to a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the target object detection method as described in the first aspect above.
[0038] According to a fifth aspect of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the target object detection method described in the first aspect.
[0039] The present application provides a method for detecting a target object, including: obtaining target point cloud data, business information and computing power value of the target vehicle's body environment; wherein the business information includes the business type; determining the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, extracting the features of the target point cloud data under each feature type respectively; wherein different feature types correspond to different viewing angles; fusing the features of the target point cloud data under all feature types to obtain fused features; processing the fused features according to the network model corresponding to the business type to detect the target object in the body environment to obtain a detection result.
[0040] While taking into account the computing power value of the target vehicle, the present application provides feature types corresponding to different perspectives, and then extracts the features of the target point cloud data under each feature type respectively. Since there are multiple feature types, the fused features obtained after the features of the target point cloud data under each feature type are fused and processed can have the advantages of the features of the target point cloud data under each feature type, thereby improving the accuracy of target object detection and thus improving the safety of vehicle autonomous driving.
[0041] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 A schematic diagram of a method for detecting a target object provided in an embodiment of the present application;
[0043] Figure 2 A schematic diagram of a BEV view provided in an embodiment of the present application;
[0044] Figure 3 A schematic diagram of an RV view provided in an embodiment of the present application;
[0045] Figure 4 A schematic diagram of a process for detecting another target object provided in an embodiment of the present application;
[0046] Figure 5 A schematic diagram of a multi-feature fusion processing flow provided in an embodiment of the present application;
[0047] Figure 6 A schematic diagram of the structure of a target object detection device provided in an embodiment of the present application;
[0048] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] Here, exemplary embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application.
[0050] The existing technology projects 3D point cloud data to 2D perspectives such as BEV or RV and uses 2D convolutional networks. Although BEV features are more consistent with the movement of vehicles on the road plane, their sparsity limits the detection of distant targets; and although RV features are dense and can quickly query neighborhood relationships, they face the problem of different scales.
[0051] However, the existing technology has problems such as poor target object detection accuracy and poor vehicle autonomous driving safety.
[0052] In order to solve the above technical problems, the present application provides a target object detection method applied in the field of autonomous driving, which is used to improve the detection accuracy of target objects and thereby improve the safety of vehicle autonomous driving.
[0053] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0054] Figure 1 A schematic diagram of a method for detecting a target object provided in an embodiment of the present application. Figure 1 As shown, the method of this embodiment includes the following steps:
[0055] S10. Obtain target point cloud data of the surrounding environment of the target vehicle, business information, and computing power value of the target vehicle; wherein the business information includes business type.
[0056] In the embodiment of the present application, the vehicle body environment is the vehicle body environment of the target vehicle. The business information may be data segmentation and / or target object detection information. The computing power value may be obtained from a processing device on the target vehicle.
[0057] S20. Determine a corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, extract features of the target point cloud data under each feature type respectively; wherein different feature types correspond to different viewing angles.
[0058] It should be understood that the feature type includes but is not limited to: single point feature type, BEV feature type, RV feature type, etc. The specific description of step S20 is as follows and will not be repeated here.
[0059] S30, fusing the features of the target point cloud data under all feature types to obtain fused features.
[0060] In the embodiment of the present application, the above-mentioned fusion feature is also called a multi-view network feature, and the image formed by the fusion feature is a Multi-View-Feature. The fusion feature fuses the features of the target point cloud data at different viewing angles. Compared with the features of the target point cloud data at each viewing angle, it has the advantages of richer information and stronger robustness, and provides data support for the network model in the subsequent step S40 to extract features with richer information and stronger robustness, thereby improving the overall performance of the network.
[0061] S40. Process the fused features according to the network model corresponding to the service type to detect the target object in the environment surrounding the vehicle body and obtain a detection result.
[0062] It should be understood that after executing step S30, the fused features can be subsequently sent as input information to the corresponding network model according to specific task requirements. Among them, the business type includes but is not limited to: target detection type, point cloud segmentation type, etc. When the business type is target detection type, the network model corresponding to the target detection type can be the target detection network PointPillars, CenterPoint, etc.; the above-mentioned targets include but are not limited to: other vehicles, lane lines, etc. When the business type is point cloud segmentation type, the network model corresponding to the point cloud segmentation type can be the point cloud segmentation network Cylinder3D, PolarNet, etc.
[0063] Taking the above-mentioned target detection network PointPillars as an example, step S40 is analyzed as follows: In this embodiment, the image Multi-View-Feature formed by the fused features is used as the network input, and the execution steps of the network PillarGenerator, point cloud feature extraction network (Pillar FeatureNetwork, PFN), backbone network Backbone, detection head DetectionHead and other structures contained in the target detection network PointPillars are sequentially executed to achieve accurate detection of the target object.
[0064] The embodiments of the present application are applied to the business of detecting target objects in autonomous driving. The embodiments of the present application are not limited by distance, and the distance scale is the same. When target objects such as speed limit signs and no right turn signs are detected in time, the autonomous driving strategy can be adjusted in time, which can achieve both the accuracy of target object detection and application.
[0065] The embodiment of the present application provides feature types corresponding to different perspectives while taking into account the computing power value of the target vehicle, and then extracts the features of the target point cloud data under each feature type respectively. Since there are multiple feature types, the fused features obtained after the features of the target point cloud data under each feature type are fused and processed can have the advantages of the features of the target point cloud data under each feature type, thereby improving the accuracy of target object detection and thus improving the safety of vehicle autonomous driving.
[0066] In a possible implementation, in the above Figure 1 Based on this, this embodiment Figure 1 Specifically, in step S10, obtaining target point cloud data of the surrounding environment of the target vehicle body includes the following steps:
[0067] S101, obtaining original point cloud data of the surrounding environment of the target vehicle body; wherein the original point cloud data is obtained by scanning the surrounding environment of the vehicle body by a laser radar device installed on the target vehicle body. It should be understood that the above laser radar device can be simply referred to as a laser radar.
[0068] S102. Preprocess the original point cloud data to obtain target point cloud data; wherein the target point cloud data includes coordinate data and reflection intensity values of each target point; wherein the target point is a reflection point of multiple laser beams emitted by a laser radar device contained in the original point cloud data, and each reflection point has its own reflection intensity value.
[0069] It should be understood that preprocessing may refer to formal unification. The embodiment of the present application does not specifically limit the type of preprocessing.
[0070] In the embodiment of the present application, step S102 can obtain the target point cloud data by establishing a three-dimensional coordinate system and projecting the original point cloud data into the three-dimensional coordinate system.
[0071] The embodiment of the present application reads the original point cloud data of the laser radar, and does not rely on the data characteristics of the mechanical, semi-solid, solid and other forms of laser radar. The point cloud data is represented by the three-dimensional coordinates (x, y, z) in the oxyz coordinate system and its reflection intensity i. The target point cloud data is represented by the symbol: Among them, 3 represents three dimensions and 1 represents the dimension of reflection intensity i.
[0072] The embodiment of the present application describes in detail the process of acquiring target point cloud data to provide data support for subsequent operations.
[0073] In a possible implementation, in the above Figure 1 Based on this, this embodiment Figure 1 Specifically, in step S20, according to the computing power value of the target vehicle, the corresponding feature type is determined, including any of the following:
[0074] When the computing power value of the target vehicle is not less than the first preset threshold, the corresponding feature type is determined to be a single point feature type, a bird's-eye view BEV feature type, and a depth map RV feature type.
[0075] When the computing power value of the target vehicle is greater than the second preset threshold and less than the first preset threshold, the corresponding feature type is determined to be at least two of the single-point feature type, the BEV feature type and the RV feature type; wherein the first preset threshold is higher than the second preset threshold.
[0076] When the computing power value of the target vehicle is not greater than the second preset threshold, the corresponding feature type is determined to be at least one of a single-point feature type, a BEV feature type, and a RV feature type.
[0077] It should be understood that the determination of different feature types can ensure that the present embodiment can freely combine features under different viewing angles. Due to the free combination of different feature types, the present embodiment can ensure the subsequent fusion of different viewing angle projection modes, thereby improving the accuracy of target object detection.
[0078] In a possible implementation, in the above Figure 1 Based on this, this embodiment Figure 1 Specifically, in step S20, the features of the target point cloud data under each feature type are extracted respectively, including at least one of the following:
[0079] The first item is to use a preset single-point feature extraction network to extract point features from the target point cloud data to obtain single-point features. It should be understood that the preset single-point feature extraction network is also called Point-NET. The embodiment of the present application can use Point-NET to extract single-point features P on each reflection point in the target point cloud data. c′ .
[0080] The second item is to project the target point cloud data from the BEV perspective to obtain the first BEV feature of the target point cloud data from the BEV perspective, and input the first BEV feature into the preset BEV feature extraction network to obtain the second BEV feature of the target point cloud data under the BEV feature type. It should be understood that the preset BEV feature extraction network is also called BEV-NET, which can use a two-dimensional convolutional network such as ResNet for feature extraction.
[0081] Specifically, the target point cloud data is projected onto the BEV perspective, and the first BEV feature obtained is recorded as Among them, chw represents the characteristics of three dimensions: channel, height, and width, and records the coordinate position of each reflection point on the BEV view. BEV . The first BEV features As the network feature input, it is input into the two-dimensional convolutional network, which extracts the second BEV feature
[0082] The third item is to project the target point cloud data from the RV perspective to obtain the first RV feature of the target point cloud data from the RV perspective, and input the first RV feature into the preset RV feature extraction network to obtain the second RV feature of the target point cloud data under the RV feature type. It should be understood that the preset RV feature extraction network is also called RV-NET, which can use a two-dimensional convolutional network for feature extraction.
[0083] Specifically, the target point cloud data is projected onto the RV perspective, and the first RV feature obtained is recorded as Among them, c″hw represents the characteristics of the three dimensions of channel number, height height and width width, and records the coordinate position of each reflection point on the RV view Coors RV . The first RV feature As the network feature input, it is input into the two-dimensional convolutional network, which extracts the second RV feature
[0084] For the same application scenario, the BEV view formed by the first BEV feature is as follows: Figure 2 As shown in , it includes roads, vehicles, height limit equipment, vehicles in front, etc. If these targets are small enough, they are regarded as a point. If they have a certain length, they are regarded as a line. The RV view formed by the first RV feature is as follows Figure 3 As shown, the road, vehicle, height limit equipment, vehicle ahead, etc. Figure 3 can be observed by the human eye.
[0085] The features extracted in different views in this embodiment can provide data support for subsequent fusion processing, thereby improving the accuracy of target object detection.
[0086] In a possible implementation, in the second item, the target point cloud data is projected from the BEV perspective, including:
[0087] The target point cloud data is projected at the BEV viewing angle using a first target projection method; wherein the first target projection method includes at least one of the following: a grid projection method and a Pillar projection method.
[0088] It should be understood that the first target projection method includes but is not limited to: a grid projection method, a Pillar projection method, etc. The grid projection method allows only the point with the highest z value in the pixel to be retained in the BEV view.
[0089] The embodiment of the present application refines the projection processing method of the BEV perspective, thereby providing data support for subsequent fusion processing.
[0090] In a possible implementation, in the third item, the target point cloud data is projected from the RV perspective, including:
[0091] The target point cloud data is projected under the RV viewing angle using a second target projection method; wherein the second target projection method includes at least one of the following: a cylindrical view projection method, a spherical view projection method, and an annular azimuth projection method.
[0092] It should be understood that the second target projection method includes but is not limited to: a cylindrical view projection method Cylindrical-View, a spherical view projection method Spherical-View, a ring-azimuth-View projection method Ring-Azimuth-View, etc.
[0093] The pixel coordinates in the RV view obtained by the above cylindrical view projection method Cylindrical-View are marked as in, Z i =z i , and (x i ,y i , z i )∈(x,y,z).
[0094] The pixel coordinates in the RV view obtained by the spherical-view projection method are marked as in,
[0095] The pixel coordinates in the RV view obtained by the above ring-azimuth-view projection method are marked as Among them, r i is the ring-id to which the point belongs,
[0096] The embodiment of the present application refines the projection processing method of the RV perspective, thereby providing data support for subsequent fusion processing.
[0097] In a possible implementation, in the above Figure 1 Based on this, this embodiment Figure 1 Specifically, step S30, fusing the features of the target point cloud data under all feature types to obtain fused features, includes the following steps:
[0098] Step S301: back-project the second BEV feature and / or the second RV feature to obtain a BEV point cloud feature and / or a RV point cloud feature.
[0099] During the back-projection process, the second BEV feature Interpolate back to the point cloud sequence corresponding to the target point cloud data to obtain the BEV point cloud features from the BEV perspective Similarly, the second RV feature Interpolate back to the point cloud sequence corresponding to the target point cloud data to obtain the RV point cloud features from the RV perspective
[0100] In the embodiment of the present application, the BEV point cloud view formed by the BEV point cloud feature is Figure 2 The BEV view in the figure is similar to that in the figure, except that the BEV point cloud view is more concise. Figure 3 The presentation of the RV view in is similar, the difference is that the RV point cloud view is more concise.
[0101] Step S302: concatenate the BEV point cloud features, RV point cloud features and / or single point features to obtain fused features.
[0102] It should be understood that BEV point cloud features, RV point cloud features, and single point features can all be used as view features. In this embodiment, in order to achieve the fusion of multi-view features, the features under each view are interpolated back to the point cloud sequence corresponding to the target point cloud data, and then spliced and fused. RV point cloud features and single point feature P c′ Fusion is performed according to the point order to obtain fusion features Later, according to the task requirements, the fusion features can be After corresponding processing, it is sent as input to the network model corresponding to the business type.
[0103] The above BEV point cloud features are more suitable for characterizing the movement of vehicles on a plane. When combined with the RV point cloud features, the fusion features can meet the application requirements while also meeting the accuracy of automatic detection. Therefore, when the fusion features fuse point cloud features from multiple different perspectives, they can be free from the constraints of the perspective projection method, improve the detection accuracy of the target, and thus improve the safety of autonomous driving.
[0104] Based on the above embodiments, the technical solution of the present application is described in more detail below in conjunction with specific embodiments.
[0105] Figure 4 A schematic diagram of another method for detecting a target object provided in an embodiment of the present application. Figure 4 It can be seen that the method of this embodiment includes the following steps:
[0106] S1. Obtain point cloud data and input the point cloud data into the target feature fusion network so that the target feature fusion network can perform preprocessing, projection, feature extraction, back-projection and fusion processing on the point cloud data to obtain fusion features.
[0107] This embodiment mainly relates to a target feature fusion network, which mainly includes multiple view feature extraction networks and a fusion network, wherein the multiple view feature extraction networks include: a BEV feature extraction network, an RV feature extraction network and a single point feature extraction network. After the multiple view feature extraction networks perform corresponding feature extraction respectively, the above fusion network can fuse the features of the point cloud data at each perspective, and provide better feature input for the network model corresponding to the subsequent business type.
[0108] S2. Input the fused features into the network model corresponding to the business type, so that the network model corresponding to the business type performs various processes on the fused features to obtain detection results.
[0109] In the embodiment of the present application, the network model corresponding to the service type can be understood as a deep network, which may include a network backbone.
[0110] It should be noted that different business types correspond to different network models. For example, data segmentation business corresponds to segmentation network, and target detection task corresponds to detection network.
[0111] Step S1 in the above-mentioned target object detection method is used as a method for laser radar multi-view network feature fusion. The method of freely combining the features of point cloud data under different feature types by computing power value can ensure the fast and efficient autonomous driving system. On this basis, the advantages of point cloud features under different perspectives are combined to provide better features for the deep network.
[0112] In the embodiment of the present application, the target feature fusion network in the above step S1 can implement a multi-feature fusion processing flow, and the multi-feature fusion processing flow is as follows: Figure 5 As shown in Figure 1, it includes preprocessing, projection under different views, extraction of different features through different feature extraction networks, fusion and other operations. Figure 5 It can be seen that the multi-feature fusion processing method includes the following steps:
[0113] S11. Acquire point cloud data, and preprocess the point cloud data to obtain preprocessed point cloud data.
[0114] S12, projecting the preprocessed point cloud data under the BEV view and the RV view respectively to obtain a first BEV feature and a first RV feature.
[0115] S13, respectively using the BEV feature extraction network, the RV feature extraction network and the single point feature extraction network to perform feature extraction to obtain a second BEV feature, a second RV feature and a single point feature.
[0116] It can be seen from the above embodiment 1 that the BEV feature extraction network is also called BEV-NET, the RV feature extraction network is also called RV-NET, and the single-point feature extraction network is also called Point-NET.
[0117] S14. Perform BEV back-projection on the second BEV feature, and perform RV back-projection on the second RV feature to obtain BEV point cloud features and RV point cloud features.
[0118] S15. Fuse the single point features in step S13 with the BEV point cloud features and RV point cloud features in step S14 to obtain fused features.
[0119] It can be seen from Example 1 that the above-mentioned fusion feature is also called a multi-view network feature, and the image formed by the fusion feature is a Multi-View-Feature.
[0120] The embodiment of the present application provides feature types corresponding to different perspectives while taking into account the computing power value of the target vehicle, and then extracts the features of the target point cloud data under each feature type respectively. Since there are multiple feature types, the fused features obtained after the features of the target point cloud data under each feature type are fused and processed can have the advantages of the features of the target point cloud data under each feature type, thereby improving the accuracy of target object detection and thus improving the safety of vehicle autonomous driving.
[0121] Figure 6 This is a schematic diagram of the structure of a target detection device provided in an embodiment of the present application. The device of this embodiment can be in the form of software and / or hardware. Figure 6 As shown, the target object detection device provided in this embodiment includes: an acquisition module 61, a determination and extraction module 62, a fusion processing module 63 and a detection module 64. Among them:
[0122] The acquisition module 61 is used to acquire the target point cloud data of the surrounding environment of the target vehicle body, business information and the computing power value of the target vehicle; wherein the business information includes the business type.
[0123] The determination extraction module 62 is used to determine the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, respectively extract the features of the target point cloud data under each feature type; wherein different feature types correspond to different viewing angles.
[0124] The fusion processing module 63 is used to perform fusion processing on the features of the target point cloud data under all feature types to obtain fusion features.
[0125] The detection module 64 is used to process the fused features according to the network model corresponding to the service type to detect the target object in the environment around the vehicle body and obtain the detection result.
[0126] In a possible implementation, the target object detection device is further used to perform any of the following:
[0127] When the computing power value of the target vehicle is not less than the first preset threshold, the corresponding feature type is determined to be a single point feature type, a BEV feature type, and a RV feature type.
[0128] When the computing power value of the target vehicle is greater than the second preset threshold and less than the first preset threshold, the corresponding feature type is determined to be at least two of the single-point feature type, the BEV feature type and the RV feature type; wherein the first preset threshold is higher than the second preset threshold.
[0129] When the computing power value of the target vehicle is not greater than the second preset threshold, the corresponding feature type is determined to be at least one of a single-point feature type, a BEV feature type, and a RV feature type.
[0130] In a possible implementation, the target object detection device is further used to perform any of the following:
[0131] First, a preset single-point feature extraction network is used to extract point features from the target point cloud data to obtain single-point features.
[0132] The second item is to project the target point cloud data from the BEV perspective to obtain the first BEV feature of the target point cloud data under the BEV perspective, and input the first BEV feature into the preset BEV feature extraction network to obtain the second BEV feature of the target point cloud data under the BEV feature type.
[0133] The third item is to project the target point cloud data from the RV perspective to obtain the first RV feature of the target point cloud data under the RV perspective, and input the first RV feature into the preset RV feature extraction network to obtain the second RV feature of the target point cloud data under the RV feature type.
[0134] In a possible implementation, the target object detection device is further used for:
[0135] The target point cloud data is projected at the BEV viewing angle using a first target projection method; wherein the first target projection method includes at least one of the following: a grid projection method and a Pillar projection method.
[0136] In a possible implementation, the target object detection device is further used for:
[0137] The target point cloud data is projected under the RV viewing angle using a second target projection method; wherein the second target projection method includes at least one of the following: a cylindrical view projection method, a spherical view projection method, and an annular azimuth projection method.
[0138] In a possible implementation, the fusion processing module 63 is further used to:
[0139] Back-projection processing is performed on the second BEV feature and / or the second RV feature to obtain a BEV point cloud feature and / or a RV point cloud feature.
[0140] The BEV point cloud features, RV point cloud features and / or single point features are spliced to obtain fused features.
[0141] In a possible implementation, the acquisition module 61 is further configured to:
[0142] The original point cloud data of the environment around the target vehicle body is obtained; wherein the original point cloud data is obtained by scanning the environment around the vehicle body by a laser radar device installed on the target vehicle body.
[0143] The original point cloud data is preprocessed to obtain target point cloud data; wherein the target point cloud data includes coordinate data and reflection intensity values of each target point; wherein the target point is a reflection point of multiple laser beams emitted by a laser radar device contained in the original point cloud data, and each reflection point has its own reflection intensity value.
[0144] The target object detection device provided in this embodiment can be used to execute the target object detection method provided in any of the above method embodiments. Its implementation principles and technical effects are similar and will not be described in detail here.
[0145] It should be noted that the user information and data involved in this application (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0146] That is to say, in the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the relevant laws and regulations and do not violate public order and good morals.
[0147] According to an embodiment of the present application, the present application also provides an electronic device and a readable storage medium.
[0148] Figure 7A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device includes a receiver 70, a transmitter 71, at least one processor 72 and a memory 73. The electronic device composed of the above components can be used to implement the above-mentioned several specific embodiments of the present application, which will not be described in detail here.
[0149] The embodiment of the present application further provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, each step of the method in the above embodiment is implemented.
[0150] An embodiment of the present application also provides a computer program product, including a computer program, which implements each step of the method in the above embodiment when executed by a processor.
[0151] Various embodiments of the systems and techniques described above in the present application can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor, which can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or electronic device.
[0153] In the context of the present application, a computer-readable storage medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may be a machine-readable signal medium or a machine-readable storage medium. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a computer-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0155] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data electronic device), or a computing system that includes middleware components (e.g., an application electronic device), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0156] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps disclosed in this application can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in this application can be achieved, and this document does not limit this.
[0157] The above specific implementations do not constitute a limitation on the protection scope of this application. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of this application should be included in the protection scope of this application.
Claims
1. A method for detecting a target object, characterized in that: include: Obtaining target point cloud data, business information and computing power value of the target vehicle's surrounding environment; wherein the business information includes business type; Determine the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, extract the features of the target point cloud data under each feature type respectively; wherein different feature types correspond to different viewing angles; Performing fusion processing on the features of the target point cloud data under all feature types to obtain fusion features; The fused features are processed according to a network model corresponding to the business type to detect a target object in the surrounding environment of the vehicle body and obtain a detection result.
2. The method according to claim 1, characterized in that: Determining the corresponding feature type according to the computing power value of the target vehicle includes any one of the following: When the computing power value of the target vehicle is not less than a first preset threshold, determining the corresponding feature type as a single point feature type, a bird's-eye view BEV feature type, and a depth map RV feature type; When the computing power value of the target vehicle is greater than a second preset threshold value and less than a first preset threshold value, determining that the corresponding feature type is at least two of a single point feature type, a BEV feature type, and an RV feature type; wherein the first preset threshold value is higher than the second preset threshold value; When the computing power value of the target vehicle is not greater than the second preset threshold, the corresponding feature type is determined to be at least one of a single-point feature type, a BEV feature type, and a RV feature type.
3. The method according to claim 1 or 2, characterized in that: The extracting the features of the target point cloud data under each feature type respectively includes at least one of the following: Using a preset single-point feature extraction network to extract point features from the target point cloud data to obtain single-point features; Performing a BEV perspective projection process on the target point cloud data to obtain a first BEV feature of the target point cloud data under the BEV perspective, and inputting the first BEV feature into a preset BEV feature extraction network to obtain a second BEV feature of the target point cloud data under the BEV feature type; The target point cloud data is projected from the RV perspective to obtain a first RV feature of the target point cloud data under the RV perspective, and the first RV feature is input into a preset RV feature extraction network to obtain a second RV feature of the target point cloud data under the RV feature type.
4. The method according to claim 3, characterized in that The projecting of the target point cloud data from a BEV perspective includes: The target point cloud data is projected at a BEV viewing angle using a first target projection method; wherein the first target projection method includes at least one of the following: a grid projection method and a Pillar projection method.
5. The method according to claim 3, characterized in that: The projecting process of the target point cloud data from the RV perspective includes: The target point cloud data is projected under the RV viewing angle using a second target projection method; wherein the second target projection method includes at least one of the following: a cylindrical view projection method, a spherical view projection method, and an annular azimuthal projection method.
6. The method according to claim 3, characterized in that The fusing of the features of the target point cloud data under all feature types to obtain fused features includes: Performing back-projection processing on the second BEV feature and / or the second RV feature to obtain a BEV point cloud feature and / or a RV point cloud feature; The BEV point cloud features, the RV point cloud features and / or the single point features are spliced to obtain fused features.
7. The method according to claim 1, characterized in that The step of acquiring target point cloud data of the surrounding environment of the target vehicle body includes: Acquire original point cloud data of the environment around the target vehicle body; wherein the original point cloud data is obtained by scanning the environment around the target vehicle body by a laser radar device disposed on the target vehicle body; The original point cloud data is preprocessed to obtain the target point cloud data; wherein the target point cloud data includes coordinate data and reflection intensity values of each target point; wherein the target point is a reflection point of multiple laser beams emitted by a laser radar device contained in the original point cloud data, and each of the reflection points has a respective reflection intensity value.
8. A target object detection device, characterized in that: include: An acquisition module, used to acquire target point cloud data of the surrounding environment of the target vehicle body, business information and computing power value of the target vehicle; wherein the business information includes business type; A determination extraction module is used to determine the corresponding feature type according to the computing power value of the target vehicle, and when there are multiple feature types, respectively extract the features of the target point cloud data under each feature type; wherein different feature types correspond to different viewing angles; A fusion processing module is used to fuse the features of the target point cloud data under all feature types to obtain fusion features; The detection module is used to process the fusion feature according to the network model corresponding to the business type to detect the target object in the environment around the vehicle body and obtain a detection result.
9. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the target object detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the target object detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Obstacle detection method and apparatus based on driverless technology and computer device
CN113678136A
Vehicle cloud computing power scheduling method and device, electronic equipment and storage medium
CN115061808A