Obstacle detection method and device

By performing feature extraction of the original point cloud of obstacles around the vehicle by BEV algorithm and semantic segmentation network, and fusing multi-view features, the problem of inaccurate obstacle recognition in the prior art is solved, and the accuracy of obstacle detection is improved.

CN120020906APending Publication Date: 2025-05-20BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311552490.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

Smart Images

  • Figure CN120020906A_ABST
    Figure CN120020906A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an obstacle detection method and device, and relates to the technical field of data processing. The method comprises the following steps: acquiring an original point cloud of an obstacle by using a sensor; performing feature extraction on the original point cloud by using a BEV algorithm to obtain a first semantic feature corresponding to each point cloud point in the original point cloud in a spatial voxel; performing feature extraction on the original point cloud by using a semantic segmentation network to obtain a second semantic feature corresponding to each point cloud point in the original point cloud in an all-round view angle; fusing the first semantic features and the second semantic features corresponding to the point cloud points to obtain multi-view features of the point cloud points in the original point cloud; and inputting the multi-view feature of each point cloud point in the original point cloud into an obstacle detection model, and determining a detection result of the obstacle through the output of the obstacle detection model. The problem that accuracy is low when obstacles around a vehicle are detected is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and more particularly, to a method and device for detecting obstacles. Background Art

[0002] In the field of autonomous driving, the recognition of various obstacles in the environment is very important. The accuracy of obstacle recognition directly affects subsequent vehicle control decisions. Existing obstacle recognition methods mostly use sensors to collect environmental data information around the vehicle, and then use the BEV (Bird Eye View) algorithm to recognize and generate obstacle information around the vehicle for the environmental data information. In the actual environment, there will be situations where the distance to the obstacle is relatively close, such as pedestrians under the branches, pedestrians beside large vehicles, etc. When these situations occur, the existing BEV (Bird Eye View) algorithm is very likely to miss the detection of pedestrians, and thus cannot generate the corresponding obstacle information, which will affect subsequent vehicle control decisions and even cause accidents.

[0003] Therefore, how to improve the accuracy of obstacle detection around the vehicle is a technical problem that needs to be solved urgently. Summary of the Invention

[0004] To solve the problem of low accuracy in detecting obstacles around the vehicle, embodiments of the present application provide a method, device, electronic device, and computer-readable storage medium for detecting obstacles.

[0005] In a first aspect, embodiments of the present application provide a method for detecting obstacles, including:

[0006] Obtaining the original point cloud of the obstacle by using a sensor;

[0007] Performing feature extraction on the original point cloud by using the BEV algorithm to obtain first semantic features respectively corresponding to each point cloud point in the original point cloud in a spatial voxel;

[0008] Performing feature extraction on the original point cloud by using a semantic segmentation network to obtain second semantic features respectively corresponding to each point cloud point in the original point cloud in a panoramic view;

[0009] Performing fusion processing on the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain multi-view features of each point cloud point in the original point cloud respectively;

[0010] Inputting the multi-view features of each point cloud point in the original point cloud into an obstacle detection model, and determining the detection result of the obstacle through the output of the obstacle detection model.

[0011] As an optional implementation manner of an embodiment of the present application, the step of using a semantic segmentation network to extract features from the original point cloud to obtain second semantic features respectively corresponding to each point cloud point in the original point cloud under the panoramic view includes:

[0012] Project the original point cloud under the panoramic view to obtain a two-dimensional image corresponding to the original point cloud, and there is a corresponding relationship between each pixel point in the two-dimensional image and each point cloud point in the original point cloud;

[0013] Input the two-dimensional image into a semantic segmentation network for semantic segmentation to obtain the second semantic feature of each pixel point among the pixel points.

[0014] As an optional implementation manner of an embodiment of the present application, before fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively, the method further includes:

[0015] Determine the point cloud points included in the spatial voxel and the point cloud points corresponding to each pixel point in the two-dimensional image, and establish a point cloud index;

[0016] Find the first semantic feature and the second semantic feature respectively corresponding to each point cloud point through the point cloud index.

[0017] As an optional implementation manner of an embodiment of the present application, the step of fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively includes:

[0018] Combine the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain a fusion feature respectively corresponding to each point cloud point in the original point cloud;

[0019] Input the fusion features respectively corresponding to each point cloud point in the original point cloud into a self-attention layer, and determine the multi-view features of each point cloud point in the original point cloud through the output of the self-attention layer.

[0020] As an optional implementation manner of an embodiment of the present application, the step of fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively includes:

[0021] Determine the weight parameter of the first semantic feature corresponding to the point cloud point and the weight parameter of the second semantic feature corresponding to the point cloud point;

[0022] Weight the first semantic feature and the second semantic feature corresponding to the point cloud points by the weight parameter of the first semantic feature and the weight parameter of the second semantic feature, and perform weighted summation to obtain the multi-view features of each point cloud point in the original point cloud respectively.

[0023] As an optional implementation manner of the embodiment of the present application, the obstacle detection model includes a backbone network and a prediction head network; inputting the multi-view features of each point cloud point in the original point cloud into the obstacle detection model, and determining the detection result of the obstacle through the output of the obstacle detection model includes:

[0024] Input the multi-view features respectively corresponding to each point cloud point in the original point cloud into the backbone network to obtain the global features of the obstacle;

[0025] Input the global features of the obstacle into the prediction head network, detect the obstacle information corresponding to the obstacle, and output the detection result of the obstacle. The obstacle information includes at least one of the category of the obstacle, the position of the obstacle, the size of the obstacle, and the orientation of the obstacle.

[0026] In a second aspect, an embodiment of the present application provides a training method for a detection model. The obstacle detection model includes: a backbone network and a prediction head network; the method includes:

[0027] Obtain the obstacle point cloud that has been labeled;

[0028] Use the BEV algorithm to extract features from the obstacle point cloud to obtain the first semantic features respectively corresponding to each point cloud point in the obstacle point cloud in the spatial voxel;

[0029] Use the semantic segmentation network to extract features from the obstacle point cloud to obtain the second semantic features respectively corresponding to each point cloud point in the obstacle point cloud from the surround view;

[0030] Fuse the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the obstacle point cloud respectively;

[0031] Input the multi-view features of each point cloud point in the obstacle point cloud into the obstacle detection model to be trained, train the obstacle detection model to be trained until the gap between the output information and the labeled information existing in the obstacle point cloud is less than a preset end condition, and obtain the trained obstacle detection model.

[0032] In a third aspect, an embodiment of the present application provides an obstacle detection device, including:

[0033] An acquisition module, configured to acquire the original point cloud of the obstacle by using a sensor;

[0034] An extraction module, configured to extract features from the original point cloud by using a BEV algorithm, so as to obtain first semantic features respectively corresponding to each point cloud point in the original point cloud in a spatial voxel;

[0035] The extraction module is further configured to extract features from the original point cloud by using a semantic segmentation network, so as to obtain second semantic features respectively corresponding to each point cloud point in the original point cloud in a surround view;

[0036] A processing module, configured to perform fusion processing on the first semantic feature and the second semantic feature corresponding to a point cloud point, so as to obtain multi-view features of each point cloud point in the original point cloud respectively;

[0037] A detection module, configured to input the multi-view features of each point cloud point in the original point cloud into an obstacle detection model, and determine a detection result of the obstacle through an output of the obstacle detection model.

[0038] As an optional implementation manner of an embodiment of the present application, the extraction module is specifically configured to project the original point cloud in the surround view to obtain a two-dimensional image corresponding to the original point cloud, and there is a corresponding relationship between each pixel point in the two-dimensional image and each point cloud point in the original point cloud;

[0039] Input the two-dimensional image into a semantic segmentation network for semantic segmentation, and obtain a second semantic feature of each pixel point among the pixel points.

[0040] As an optional implementation manner of an embodiment of the present application, the processing module is further configured to determine point cloud points included in the spatial voxel and point cloud points corresponding to each pixel point in the two-dimensional image, and establish a point cloud index before performing fusion processing on the first semantic feature and the second semantic feature corresponding to a point cloud point to obtain multi-view features of each point cloud point in the original point cloud respectively;

[0041] Search for the first semantic feature and the second semantic feature respectively corresponding to each point cloud point through the point cloud index.

[0042] As an optional implementation manner of an embodiment of the present application, the processing module is specifically configured to combine the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain a fusion feature respectively corresponding to each point cloud point in the original point cloud;

[0043] Input the fusion feature respectively corresponding to each point cloud point in the original point cloud into a self-attention layer, and determine the multi-view feature of each point cloud point in the original point cloud through an output of the self-attention layer.

[0044] As an optional implementation manner of an embodiment of the present application, the processing module is specifically configured to determine a weight parameter of a first semantic feature corresponding to the point cloud point and a weight parameter of a second semantic feature corresponding to the point cloud;

[0045] Perform weighted summation on the first semantic feature and the second semantic feature corresponding to the point cloud point through the weight parameter of the first semantic feature and the weight parameter of the second semantic feature, respectively, to obtain multi-view features of each point cloud point in the original point cloud.

[0046] As an optional implementation manner of an embodiment of the present application, the detection module is specifically configured to input the multi-view features respectively corresponding to each point cloud point in the original point cloud into the backbone network to obtain the global feature of the obstacle;

[0047] Input the global feature of the obstacle into the prediction head network, detect the obstacle information corresponding to the obstacle, and output the detection result of the obstacle. The obstacle information includes at least one of the category of the obstacle, the position of the obstacle, the size of the obstacle, and the orientation of the obstacle.

[0048] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor. The memory is used to store a computer program, and the processor is used to execute the obstacle detection method according to the first aspect or any optional implementation manner of the first aspect when calling the computer program.

[0049] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the obstacle detection method according to the first aspect or any optional implementation manner of the first aspect is implemented.

[0050] The technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0051] The obstacle detection method provided by the embodiments of this application includes: obtaining the original point cloud of the obstacle using a sensor; extracting features from the original point cloud using the BEV algorithm to obtain the first semantic features corresponding to each point cloud point in the original point cloud in the spatial voxel; extracting features from the original point cloud using a semantic segmentation network to obtain the second semantic features corresponding to each point cloud point in the original point cloud in the surround view; fusing the first semantic features and the second semantic features corresponding to the point cloud points to obtain the multi-view features of each point cloud point in the original point cloud; inputting the multi-view features of each point cloud point in the original point cloud into an obstacle detection model, and determining the detection result of the obstacle through the output of the obstacle detection model. In the embodiments of this application, the first semantic feature and the second semantic feature of each point cloud point are obtained. Compared with using the voxel feature to represent the features of all point cloud points in the voxel, all point cloud points in each voxel are given the same semantic feature, which makes the semantic features given to some point cloud points not match the actual situation, thus causing inaccurate obstacle recognition. Therefore, on the one hand, the embodiments of this application improve the accuracy of the first semantic feature and the second semantic feature of the obtained point cloud points. On the other hand, the first semantic feature and the second semantic feature of each point cloud point in the embodiments of this application are semantic features from different perspectives. Therefore, more accurate multi-view features can be obtained after fusing the initial semantic features. Based on the above two aspects, the detection accuracy of the point cloud can be improved, and thus the recognition accuracy of the obstacle can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0053] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 is a flowchart of the obstacle detection method provided according to one or more embodiments of this application;

[0055] Figure 2 is a flowchart of the obstacle detection method provided according to one or more embodiments of this application;

[0056] Figure 3 is a structural block diagram of the obstacle detection device provided according to one or more embodiments of this application;

[0057] Figure 4The internal structure diagram of an in-vehicle terminal provided according to one or more embodiments of the present application. Detailed implementation manners

[0058] To make the objectives, implementation manners, and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0059] Based on the exemplary embodiments described in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope protected by the claims of the present application. In addition, although the disclosed content in the present application is introduced according to one or several exemplary instances, it should be understood that each aspect of these disclosed contents can also constitute a complete implementation manner alone. It should be noted that the brief description of terms in the present application is only for the convenience of understanding the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.

[0060] First, the application scenario of the embodiments of the present application is described: LiDAR plays a very important role in driverless driving. During the scanning process of LiDAR, it will identify the obstacles around the vehicle and clarify the positions of the obstacles in space to ensure the safe driving of the vehicle.

[0061] In traditional detection methods, the BEV (Bird Eye View) perception algorithm can well fuse the features of multiple sensors. Usually, voxels are constructed to perform feature processing on the point cloud, and the voxel features are assigned to all the point clouds in the voxel, that is, all the point cloud features in a voxel are the same and are all voxel features. On the one hand, when the objects around the vehicle are relatively dense, it is easy to cause missed detections, such as missing pedestrians under the branches, missing the whole row of pedestrians, missing pedestrians next to large vehicles, etc., and the detection accuracy of obstacles is relatively low.

[0062] Based on this, this embodiment provides a method for detecting obstacles. By obtaining the first semantic feature of each point cloud in the voxel in the bird's-eye view and the second semantic feature in the RV (Range View), compared with using the voxel feature to represent the features of all the point clouds in the voxel, the accuracy of the obtained point cloud features can be improved. The first semantic feature and the second semantic feature are semantic features from different perspectives. Therefore, after fusing the initial semantic features including the first semantic feature and the second semantic feature, more accurate multi-view features can be obtained, the detection accuracy of the point cloud can be improved, and thus the recognition accuracy of the obstacles can be improved.

[0063] The vehicle obstacle detection method provided by the embodiments of the present application can be executed by the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be, but is not limited to, an in-vehicle terminal (such as an in-vehicle computer), an in-vehicle infotainment product, etc. The embodiments of the present application do not make specific limitations.

[0064] The following are several specific embodiments to elaborate in detail on the obstacle detection method provided by the embodiments of the present application.

[0065] Figure 1 It is a flowchart of the obstacle detection method provided by an embodiment of the present application. Referring to Figure 1 As shown, the obstacle detection method provided by this embodiment includes the following steps:

[0066] S11. Use a sensor to obtain the original point cloud of the obstacle.

[0067] The original point cloud includes multiple point cloud points. The original point cloud contains the position information of each point cloud point on the X-axis, Y-axis, and Z-axis in a specific coordinate system (such as a vehicle coordinate system, etc.), as well as the reflection intensity. The sensor (such as a lidar) is set on the vehicle and is used to obtain the point cloud data of the obstacles around the vehicle. This point cloud data includes information such as point cloud, point cloud position (such as spatial coordinates), point cloud quantity, timestamp, etc.

[0068] S12. Use the BEV algorithm to extract features from the original point cloud to obtain the first semantic features corresponding to each point cloud point in the spatial voxel of the original point cloud, and use a semantic segmentation network to extract features from the original point cloud to obtain the second semantic features corresponding to each point cloud point in the surround view perspective.

[0069] First, the original point cloud can be converted into a point cloud in the BEV perspective and a point cloud in the RV (surround view) perspective. Then, by extracting features from the point cloud in the BEV perspective, the first semantic features corresponding to each point cloud point in the spatial voxel are obtained, and by extracting features from the point cloud in the RV perspective, the second semantic features corresponding to each point cloud point in the RV perspective are obtained.

[0070] Exemplarily, divide the space where the original point cloud is located into spatial voxels, determine the spatial voxels to which each point cloud point in the original point cloud belongs in combination with the spatial position, use the BEV algorithm to determine the first semantic features corresponding to each spatial voxel, and then determine the first semantic features corresponding to each point cloud point according to the spatial voxel to which the point cloud point belongs.

[0071] Under the surround viewing angle, the original point cloud is projected to obtain a two-dimensional image corresponding to the original point cloud, and each pixel in the two-dimensional image corresponds to each point cloud point in the original point cloud; the two-dimensional image is input into the semantic segmentation network for semantic segmentation to obtain the second semantic feature of each pixel in the pixel.

[0072] For example, there are many ways to divide the space where the original point cloud is located into voxels, such as dividing a space into N cubes; thereafter, since the position (or coordinates) of each point cloud point in the space is known, the point cloud points contained in each spatial voxel in the space can be determined, and then the first semantic features (such as obstacle information, etc.) of each spatial voxel in the space can be obtained by using the BEV algorithm (such as the PointNet++ network model in the algorithm). In addition, the spatial voxel in which the point cloud point is located is also obtained before, so the first semantic features of the spatial voxel can be assigned to the point cloud point in the spatial voxel, so that the first semantic features corresponding to each point cloud point are obtained.

[0073] Exemplarily, there are many ways to project the original point cloud onto a preset projection surface, such as presetting a projection surface in the space where the original point cloud is located, and then projecting each point cloud point in the original point cloud onto the projection surface in a set direction or a direction perpendicular to the projection surface, thereby obtaining a corresponding two-dimensional image. At this time, the correspondence between each pixel point in the two-dimensional image and each point cloud point has been determined during the projection process, and then the two-dimensional image is input into a trained semantic segmentation network, and the second semantic features corresponding to each pixel point in the two-dimensional image are output. In addition, each pixel point in the two-dimensional image has a corresponding relationship with each point cloud point in the original point cloud, so the second semantic features corresponding to each point cloud point are obtained.

[0074] For the point cloud under the BEV perspective and the point cloud under the RV perspective, establish the correspondence between the point cloud points, that is, find the corresponding pixel points of the point cloud points under the BEV perspective in the two-dimensional image under the RV perspective, so that the first semantic feature and the second semantic feature of the same point cloud point can be attributed to the corresponding point cloud point. Exemplarily, determine the point cloud points contained in the spatial voxels and the point cloud points corresponding to each pixel point in the two-dimensional point cloud, and establish a point cloud index; when determining the first semantic feature and the second semantic feature of any point cloud point, the first semantic feature and the second semantic feature corresponding to the point cloud point can be found through the point cloud index.

[0075] That is, in the embodiments of the present application, the first semantic features corresponding to each point cloud point are obtained. In the traditional method, the point cloud data is often processed in units of spatial voxels. First, the original point cloud data is processed into multiple spatial voxels, and then the spatial voxels occupied by obstacles are determined. The semantic feature of an obstacle voxel is assigned to each point cloud point in the spatial voxel, that is, the same semantic feature is assigned to all point cloud points in the spatial voxel of each obstacle. However, in practice, the obstacles corresponding to some point cloud points are inconsistent with the obstacles corresponding to the spatial voxels. For example, a certain spatial voxel includes two different obstacles, and there are more point cloud points of one obstacle. When actually assigning values, the semantic features assigned to the small number of point cloud points corresponding to the other obstacle do not match the actual situation, which may cause inaccurate subsequent obstacle recognition. However, when the semantic features of each point cloud point are obtained in the embodiments of the present application, two methods are used, so that the obstacles corresponding to the point cloud points can be more accurately recognized in the subsequent process.

[0076] Among them, the semantic segmentation network is obtained through supervised training, that is, by labeling the point cloud data and sampling the labeled point cloud data to train the semantic segmentation network until the training end condition is met, and a trained semantic segmentation network is obtained. The initial semantic segmentation network can be a Unet network, an FCN network, etc., and the present embodiment does not make specific limitations.

[0077] S13. Fuse the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively.

[0078] Before fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively, the first semantic feature and the second semantic feature corresponding to the point cloud point are found through point cloud indexing.

[0079] Exemplarily, fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud respectively can be achieved by combining the first semantic feature and the second semantic feature, and inputting the combined feature into a self-attention model for fusion to obtain the multi-view features of each point cloud point. For example, it is implemented in the following manner: combine the first semantic feature and the second semantic feature corresponding to each point cloud point respectively to obtain the fusion feature corresponding to each point cloud point in the original point cloud; input the fusion feature corresponding to each point cloud point in the original point cloud into the self-attention model, and determine the multi-view features of each point cloud point in the original point cloud through the output of the self-attention model.

[0080] Exemplarily, if the first semantic feature of the point cloud point O is a 1*M matrix and the second semantic feature of the point cloud point O is a 1*N matrix, then the first semantic feature and the second semantic feature of the point cloud point O are combined to obtain a fusion feature. The fusion feature can be obtained by existing methods such as multiplying the two features. The fusion feature of the point cloud point O is input into the self-attention model, and the self-attention model learns which features are more important and outputs multi-view features. The self-attention model can use the existing feature attention model, and the learning mechanism and the related usage process of the model are not elaborated in detail in the embodiments of the present application.

[0081] Alternatively, the multi-view features of each point cloud point in the original point cloud can be obtained through the first semantic feature and the second semantic feature respectively corresponding to each point cloud point in the original point cloud, and can be implemented in the following manner: determining the weight parameter of the first semantic feature corresponding to the point cloud point and the weight parameter of the second semantic feature corresponding to the point cloud; performing weighted summation on the first semantic feature and the second semantic feature corresponding to the point cloud point through the weight parameter of the first semantic feature and the weight parameter of the second semantic feature, such as substituting the weights corresponding to the matrices during the multiplication of the two matrices, so as to obtain the multi-view features of each point cloud point in the original point cloud respectively.

[0082] S14. Input the multi-view features of each point cloud point in the original point cloud into the obstacle detection model, and determine the detection result of the obstacle through the output of the obstacle detection model.

[0083] Wherein, the obstacle detection model includes a backbone network and a prediction head network, and the prediction head network is spliced after the backbone network. Input the multi-view features respectively corresponding to each point cloud point in the original point cloud into the backbone network to obtain the global feature of the obstacle; input the global feature of the obstacle into the prediction head network to detect the obstacle information corresponding to the obstacle and output the detection result of the obstacle. The obstacle information includes at least one of the category of the obstacle, the position of the obstacle, the size of the obstacle, and the orientation of the obstacle.

[0084] That is, at least one of the obstacle information such as the category of the obstacle, the position of the obstacle, the size of the obstacle, and the orientation of the obstacle is output through the prediction head network. The obstacle may be a pedestrian, another vehicle, a tree, etc.

[0085] In addition, the obstacle detection model can be obtained through the following training method:

[0086] Obtain the labeled obstacle point cloud; use the BEV algorithm to extract features from the obstacle point cloud to obtain the first semantic features corresponding to each point cloud point in the obstacle point cloud in the spatial voxel; use the semantic segmentation network to extract features from the obstacle point cloud to obtain the second semantic features corresponding to each point cloud point in the obstacle point cloud in the surround view; fuse the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the obstacle point cloud; input the multi-view features of each point cloud point in the obstacle point cloud into the obstacle detection model to be trained, and train the obstacle detection model to be trained until the gap between the output information and the labeled information existing in the obstacle point cloud is less than the preset end condition to obtain the trained obstacle detection model.

[0087] Among them, the labeled obstacle point cloud contains not only the original information carried by the sensor when collecting the point cloud, but also the labeled obstacle information.

[0088] The process of extracting the first semantic feature and the second semantic feature from the labeled obstacle point cloud, and the process of fusing the first semantic feature and the second semantic feature should be consistent with the previous process of processing the original point cloud.

[0089] The obstacle detection method provided by the embodiment of the present application includes: using a sensor to obtain the original point cloud of the obstacle; using the BEV algorithm to extract features from the original point cloud to obtain the first semantic features corresponding to each point cloud point in the original point cloud in the spatial voxel; using the semantic segmentation network to extract features from the original point cloud to obtain the second semantic features corresponding to each point cloud point in the original point cloud in the surround view; fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud; inputting the multi-view features of each point cloud point in the original point cloud into the obstacle detection model, and determining the detection result of the obstacle through the output of the obstacle detection model. In the embodiment of the present application, the first semantic feature and the second semantic feature of each point cloud point are obtained. Compared with using the voxel feature to represent the features of all point cloud points in the voxel, all point cloud points in each voxel are given the same semantic feature, which makes the semantic feature given to some point cloud points inconsistent with the actual situation, resulting in inaccurate obstacle recognition. Therefore, on the one hand, the embodiment of the present application improves the accuracy of the first semantic feature and the second semantic feature of the obtained point cloud points. On the other hand, the first semantic feature and the second semantic feature of each point cloud point in the embodiment of the present application are semantic features from different perspectives. Therefore, more accurate multi-view features can be obtained after fusing the initial semantic features. Based on the above two aspects, the detection accuracy of the point cloud can be improved, and thus the recognition accuracy of the obstacle can be improved.

[0090] Refer to Figure 2As shown Figure 2 is a schematic flowchart of a method for detecting obstacles provided in another embodiment of the present application. Based on the embodiment shown Figure 1 in the above embodiment Figure 1 Step S13 in the embodiment shown may include the following steps S21 - S22. In this embodiment, the same or similar steps as those in the embodiment shown Figure 1 will not be elaborated again, and reference can be made to the explanation and description of the embodiment shown Figure 1 for details

[0091] S11. Use a sensor to obtain the original point cloud of the obstacle

[0092] S12. Use the BEV algorithm to extract features from the original point cloud to obtain the first semantic features respectively corresponding to each point cloud point in the spatial voxel of the original point cloud, and use a semantic segmentation network to extract features from the original point cloud to obtain the second semantic features respectively corresponding to each point cloud point in the panoramic view

[0093] S21. Combine the first semantic feature and the second semantic feature respectively corresponding to each point cloud point to obtain the fusion feature respectively corresponding to each point cloud point in the original point cloud

[0094] Among them, the first semantic feature and the second semantic feature of each point cloud point can be determined based on the point cloud index. The point cloud index includes the identifier of the point cloud point. Exemplarily, when the identifier of point cloud point A is identifier A, determine the first semantic feature corresponding to identifier A and the second semantic feature corresponding to identifier A, and combine the first semantic feature and the second semantic feature corresponding to identifier A to obtain the fusion feature of point cloud point A; when the identifier of point cloud point B is identifier B, determine the first semantic feature corresponding to identifier B and the second semantic feature corresponding to identifier B, and combine the first semantic feature and the second semantic feature corresponding to identifier B to obtain the fusion feature of point cloud point B

[0095] S22. Input the fusion features respectively corresponding to each point cloud point in the original point cloud into the self - attention model, and determine the multi - view features of each point cloud point in the original point cloud through the output of the self - attention model

[0096] Combined with the above example, input the fusion features of each point cloud point obtained into the self - attention model, extract important features through the self - attention mechanism, and output the multi - view features of each point cloud point in the original point cloud

[0097] S14. Input the multi - view features of each point cloud point in the original point cloud into the obstacle detection model, and determine the detection result of the obstacle through the output of the obstacle detection model

[0098] In this embodiment, through the self-attention model, the important features and correlations in the first semantic feature and the second semantic feature corresponding to each point cloud point are obtained, which helps to accurately judge the type of obstacle.

[0099] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application also provides a detection device for obstacles provided in the above embodiment. This device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this device embodiment. However, it should be clear that the detection device for obstacles in this embodiment can correspondingly implement all the contents in the foregoing method embodiment.

[0100] Figure 3 It is a schematic structural diagram of a detection device for obstacles provided in an embodiment of the present application, as Figure 3 shown. The detection device 300 for obstacles provided in this embodiment includes:

[0101] An acquisition module 310, configured to acquire the original point cloud of the obstacle by using a sensor;

[0102] An extraction module 320, configured to perform feature extraction on the original point cloud by using the BEV algorithm to obtain the first semantic feature corresponding to each point cloud point in the spatial voxel in the original point cloud;

[0103] The extraction module 320 is further configured to perform feature extraction on the original point cloud by using a semantic segmentation network to obtain the second semantic feature corresponding to each point cloud point in the original point cloud from the surround view perspective;

[0104] A processing module 330, configured to perform fusion processing on the first semantic feature and the second semantic feature corresponding to the point cloud point to respectively obtain the multi-view feature of each point cloud point in the original point cloud;

[0105] A detection module 340, configured to input the multi-view feature of each point cloud point in the original point cloud into an obstacle detection model, and determine the detection result of the obstacle through the output of the obstacle detection model.

[0106] As an optional implementation manner of an embodiment of the present application, the extraction module 320 is specifically configured to project the original point cloud from the surround view perspective to obtain a two-dimensional image corresponding to the original point cloud, and there is a corresponding relationship between each pixel point in the two-dimensional image and each point cloud point in the original point cloud; input the two-dimensional image into a semantic segmentation network for semantic segmentation to obtain the second semantic feature of each pixel point among the pixel points.

[0107] As an alternative implementation manner of the embodiment of the present application, the processing module 330 is further configured to determine the point cloud points included in the spatial voxel and the point cloud points corresponding to each pixel point of the two-dimensional image, and establish a point cloud index before fusing the first semantic feature and the second semantic feature corresponding to the point cloud points to obtain the multi-view features of each point cloud point in the original point cloud; and find the first semantic feature and the second semantic feature corresponding to each point cloud point through the point cloud index.

[0108] As an alternative implementation manner of the embodiment of the present application, the processing module 330 is specifically configured to combine the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the fused feature corresponding to each point cloud point in the original point cloud; and input the fused feature corresponding to each point cloud point in the original point cloud into the self-attention layer, and determine the multi-view features of each point cloud point in the original point cloud through the output of the self-attention layer.

[0109] As an alternative implementation manner of the embodiment of the present application, the processing module 330 is specifically configured to determine the weight parameter of the first semantic feature corresponding to the point cloud point and the weight parameter of the second semantic feature corresponding to the point cloud; and perform weighted summation on the first semantic feature and the second semantic feature corresponding to the point cloud point through the weight parameter of the first semantic feature and the weight parameter of the second semantic feature to obtain the multi-view features of each point cloud point in the original point cloud.

[0110] As an alternative implementation manner of the embodiment of the present application, the detection module 340 is specifically configured to input the multi-view features corresponding to each point cloud point in the original point cloud into the backbone network to obtain the global feature of the obstacle; input the global feature of the obstacle into the prediction head network to detect the obstacle information corresponding to the obstacle, and output the detection result of the obstacle, where the obstacle information includes at least one of the category of the obstacle, the position of the obstacle, the size of the obstacle, and the orientation of the obstacle.

[0111] The obstacle detection device provided by the embodiment of the present application can execute the obstacle detection method provided by the above method embodiment, and its implementation principle and technical effect are similar, and will not be described in detail here. Each module in the above obstacle detection device can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0112] In one embodiment, an electronic device is provided, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the obstacle detection methods described in the above method embodiments are implemented.

[0113] Exemplarily, Figure 4 FIG. is a schematic structural diagram of the electronic device provided in the embodiment of the present application. As Figure 4 shown, the electronic device provided in this embodiment includes: a memory 41 and a processor 42. The memory 41 is used to store a computer program; the processor 42 is used to execute the steps in the obstacle detection method provided in the above method embodiment when calling the computer program. The implementation principle and technical effect are similar and will not be elaborated here. Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the vehicle-mounted terminal to which the solution of the present application is applied. The specific vehicle-mounted terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0114] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the obstacle detection methods described in the above method embodiments are implemented.

[0115] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, a database, or other media used in the various embodiments provided in the present application may include at least one of non-volatile and volatile memories. The non-volatile memory may include a read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. The volatile memory may include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM).

[0116] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0117] For ease of explanation, the above description has been made in connection with specific embodiments. However, the above discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different modified embodiments suitable for specific use considerations.

Claims

1. A method for detecting an obstacle, characterized in that: include: Use sensors to obtain the original point cloud of obstacles; Using the BEV algorithm to extract features from the original point cloud, obtaining first semantic features corresponding to each point cloud point in the original point cloud in the spatial voxel; Using a semantic segmentation network to extract features from the original point cloud, obtaining second semantic features corresponding to each point cloud point in the original point cloud under a surround viewing angle; The first semantic feature and the second semantic feature corresponding to the point cloud point are fused to obtain multi-view features of each point cloud point in the original point cloud; The multi-view features of each point cloud point in the original point cloud are input into an obstacle detection model, and the detection result of the obstacle is determined through the output of the obstacle detection model.

2. The method according to claim 1, characterized in that The extracting features of the original point cloud by using a semantic segmentation network to obtain second semantic features corresponding to each point cloud point in the original point cloud under a surround viewing angle includes: Under the surround viewing angle, the original point cloud is projected to obtain a two-dimensional image corresponding to the original point cloud, and each pixel point in the two-dimensional image corresponds to each point cloud point in the original point cloud; The two-dimensional image is input into a semantic segmentation network for semantic segmentation to obtain a second semantic feature of each pixel in the pixel points.

3. The method according to claim 2, characterized in that Before fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain the multi-view features of each point cloud point in the original point cloud, the method further includes: Determine the point cloud points contained in the spatial voxel and the point cloud points corresponding to each pixel point in each pixel point of the two-dimensional image, and establish a point cloud index; The first semantic feature and the second semantic feature respectively corresponding to each point cloud point are searched through the point cloud index.

4. The method according to any one of claims 1 to 3, characterized in that: The first semantic feature and the second semantic feature corresponding to the point cloud point are fused to obtain the multi-view features of each point cloud point in the original point cloud, including: Combining the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain fusion features corresponding to each point cloud point in the original point cloud; The fused features corresponding to each point cloud point in the original point cloud are respectively input into the self-attention model, and the multi-view features of each point cloud point in the original point cloud are determined through the output of the self-attention model.

5. The method according to any one of claims 1 to 3, characterized in that: The first semantic feature and the second semantic feature corresponding to the point cloud point are fused to obtain the multi-view features of each point cloud point in the original point cloud, including: Determining a weight parameter of a first semantic feature corresponding to the point cloud point and a weight parameter of a second semantic feature corresponding to the point cloud point; The first semantic feature and the second semantic feature corresponding to the point cloud point are weightedly summed by using the weight parameter of the first semantic feature and the weight parameter of the second semantic feature to obtain multi-view features of each point cloud point in the original point cloud.

6. The method according to claim 1, characterized in that The obstacle detection model includes a backbone network and a prediction head network; the multi-view features of each point cloud point in the original point cloud are input into the obstacle detection model, and the obstacle detection result is determined by the output of the obstacle detection model, including: Inputting the multi-view features corresponding to each point cloud point in the original point cloud into the backbone network to obtain the global features of the obstacle; The global feature of the obstacle is input into the prediction head network, obstacle information corresponding to the obstacle is detected, and the detection result of the obstacle is output, wherein the obstacle information includes at least one of the category of the obstacle, the position of the obstacle, the size of the obstacle, and the direction of the obstacle.

7. A method for training an obstacle detection model, characterized in that: The obstacle detection model includes: a backbone network and a prediction head network; The method comprises: Obtain the marked obstacle point cloud; Using the BEV algorithm to extract features from the obstacle point cloud, obtaining first semantic features corresponding to each point cloud point in the obstacle point cloud in the spatial voxel; Using a semantic segmentation network to extract features from the obstacle point cloud, obtaining second semantic features corresponding to each point cloud point in the obstacle point cloud under a surround viewing angle; The first semantic feature and the second semantic feature corresponding to the point cloud point are fused to obtain multi-view features of each point cloud point in the obstacle point cloud; The multi-view features of each point cloud point in the obstacle point cloud are input into the obstacle detection model to be trained, and the obstacle detection model to be trained is trained until the difference between the output information and the annotation information existing in the obstacle point cloud is less than a preset end condition, thereby obtaining a trained obstacle detection model.

8. An obstacle detection device, characterized in that: include: An acquisition module is used to acquire the original point cloud of obstacles using sensors; An extraction module, used for performing feature extraction on the original point cloud using a BEV algorithm to obtain first semantic features corresponding to each point cloud point in the original point cloud in a spatial voxel; The extraction module is further used to extract features from the original point cloud using a semantic segmentation network to obtain second semantic features corresponding to each point cloud point in the original point cloud under a surround viewing angle; A processing module, used for fusing the first semantic feature and the second semantic feature corresponding to the point cloud point to obtain multi-view features of each point cloud point in the original point cloud; The detection module is used to input the multi-view features of each point cloud point in the original point cloud into the obstacle detection model, and determine the detection result of the obstacle through the output of the obstacle detection model.

9. An electronic device, comprising: A memory and a processor, wherein the memory stores a computer program, wherein the processor implements the obstacle detection method described in any one of claims 1 to 6 when executing the computer program, or implements the obstacle detection model training method described in claim 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the obstacle detection method described in any one of claims 1 to 6 is implemented, or, when the computer program is executed by a processor, the obstacle detection model training method described in claims 1 to 7 is implemented.

11. A vehicle, characterized in that: The vehicle is equipped with the obstacle detection device according to claim 8, the electronic device according to claim 9, or the storage medium according to claim 10.