Target detection method and device, vehicle and computer readable storage medium

By sampling and feature fusion of the original point cloud, and combining point cloud points and BEV features, the problem of low target detection accuracy in existing technologies is solved, and higher detection accuracy is achieved.

CN117037098BActive Publication Date: 2025-10-21CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311013309.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2025-10-21
Estimated Expiration
2043-08-11

AI Technical Summary

Technical Problem

The accuracy of target detection based on point cloud features in existing technologies is relatively low.

Method used

By sampling the original point cloud, a set of key points of the target object is obtained. The third point feature of the key points is determined by combining the first point feature of the point cloud points and the BEV feature from the bird's-eye view. The accuracy of target detection is improved by voxel segmentation and attention network processing.

Benefits of technology

The accuracy of target detection has been improved by fusing information from point cloud points and BEV features, which enhances the feature richness and diversity of key points and improves the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117037098B_ABST
    Figure CN117037098B_ABST
Patent Text Reader

Abstract

The present disclosure provides a target detection method and device, a vehicle and a computer readable storage medium, relating to the technical field of computer vision, the method comprising: sampling an original point cloud to obtain a key point set of a target object, the original point cloud comprising a point cloud of the target object; determining a second point feature of each key point in the key point set according to a first point feature of part of point cloud points in the original point cloud; determining a second BEV feature of each key point in the key point set according to a first bird's eye view (BEV) feature of part of point cloud points in the original point cloud under a bird's eye view; determining a third point feature of each key point according to the second point feature and the second BEV feature; and determining a detection result of the target object according to the third point feature of each key point. In this way, the accuracy of the detection result of the target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and in particular to a target detection method, device, vehicle, and computer-readable storage medium. Background Art

[0002] In fields such as autonomous driving, it is necessary to detect target objects in the surrounding environment (such as the surrounding environment of a vehicle).

[0003] In the related art, point cloud data of the surrounding environment is acquired, and features are extracted from the collected point cloud data to detect target objects based on the point features of the extracted point cloud points. Summary of the Invention

[0004] In the related art, the accuracy of the detection result obtained by performing target detection on the target object only based on the point features of the point cloud points is low.

[0005] In order to solve the above problems, the embodiments of the present disclosure propose the following solutions.

[0006] According to one aspect of an embodiment of the present disclosure, a target detection method is provided, including: sampling an original point cloud to obtain a key point set of a target object, the original point cloud including a point cloud of the target object; determining a second point feature of each key point in the key point set based on a first point feature of some point cloud points in the original point cloud; determining a second BEV feature of each key point in the key point set based on a first bird's-eye view (BEV) feature of some point cloud points in the original point cloud under a bird's-eye view perspective; determining a third point feature of each key point based on the second point feature and the second BEV feature; and determining a detection result of the target object based on the third point feature of each key point.

[0007] In some embodiments, the method further includes: determining a target detection area including a point cloud of the target object from the original point cloud; wherein the key point set is obtained by sampling only the target detection area.

[0008] In some embodiments, the method further includes: performing voxel segmentation on the target detection area to obtain multiple first voxels; wherein, determining the detection result of the target object based on the third point feature of each key point includes: determining the feature of each first voxel based on the third point feature of each key point in the multiple first voxels; determining the feature of the target detection area based on the features of the multiple first voxels; and determining the detection result of the target object based on the features of the target detection area.

[0009] In some embodiments, performing voxel segmentation on the target detection region to obtain a plurality of first voxels includes: performing voxel segmentation on the target detection region at multiple scales to obtain multiple groups of first voxels corresponding to the multiple scales.

[0010] In some embodiments, the first voxels in each group of first voxels are the same size.

[0011] In some embodiments, the multiple scales include a first scale and a second scale other than the first scale, the first scale is the smallest scale among the multiple scales, the features of each first voxel in each group of first voxels constitute a voxel feature matrix, the voxel feature matrix corresponding to the first scale is the first voxel feature matrix, and the voxel feature matrix corresponding to the second scale is the second voxel feature matrix; determining the features of the target detection area based on the features of the multiple first voxels includes: upsampling the second voxel feature matrix to obtain a third voxel feature matrix, the size of the third voxel feature matrix is ​​the same as the size of the first voxel feature matrix; determining a fourth voxel feature matrix based on the third voxel feature matrix and the first voxel feature matrix, the elements of any position in the fourth voxel feature matrix include the elements of any position in the third voxel feature matrix and the elements of any position in the first voxel feature matrix; determining the features of the target detection area based on the elements in the fourth voxel feature matrix.

[0012] In some embodiments, the fourth voxel feature matrix includes n matrices, each matrix is ​​an n×n matrix, and the n matrices are arranged in order from 1 to n in the first direction. According to the elements in the fourth voxel feature matrix, determining the characteristics of the target detection area includes: arranging the elements of each matrix row by row or column by column to obtain a first sequence of each matrix; arranging n first sequences of the n matrices in order from 1 to n or from n to 1 to obtain a one-dimensional feature matrix as the characteristic of the target detection area.

[0013] In some embodiments, determining the third point feature of each key point based on the second point feature and the second BEV feature includes: determining the fourth point feature of each key point based on the second point feature and the second BEV feature; determining the probability value that the fourth point feature of each key point belongs to the foreground feature; determining the third point feature of the key point based on the fourth point feature of each key point and the probability value that the fourth point feature of the key point belongs to the foreground feature.

[0014] In some embodiments, the method further includes: dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding one to one to the multiple scales, the voxel feature matrix corresponding to each scale being composed of the initial point features of each point cloud point belonging to a group of second voxels corresponding to the scale; inputting the multiple voxel feature matrices corresponding to the multiple scales into the attention network to obtain an attention feature matrix, the elements at any position in the attention feature matrix are obtained by attention encoding the elements at any position in each voxel feature matrix in the multiple voxel feature matrices; wherein, the first point feature is determined based on the elements in the attention feature matrix.

[0015] In some embodiments, the method further includes: determining a coding feature matrix based on the attention feature matrix and the multiple voxel feature matrices; wherein the first point feature is determined based on elements in the coding feature matrix.

[0016] In some embodiments, the element at any position in the encoding feature matrix includes the element at any position in each voxel feature matrix in the multiple voxel feature matrices and the element at any position in the attention feature matrix.

[0017] In some embodiments, the method further includes: dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding one to one to the multiple scales, the voxel feature matrix corresponding to each scale being composed of point features of each point cloud point belonging to a group of second voxels corresponding to the scale; inputting the multiple voxel feature matrices corresponding to the multiple scales into the attention network to obtain an attention feature matrix, the elements at any position in the attention feature matrix being obtained by attention encoding the elements at any position in each voxel feature matrix in the multiple voxel feature matrices; performing one-stage target detection based on the attention feature matrix to determine a target detection area of ​​the point cloud including the target object from the original point cloud; wherein the key point set is obtained by sampling the target detection area.

[0018] In some embodiments, performing a one-stage target detection based on the attention feature matrix to determine a target detection area of ​​a point cloud including the target object from the original point cloud includes: determining a coding feature matrix based on the attention feature matrix and the multiple voxel feature matrices; performing a one-stage target detection based on the coding feature matrix to determine a target detection area of ​​a point cloud including the target object from the original point cloud.

[0019] In some embodiments, the method further includes: dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding one-to-one to the multiple scales, wherein the voxel feature matrix corresponding to each scale is composed of point features of each point cloud point in a group of second voxels corresponding to the scale; using a feature pyramid network to process the multiple voxel feature matrices corresponding to the multiple scales to obtain multiple BEV feature matrices corresponding to the multiple scales; and determining the first BEV feature based on the multiple BEV feature matrices.

[0020] In some embodiments, the method further includes: determining the four pixel points that are closest to the key pixel points corresponding to each key point in the key point set at a bird's-eye view, and the key pixel points are located within a rectangle formed by the four pixel points; wherein, determining the second BEV feature of each key point in the key point set based on the first BEV feature of some point cloud points in the original point cloud at a bird's-eye view includes: determining the second BEV feature of each key point in the key point set based on the first BEV feature of the four pixel points.

[0021] According to another aspect of an embodiment of the present disclosure, there is provided a target detection apparatus, comprising: a module configured to execute the method described in any one of the above embodiments.

[0022] According to another aspect of an embodiment of the present disclosure, a target detection device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the method described in any one of the above embodiments based on instructions stored in the memory.

[0023] According to another aspect of the embodiments of the present disclosure, a vehicle is provided, comprising: the target detection device described in any one of the above embodiments.

[0024] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, comprising computer program instructions, wherein when the computer program instructions are executed by a processor, the method described in any one of the above embodiments is implemented.

[0025] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method described in any one of the above embodiments is implemented.

[0026] In the embodiment of the present disclosure, the original point cloud is sampled to obtain a set of key points of the target object, the second point feature of each key point is determined based on the first point feature of some point cloud points in the original point cloud, the second BEV feature of each key point is determined based on the first BEV feature of some point cloud points in the original point cloud at a bird's-eye view, the third point feature of each key point is determined based on the second feature point feature and the second BEV feature of each key point, and the detection result of the target object is determined based on the third point feature of each key point. In this way, since the features (i.e., the third point features) of each key point used to detect the target object are determined based on the point features and BEV features of some point cloud points, the features of each key point not only incorporate the information contained in the point features of the point cloud points (such as position and geometric structure information), but also incorporate the information contained in the BEV features (such as context information), thereby improving the richness and diversity of the features of each key point, thereby improving the accuracy of the target object detection results obtained by performing target detection based on the features of each key point.

[0027] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] Figure 1 It is a flowchart of a target detection method according to some embodiments of the present disclosure.

[0030] Figure 2 It is a schematic diagram of the matrix composed of the first point features of some point cloud points.

[0031] Figure 3 It is a schematic diagram of different sampling methods for obtaining the key point set of the target object.

[0032] Figure 4 is a two-dimensional schematic diagram of a plurality of first voxels according to some embodiments of the present disclosure.

[0033] Figure 5 4 is a flowchart of some implementation methods of step S2.

[0034] Figure 6 108 is a flowchart of some implementation methods.

[0035] Figure 7Schematic diagram of a flow chart of a target detection method according to some other embodiments of the present disclosure.

[0036] Figure 8 and Figure 9 It is a schematic diagram of a dynamic voxelization method at a certain scale according to some embodiments of the present disclosure.

[0037] Figure 10 Schematic diagram of a method for obtaining a coding feature matrix according to some embodiments of the present disclosure.

[0038] Figure 11 It is a flowchart of a target detection method according to some other embodiments of the present disclosure.

[0039] Figure 12 is a schematic structural diagram of a target detection device according to some embodiments of the present disclosure.

[0040] Figure 13 2 is a schematic structural diagram of a target detection device according to other embodiments of the present disclosure. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0042] Unless specifically stated otherwise, the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0043] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0044] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0045] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0046] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0047] Figure 1It is a flowchart of a target detection method according to some embodiments of the present disclosure.

[0048] In step 102 , the original point cloud is sampled to obtain a set of key points of the target object.

[0049] Here, the original point cloud includes the point cloud of the target object. It should be understood that the key point set of the target object includes key point cloud points (i.e., key points) in the point cloud of the target object. The target object is, for example, an object around a vehicle. Figure 1 The object detection method shown can be used, for example, to detect target objects around a vehicle.

[0050] In some embodiments, the original point cloud may further include point clouds of other objects in addition to the point cloud of the target object.

[0051] In some embodiments, the target object may include one or more objects. In the case where the target object includes multiple objects, multiple key point sets of the multiple objects may be obtained after sampling the original point cloud, and target detection may be performed based on the multiple key point sets to obtain detection results for the multiple objects.

[0052] In some embodiments, a farthest point sampling algorithm may be used to sample the original point cloud to obtain a set of key points of the target object.

[0053] In step 104 , a second point feature of each key point in the key point set is determined based on the first point features of some point cloud points in the original point cloud.

[0054] In some embodiments, the initial point features of some of the point cloud points in the original point cloud can be used as the first point features. For example, the initial point features of each point cloud point include the position data and reflection intensity of the point cloud point (e.g., the reflection intensity of the laser beam emitted by a lidar after being reflected on the surface of an object of different materials). For example, the initial point feature of the i-th point cloud point in the original point cloud is (xi, yi, zi, ri), where (xi, yi, zi) represents the three-dimensional position coordinates of the i-th point cloud point, and ri represents the reflection intensity of the i-th point cloud point.

[0055] It should be noted that the initial point feature (xi, yi, zi, ri) of the i-th point cloud point includes four eigenvalues ​​(i.e., xi, yi, zi, and ri) ​​of four feature dimensions (also called channel dimensions), so the feature length of the initial point feature of the i-th point cloud point is 4.

[0056] As some implementation methods, a preset number of partial point cloud points can be collected within a spherical area with each key point as the center and r as the radius, and the feature composed of the maximum eigenvalue of the first point feature of this partial point cloud point in each feature dimension is determined as the second point feature of the key point. For example, (N, d) represents the matrix composed of the first point features of the partial point cloud points corresponding to each key point, where N represents the number of this partial point cloud points (i.e., the number of rows of the matrix), and d represents the length of the first point feature of this partial point cloud point (i.e., the number of columns of the matrix). Each column element in the matrix (N, d) includes N eigenvalues ​​of N point cloud points in a corresponding feature dimension, the maximum eigenvalue in each column element is determined, and the feature composed of the maximum eigenvalue in each column element is determined as the second point feature of the key point.

[0057] Figure 2 It is a schematic diagram of the matrix composed of the first point features of some point cloud points.

[0058] like Figure 2 As shown, Figure 2 Matrix 20 composed of the first-point features of a portion of the point cloud is schematically shown. Three point cloud points are collected within a spherical region centered at a key point and with a radius of r (i.e., the number of points in the portion of the point cloud is 3). The first-point features of these three point cloud points are (5, 4, 6, 2), (8, 9, 2, 4), and (1, 7, 3, 9), respectively. The length of the first-point feature of each point cloud point is 4 (i.e., the first-point feature of each point cloud point includes four eigenvalues ​​of four feature dimensions). Therefore, the matrix 20 composed of the first-point features of these three point cloud points is a 3×4 matrix, which can be represented by (3, 4).

[0059] Each column of elements in matrix 20 includes the three eigenvalues ​​of these three point cloud points under the same feature dimension. The maximum values ​​in each column of elements are 8, 9, 6, and 9 respectively. Therefore, it can be determined that the second point feature of the key point corresponding to these three point cloud points is (8, 9, 6, 9).

[0060] In step 106 , a second Bird's Eye View (BEV) feature of each key point in the key point set is determined based on the first BEV feature of some point cloud points in the original point cloud at a bird's eye view.

[0061] In some embodiments, the partial point cloud points used to determine the second point feature of each key point in step 104 and the partial point cloud points used to determine the second BEV feature of each key point in step 106 may be the same or different. For example, the two partial point cloud points may include the same partial point cloud points.

[0062] In some embodiments, the four pixel points closest to the key pixel point corresponding to each key point in the key point set from a bird's-eye view can be determined, and the key pixel point is located within the rectangle formed by these four pixel points. Then, the second BEV feature of the key point can be determined based on the first BEV feature of these four pixel points.

[0063] For example, the four pixels closest to the key pixel corresponding to each key point can be determined based on the coordinates of the pixels corresponding to each point cloud point in the original point cloud from a bird's-eye view. The four pixels corresponding to each key point form a rectangle, that is, any one of the four pixels has the same horizontal coordinate as one of the four pixels and the same horizontal coordinate as another of the four pixels. The distance between this arbitrary pixel and the remaining pixel of the four pixels other than this one and this other pixel is the farthest, and the key pixel corresponding to this key point is located within the rectangle formed by the corresponding four pixels. The first BEV features of the four pixels corresponding to each key point can be processed using a bilinear interpolation algorithm to determine the second BEV feature of this key point.

[0064] It should be understood that each key point in the key point set is a point cloud point in the original point cloud, and each point cloud point in the original point cloud corresponds to a pixel point from a bird's-eye view. The first BEV feature of each point cloud point from a bird's-eye view is the first BEV feature of the pixel point corresponding to the point cloud point.

[0065] In step 108 , a third point feature of each key point is determined based on the second point feature and the second BEV feature.

[0066] In some embodiments, the second point feature and the second BEV feature of each key point can be concatenated in the feature dimension (i.e., the channel dimension) to obtain the third point feature of the key point. For example, the second point feature of a key point is (1, 2, 3, 4) and the second BEV feature is (4, 5, 6, 7). After concatenating these two features, the third point feature of the key point can be obtained as (1, 2, 3, 4, 4, 5, 6, 7). The feature length of the third point feature of the key point is 8.

[0067] It should be noted that, unless otherwise specified, the "splicing" mentioned below refers to feature splicing in the feature dimension (i.e., channel dimension), that is, the length of the spliced ​​feature is the sum of the feature lengths of each spliced ​​feature.

[0068] In step 110 , a detection result of the target object is determined based on the third point feature of each key point.

[0069] In some embodiments, a first-stage target detection can be performed to determine the target detection area, and then the characteristics of the target detection area corresponding to the target object are determined based on the third-point characteristics of each key point, and a second-stage target detection is performed based on the characteristics of the target detection area to determine the detection result of the target object.

[0070] In some embodiments, the features of the target detection area corresponding to the target object can be determined based on the third point features of each key point in the key point set, and the features of the target detection area are input into the corresponding neural network model for detection to obtain the detection result of the target object. For example, the detection result of the target object can include the classification results of the foreground information and background information of the target object and the detection box prediction result. For example, after the features of the target detection area are input into the corresponding neural network model, a classification vector (i.e., a classification result) can be generated to represent the number of foreground points in the target detection area and the probability value that the features of the target detection area belong to the background features; and a detection box prediction vector (x, y, z, l, w, h, θ) with a length of 7 can be generated (i.e., a detection box prediction result), where (x, y, z) represents the three-dimensional coordinates of the target object, (l, w, h) represents the three-dimensional length of the target object, and θ represents the orientation of the target object. The classification task in the neural network model can be achieved by applying a focal loss constraint, and the detection box prediction task in the neural network model can be achieved by applying a smooth-l1-loss constraint.

[0071] In the above embodiment, the original point cloud is sampled to obtain a set of key points of the target object, the second point feature of each key point is determined based on the first point feature of some point cloud points in the original point cloud, the second BEV feature of each key point is determined based on the first BEV feature of some point cloud points in the original point cloud from a bird's-eye view, the third point feature of each key point is determined based on the second feature point feature and the second BEV feature of each key point, and the detection result of the target object is determined based on the third point feature of each key point. In this way, since the features of each key point used to detect the target object (i.e., the third point feature) are determined based on the point features and BEV features of some point cloud points, the features of each key point not only incorporate the information contained in the point features of the point cloud points (such as position and geometric structure information), but also incorporate the information contained in the BEV features (such as context information), thereby improving the richness and diversity of the features of each key point, thereby improving the accuracy of the target object detection results obtained by performing target detection based on the features of each key point.

[0072] In some embodiments, a target detection region of a point cloud containing a target object can be determined from the original point cloud. For example, a target detection region corresponding to the target object can be determined by performing a first-stage target detection on the original point cloud. This will be further described below.

[0073] In these embodiments, as a first sampling method, the key point set of the target object can be obtained by sampling the original point cloud, that is, by sampling the target detection area and other areas except the target detection area.

[0074] As a second sampling method, the key point set of the target object can be obtained by sampling only the target detection area. For example, the farthest point sampling algorithm is used to sample only the target detection area to obtain the key point set of the target object.

[0075] It should be understood that sampling only the target detection area means sampling only the point cloud of the target object within the target detection area.

[0076] Figure 3 It is a schematic diagram of different sampling methods for obtaining the key point set of the target object.

[0077] like Figure 3 As shown in sub-figures (1) to (3) in the figure, the solid rectangular box a represents the target object a, the solid rectangular box b represents the target object b, the dotted circular box A represents the target detection area corresponding to the target object a, and the dotted circular box B represents the target detection area corresponding to the target object b.

[0078] After performing one-stage target detection on the original point cloud, the target detection area A corresponding to the target object a, the target detection area B corresponding to the target object b, and other areas C except these two target detection areas can be determined, such as Figure 3 As shown in sub-figure (1) in .

[0079] In the first sampling mode, Figure 3 As shown in sub-figure (2) in , the key point set obtained after sampling includes the key points in the target detection areas A and B, as well as the key points in other areas C. In the second sampling method, Figure 3 As shown in sub-figure (3) in , the key point set obtained after sampling only includes the key points in the target detection areas A and B.

[0080] It should be noted that the points in the point cloud of target objects a and b are foreground points, and the other points (i.e., points outside the solid rectangular boxes a and b) are background points. In target detection, the richness of the foreground point information is positively correlated with the accuracy of the detection results obtained.

[0081] Since the second sampling method only samples the point clouds in the target detection areas A and B, the sampling of background points (such as the point clouds in other areas C) is reduced compared to the first sampling method, and more sufficient sampling of foreground points is achieved. Therefore, the second sampling method can ensure that most of the points in the key point set obtained by sampling for subsequent target detection are foreground points, thereby increasing the proportion of foreground points in the key point set used for subsequent target detection, thereby improving the accuracy of the subsequent detection results.

[0082] In some embodiments, the target detection region may be voxel-segmented to obtain a plurality of first voxels. It should be understood that voxel-segmenting the target detection region means dividing the point cloud points in the target detection region into a plurality of first voxels.

[0083] In these embodiments, as an implementation of step 110 , the detection result of the target object may be determined according to the following steps S1 to S3 .

[0084] S1: Determine a feature of each first voxel based on the third point feature of each key point in the plurality of first voxels.

[0085] In some embodiments, a preset number of key points are collected within a spherical area with the voxel center of each first voxel as the sphere center and r as the radius, and the feature composed of the maximum eigenvalue of the third point feature of these key points in each feature dimension is determined as the feature of the first voxel. This method is similar to Figure 2 The method shown is similar, for details, please refer to Figure 2 Description of the illustrated embodiment.

[0086] It should be noted that if the number of key points that can be collected in a first voxel is less than the preset number, the features of the preset number of points can be added by padding. For example, if the preset number is 10 and the feature length of the key point is 4, if there is no key point in a first voxel or the number of key points is less than 10, a number of point features (x, y, z, 0) can be randomly added to bring the number of point features to 10, where (x, y, z) represents the coordinates of a point at any position in the first voxel.

[0087] Figure 4 is a two-dimensional schematic diagram of a plurality of first voxels according to some embodiments of the present disclosure.

[0088] like Figure 4 As shown in subgraph (1) in Figure 4 The sub-figure (1) in FIG shows schematically the schematic diagram of the 27 first voxels obtained after the target detection area is segmented into 3×3×3 voxels on the two-dimensional plane. Figure 4As shown in the sub-figure (2) in the figure, in the spherical area ( Figure 4 A preset number of key points are collected within the sub-figure (the circular area R in the sub-figure (2) is a two-dimensional schematic diagram of this spherical area) to determine the feature of the first voxel based on the third point features of the preset number of key points (if the number is less than the preset number, the value is supplemented).

[0089] S2: Determine features of the target detection area according to features of the plurality of first voxels.

[0090] In some embodiments, features of multiple first voxels may be spliced ​​together to obtain features of the target detection area.

[0091] S3: Determine the detection result of the target object based on the characteristics of the target detection area.

[0092] In some embodiments, the features of the target detection area can be input into a corresponding neural network model for detection to obtain a detection result of the target object. The specific description of step S3 can be found in the description of the relevant embodiment of the aforementioned step 110, which will not be repeated here.

[0093] In the above embodiment, the target detection region is segmented into voxels, and the characteristics of each first voxel are determined based on the third point features of each key point in the plurality of first voxels. Furthermore, the characteristics of the target detection region are determined based on the characteristics of the plurality of first voxels, so that the detection result of the target object is determined based on the characteristics of the target detection region. In this way, because the characteristics of each first voxel can accurately represent the characteristics of the local region corresponding to the first voxel, using the characteristics of each first voxel as an intermediate feature in determining the characteristics of the target detection region can improve the accuracy of the determined characteristics of the target detection region, thereby further improving the accuracy of the obtained target object detection result.

[0094] As some implementations of step S1, the target detection area can be subjected to voxel segmentation at multiple scales to obtain multiple groups of first voxels corresponding to each other at the multiple scales. It should be understood that the different groups of first voxels obtained after voxel segmentation at multiple different scales have different resolutions of voxel maps from a bird's-eye view, wherein the scale is inversely correlated with the resolution (i.e., the larger the scale, the lower the resolution). The number of first voxels in each group of first voxels is multiple, and the features of the multiple first voxels used to determine the features of the target detection area include the features of all first voxels in the multiple groups of first voxels.

[0095] In this way, the features of multiple first voxels used to determine the characteristics of the target detection area are integrated with the feature information of key points extracted at different scales. Compared with feature extraction and fusion at only a single scale, the richness of information used to determine the characteristics of the target detection area is improved, and the granularity of the information contained in the characteristics of the determined target detection area is smaller (that is, the information contained is more comprehensive and sufficient), thereby further improving the accuracy of the detection results obtained by using the characteristics of the target detection area for target detection.

[0096] In some embodiments, the first voxels in each group of first voxels are of the same size. For example, each scale of voxel segmentation divides the target detection region into n equal parts in the three-dimensional direction (i.e., n×n×n equidistant voxel segmentation) to obtain n×n×n first voxels. The value of n corresponding to different scales of voxel segmentation is different, wherein the larger the scale, the smaller the value of n. The first voxels in a group of first voxels obtained under the voxel segmentation of the same scale are all of the same size, while the first voxels in different groups of first voxels obtained under voxel segmentation of different scales are different from each other. For example, three groups of first voxels obtained after performing three different scales of voxel segmentation on the target detection region may include 2×2×2 first voxels, 3×3×3 first voxels, and 6×6×6 first voxels, respectively.

[0097] In this way, through equidistant voxel segmentation at each scale, the accuracy of the intermediate features (i.e., the features of multiple first voxels) used to determine the features of the target detection area is further improved while ensuring the richness of the intermediate features, thereby further improving the accuracy of the detection results obtained by using the features of the target detection area for target detection.

[0098] In some embodiments, the multiple scales include a first scale and a second scale other than the first scale, wherein the first scale is the smallest scale among the multiple scales, and the features of each first voxel in each group of first voxels constitute a voxel feature matrix. The voxel feature matrix corresponding to the first scale is the first voxel feature matrix, and the voxel feature matrix corresponding to the second scale is the second voxel feature matrix. For example, if a group of first voxels includes 3×3×3 first voxels, then these 3×3×3 first voxels constitute a 3×3×3 voxel feature matrix (i.e., the voxel feature matrix includes three 3×3 two-dimensional matrices), and the feature of any first voxel in these 3×3×3 first voxels is an element in the voxel feature matrix.

[0099] In these embodiments, as some implementations of step S2, the following can be used: Figure 5 The steps shown determine the characteristics of the target detection area. Figure 5 4 is a flowchart of some implementation methods of step S2.

[0100] In step 502 , the second voxel feature matrix is ​​upsampled to obtain a third voxel feature matrix.

[0101] Here, the size of the third voxel feature matrix is ​​the same as that of the first voxel feature matrix.

[0102] In some embodiments, each scale other than the first scale among the multiple scales is a second scale. The second voxel feature matrix corresponding to each scale other than the first scale is upsampled to obtain a third voxel feature matrix corresponding to the scale having the same size as the first voxel feature matrix.

[0103] In some embodiments, the second voxel feature matrix can be upsampled by linear interpolation to obtain a third voxel feature matrix. For example, (n, n, n, d) represents the voxel feature matrix corresponding to each scale, (n, n, n) represents n×n×n equidistant voxel segmentation (i.e., dividing the target detection area into n equal parts in the three-dimensional direction), the value of n corresponding to different scales is different, and d represents the characteristic length of the first voxel. The three groups of first voxels obtained after voxel segmentation of the target detection area at three different scales can include 2×2×2 first voxels, 3×3×3 first voxels and 6×6×6 first voxels respectively. Since the voxel segmentation scale corresponding to 6×6×6 first voxels is the smallest, the second voxel feature matrix (2, 2, 2, d) composed of 2×2×2 first voxels and the second voxel feature matrix (3, 3, 3, d) composed of 3×3×3 first voxels can be upsampled to obtain two third voxel feature matrices of 6×6×6 (6, 6, 6, d).

[0104] In step 504 , a fourth voxel feature matrix is ​​determined based on the third voxel feature matrix and the first voxel feature matrix.

[0105] Here, the elements at any position in the fourth voxel feature matrix include the elements at any position in the third voxel feature matrix and the elements at any position in the first voxel feature matrix.

[0106] In some embodiments, the third voxel feature matrix and the first voxel feature matrix may be concatenated to obtain a fourth voxel feature matrix.

[0107] For example, the elements at any position (e.g., layer 1, row 1, column 1) in the fourth voxel feature matrix may include the elements at that position (layer 1, row 1, column 1) in the third voxel feature matrix and the elements at that position (layer 1, row 1, column 1) in the first voxel feature matrix. That is, the characteristic length of the elements at each position in the fourth voxel feature matrix is ​​the sum of the characteristic lengths of the elements at that position in the third voxel feature matrix and the characteristic lengths of the elements at that position in the first voxel feature matrix. For example, if the characteristic length of the elements at that position in the third voxel feature matrix is ​​5 and the characteristic length of the elements at that position in the first voxel feature matrix is ​​5, then the characteristic length of the elements at that position in the fourth voxel feature matrix is ​​10.

[0108] In step 506 , the features of the target detection area are determined based on the elements in the fourth voxel feature matrix.

[0109] In some embodiments, the elements in the fourth voxel feature matrix may be arranged in a preset arrangement to obtain a one-dimensional feature matrix serving as the features of the target detection area.

[0110] In the above embodiment, the voxel feature matrix corresponding to another scale other than the voxel feature matrix corresponding to the minimum scale is upsampled so that the size of the voxel feature matrix obtained after upsampling is the same as the size of the voxel feature matrix corresponding to the minimum scale. The voxel feature matrix used to determine the features of the target detection area is determined based on the voxel feature matrix obtained after upsampling and the voxel feature matrix corresponding to the minimum scale. In this way, on the basis of retaining the features in the voxel feature matrix under voxel segmentation at another scale, the features in the voxel feature matrix under voxel segmentation at another scale are further refined by upsampling according to the size of the voxel feature matrix corresponding to the minimum scale, so that the granularity of the information contained in the features of the target detection area determined based on these refined voxel feature matrices is smaller (i.e., the information contained is more comprehensive and sufficient), further improving the accuracy of the features of the target detection area, and thus further improving the accuracy of the subsequent detection results.

[0111] In some embodiments, the fourth voxel feature matrix includes n matrices, each matrix is ​​an n×n matrix, and the n matrices are arranged in order from 1 to n in a first direction. For example, the first direction can be any direction in three dimensions.

[0112] As some implementations, the elements of each of the n matrices can be arranged row by row or column by column to obtain a first sequence for each matrix, and the n first sequences of the n matrices can be arranged in order from 1 to n or from n to 1 to obtain a one-dimensional feature matrix that serves as the feature of the target detection area. For example, if n is 3, the three positions in the first row of each matrix are recorded as 1, 2, and 3, the three positions in the second row are recorded as 4, 5, and 6, and the three positions in the third row are recorded as 7, 8, and 9. The 3×3 elements of each of the three matrices can be arranged row by row to obtain a first sequence for each matrix, for example, the 9 elements in the first sequence of each matrix are arranged in the positional order of 123-456-789; or the 3×3 elements of each of the three matrices can be arranged column by column to obtain a first sequence for each matrix, for example, the 9 elements in the first sequence of each matrix are arranged in the positional order of 147-258-369. Afterwards, since these three matrices are arranged in order from 1 to 3 in a certain direction, the three first sequences of these three matrices can be arranged in order from the first matrix to the third matrix or from the third matrix to the first matrix to obtain a one-dimensional feature matrix as a feature of the target detection area.

[0113] In this way, each element in the determined features of the target detection area is arranged in the order of its position in the corresponding voxel feature matrix, which further improves the accuracy of the features of the target detection area, thereby further improving the accuracy of the subsequent detection results.

[0114] Figure 6 108 is a flowchart of some implementation methods.

[0115] In step 602 , a fourth point feature of each key point is determined based on the second point feature and the second BEV feature.

[0116] In some embodiments, the concatenation of the second feature and the second BEV feature of each key point can be used as the fourth feature of the key point. For example, if the second feature of a key point is (1, 2, 3, 4) and the second BEV feature is (4, 5, 6, 7), the concatenation of these two features yields the fourth feature of the key point (1, 2, 3, 4, 4, 5, 6, 7).

[0117] In step 604 , the probability value of the fourth feature of each key point belonging to the foreground feature is determined.

[0118] In some embodiments, the fourth point feature of each key point can be input into a neural network model with foreground and background feature classification constraints to determine the probability value of the fourth point feature of each key point belonging to the foreground feature (i.e., the probability value of each key point belonging to the foreground point).

[0119] In step 606 , the third feature of each key point is determined based on the fourth feature of the key point and the probability value that the fourth feature of the key point belongs to the foreground feature.

[0120] In some embodiments, the probability value of each key point's fourth feature belonging to a foreground feature can be used as a weight for each feature value in the key point's fourth feature, and each feature value in the key point's fourth feature can be multiplied by the weight to determine the key point's third feature. For example, if the fourth feature of a key point is (1, 2, 3, 4, 4, 5, 6, 7), and the probability value of the key point's fourth feature belonging to a foreground feature is 0.7, then the key point's third feature can be determined to be (0.7, 1.4, 2.1, 2.8, 2.8, 3.5, 4.2).

[0121] In this way, by determining the probability value of the feature after the fusion of the second-point feature and the second BEV feature of each key point (i.e., the fourth-point feature) belonging to the foreground feature, and determining the third-point feature of each key point based on the fourth-point feature of each key point and the corresponding probability value, it is possible to strengthen the learning of the foreground information in the third-point feature of each key point during the target detection process based on the third-point feature of each key point, thereby effectively improving the accuracy of the detection results obtained by subsequent target detection.

[0122] Figure 7 Schematic diagram of a flow chart of a target detection method according to some other embodiments of the present disclosure.

[0123] and Figure 1 Compared with the embodiment shown, Figure 7 The method shown further includes steps 702 to 706 .

[0124] In step 702 , dynamic voxelization at multiple scales is performed on the original point cloud to obtain multiple groups of second voxels in one-to-one correspondence at the multiple scales.

[0125] Here, the voxel feature matrix corresponding to each scale is composed of point features of each point cloud point belonging to a set of second voxels corresponding to the scale.

[0126] The following combination Figure 8 and Figure 9 The method of dynamic voxelization at a certain scale is described. Figure 8 and Figure 9 is a schematic diagram of a dynamic voxelization method at a certain scale according to some embodiments of the present disclosure, wherein: Figure 9 Includes subgraphs (1) to (3).

[0127] For example, the initial point feature of the i-th point cloud point in the original point cloud is (xi, yi, zi, ri), where (xi, yi, zi) represents the three-dimensional position coordinates of the i-th point cloud point, and ri represents the reflection intensity of the i-th point cloud point.

[0128] According to the three-dimensional position coordinates in the initial point features of each point cloud point in the original point cloud, the three-dimensional space range corresponding to the original point cloud can be determined, and the minimum three-dimensional position coordinates (xmin, ymin, zmin) and maximum three-dimensional position coordinates (xmax, ymax, zmax) in this three-dimensional space range can be determined, where xmin and xmax respectively represent the minimum and maximum values ​​in the x-axis direction in this three-dimensional space range; ymin and ymax represent the minimum and maximum values ​​in the y-axis direction in this three-dimensional space range, and zmin and zmax represent the minimum and maximum values ​​in the z-axis direction in this three-dimensional space range.

[0129] A cube can be determined from the three-dimensional space according to the minimum three-dimensional position coordinates and the maximum three-dimensional position coordinates. For example, the cube determined according to the minimum three-dimensional position coordinates and the maximum three-dimensional position coordinates can be as follows: Figure 8 As shown, the point indicated by 801 is the spatial point corresponding to the minimum three-dimensional position coordinate, and the point indicated by 802 is the spatial point corresponding to the maximum three-dimensional position coordinate.

[0130] Afterwards, the original point cloud is divided into two-dimensional voxels from a bird's-eye view according to the minimum three-dimensional position coordinates, the maximum three-dimensional position coordinates and the preset voxel size (vl, vw), and the resolution of the resulting voxel map is ((xmax-xmin) / vl, (ymax-ymin) / vw).

[0131] It should be understood that different set voxel sizes result in different scales of dynamic voxelization, that is, dynamic voxelization of the original point cloud at multiple different scales means dividing the original point cloud into two-dimensional voxels from a bird's-eye view according to multiple different voxel sizes, so that the resolution of the obtained voxel map is different.

[0132] For example, the result of 2D voxel segmentation of the original point cloud from a bird's eye view is as follows: Figure 9 shown. Figure 9 Sub-figure (1) in the figure schematically shows four second voxels (i.e., a group of second voxels) obtained after two-dimensional voxel division of the original point cloud from a bird's-eye view, namely, second voxel 901, second voxel 902, second voxel 903, and second voxel 904. The voxel diagram of these four second voxels from a bird's-eye view is shown in FIG. Figure 9 As shown in subgraph (2) in Figure 9In the sub-graph (2), the length of the voxel map is W = ymax-ymin, and the width is L = xmax-xmin. The length and width of each voxel (V1, V2, V3 and V4) in the bird's-eye view are respectively preset vl and vw. Figure 9 Sub-figure (3) in illustratively shows the length and width of the voxel V1 corresponding to the second voxel 901 .

[0133] After the original point cloud is dynamically voxelized, each point in the original point cloud will be divided into corresponding voxels, and the number of point cloud points contained in each second voxel can be different. Figure 9 The number of point cloud points in each of the four second voxels obtained under the voxel division shown is different, and the point features of all the point cloud points in these four second voxels can form a corresponding voxel feature matrix.

[0134] Based on the position of each point cloud point in the corresponding voxel, the voxel index of each point cloud point can be determined, so that the feature of the point cloud point can be found and encoded based on the voxel index. For example, the voxel index of the i-th point cloud point Ii = [(xi-xmin) / vl]×(W / vw)+(yi-ymin) / vw, where xi and yi are the coordinate data of the i-th point cloud point from a bird's-eye view.

[0135] It should be understood that the dynamic voxelization at each different scale can refer to the above Figures 8 and 9 Proceed in the manner shown.

[0136] In step 704, a plurality of voxel feature matrices corresponding to a plurality of scales are input into the attention network to obtain an attention feature matrix.

[0137] Here, an element at any position in the attention feature matrix is ​​obtained by attention encoding an element obtained after attention encoding an element at any position in each voxel feature matrix in multiple voxel feature matrices.

[0138] In some embodiments, the attention network may include multiple attention sub-networks. After inputting multiple voxel feature matrices corresponding to multiple scales into the multiple attention sub-networks one by one, attention encoding of each voxel feature matrix in the multiple voxel feature matrices can be achieved. Afterwards, the multiple voxel feature matrices after attention encoding output by the multiple attention sub-networks are spliced, and the spliced ​​multiple voxel feature matrices are attention encoded to obtain an attention feature matrix. Here, the elements of any position in the intermediate matrix obtained after splicing the multiple voxel feature matrices include the elements obtained after attention encoding the elements of any position (i.e., the same position) in each voxel feature matrix in the multiple voxel feature matrices, and the elements of any position in the attention feature matrix are all elements obtained by attention encoding the elements of any position (i.e., the same position) in the intermediate matrix. This will be further explained later.

[0139] In step 706 , a first-stage target detection is performed based on the attention feature matrix to determine a target detection region of the point cloud including the target object from the original point cloud.

[0140] here, Figure 7 Steps 702 to 706 may be performed before step 102 , and the key point set in step 102 may be obtained by sampling the target detection area determined in step 706 .

[0141] In some embodiments, a feature pyramid network may be used to process the attention feature matrix to implement one-stage target detection, thereby determining a target detection region of a point cloud including a target object from an original point cloud.

[0142] In the above embodiment, the original point cloud is dynamically voxelized at multiple scales to obtain multiple groups of second voxels. Multiple voxel feature matrices corresponding to the multiple scales composed of the multiple groups of second voxels are input into the attention network to obtain an attention feature matrix. Based on the attention feature matrix, first-stage target detection is performed to determine the target detection region of the point cloud including the target object from the original point cloud. The determined target detection region is then sampled to obtain a set of key points for subsequent target detection. In this way, since each element in the attention feature matrix used for first-stage target detection is not only attention-encoded in the corresponding voxel feature matrix, but also further attention-encoded based on the combination of the attention-encoded elements, each element in the attention feature matrix not only fuses the point feature information at each scale separately, but also further integrates the point feature information at multiple different scales. This improves the information richness of the features used for first-stage target detection, thereby improving the accuracy of first-stage target detection. Furthermore, since the accuracy of the target detection region determined by the first-stage target detection is improved, the accuracy of subsequent target detection (e.g., second-stage target detection) based on the key point set obtained by sampling the target detection region is further improved.

[0143] In addition, it should be noted that since the dynamic voxelization method does not require limiting the number of point cloud points in each voxel, the number of point cloud points included in the original point cloud does not change before and after dynamic voxelization, avoiding the problem of losing some point cloud points in the original point cloud due to limiting the number of point cloud points in each voxel. The voxel feature matrix after voxelization more comprehensively retains all point information in the original point cloud, thereby helping to improve the accuracy of subsequent target detection. Moreover, dynamic voxelization of the original point cloud at different scales, compared to dynamic voxelization at only a single scale, can fully integrate the positional relationship between points in three-dimensional space when encoding the voxel feature matrix after voxelization, thereby helping to improve the richness of the encoded features.

[0144] In some embodiments, in step 706, a coding feature matrix can be first determined based on the attention feature matrix and multiple voxel feature matrices, and then a first-stage target detection can be performed based on the coding feature matrix to determine the target detection area of ​​the point cloud including the target object from the original point cloud. For example, a feature pyramid network can be used to process the coding feature matrix to achieve first-stage target detection. In this way, since the coding feature matrix used for first-stage target detection not only integrates the point feature information of the point cloud points after attention encoding at various scales, but also integrates the original point feature information of the point cloud points at various scales (i.e., the point feature information before attention encoding) at various scales, the feature information contained in each element in the coding feature matrix is ​​more sufficient, further improving the richness of information in the features used for first-stage target detection, thereby further improving the accuracy of first-stage target detection, and thus further improving the accuracy of subsequent target detection based on the key point set obtained by sampling the target detection area.

[0145] In some embodiments, the element at any position in the encoding feature matrix includes the element at that position in each of the multiple voxel feature matrices and the element at that position in the attention feature matrix. For example, the attention feature matrix and the multiple voxel feature matrices can be concatenated to obtain the encoding feature matrix, and the characteristic length of the element at any position in the encoding feature matrix can be the sum of the characteristic lengths of the elements at the same position in each of the multiple voxel feature matrices and the characteristic lengths of the elements at the same position in the attention feature matrix.

[0146] Figure 10 Schematic diagram of a method for obtaining a coding feature matrix according to some embodiments of the present disclosure.

[0147] like Figure 10 As shown, Figure 10 The three voxel feature matrices A, B and C corresponding to the three scales, as well as the three attention sub-networks a, b and c in the attention network are schematically shown. The attention sub-network a includes a first branch network and a second branch network, wherein the structures of the attention sub-networks b and c are the same as that of the attention sub-network a. The specific structures of the attention sub-networks b and c are not shown here.

[0148] The following describes the attention encoding process by taking the voxel feature matrix A as an example to obtain the voxel feature matrix after attention encoding after inputting it into the attention sub-network a.

[0149] After the voxel feature matrix A is input into the attention sub-network a, it is first processed by the first and second branch networks. In the first branch network, max pooling and average pooling operations are first performed on the voxel feature matrix A in the row direction (i.e., channel-dimensional operations). The pooled features are then passed through a fully connected layer for feature compression, activated by an activation function, and decompressed by another fully connected layer, completing the row-wise attention encoding of the voxel feature matrix A. In the second branch network, max pooling and average pooling operations are first performed on the voxel feature matrix A in the column direction (i.e., point-wise operations). The pooled features are then passed through a fully connected layer for feature compression, activated by an activation function, and decompressed by another fully connected layer, completing the column-wise attention encoding of the voxel feature matrix A. The row-wise attention encoding results output by the first branch network and the column-wise attention encoding results output by the second branch network are then concatenated to obtain the voxel feature matrix A' after the attention sub-network a performs attention encoding on the voxel feature matrix A.

[0150] It can be understood that the process of obtaining the attention-encoded voxel feature matrices B' and C' after the voxel feature matrices B and C are input into the attention sub-networks b and c respectively is similar to the process of obtaining the voxel feature matrix A' mentioned above, and will not be repeated here.

[0151] Next, the voxel feature matrices A', B', and C' are concatenated to obtain the concatenated intermediate matrix D. The intermediate matrix D is input into the fully connected layer in the attention network for attention encoding, and the matrix D' obtained after attention encoding is multiplied by the intermediate matrix D to obtain the attention feature matrix E output by the attention network.

[0152] Afterwards, the attention feature matrix E is concatenated with the original three voxel feature matrices A, B, and C that have not been encoded by attention to obtain the encoded feature matrix F.

[0153] In this way, the element at any position in the encoding feature matrix includes the element at the same position in each of the multiple voxel feature matrices and the element at the same position in the attention feature matrix. For example, if the feature length of each element in voxel feature matrices A, B, and C is 4, and the feature length of each element in the attention feature matrix is ​​12, then the feature length of each element in the encoding feature matrix is ​​24.

[0154] In this way, each element in the encoding feature matrix more fully retains the multi-scale original point feature information, further improving the information richness in the features used for one-stage target detection, thereby further improving the accuracy of one-stage target detection, and further improving the accuracy of subsequent target detection based on the key point set obtained by sampling the target detection area.

[0155] Different methods for obtaining the first point features of some point cloud points used to determine the second point features of each key point in step 702 are described below with reference to some embodiments.

[0156] In the first embodiment, step 702 and step 704 may be performed before step 104. The first point features of the partial point cloud points used to determine the second point features of each key point in step 702 may be determined based on the elements in the attention feature matrix obtained in step 704. In this embodiment, for example, a preset number of partial point cloud points may be collected within a spherical area with each key point as the center and r as the radius; the position of each point cloud point in the corresponding voxel feature matrix is ​​determined based on the voxel index of each point cloud point in this partial point cloud point at different scales, and then the position of each point cloud point in the attention feature matrix is ​​determined; for each point cloud point, the element located at the position of the point cloud point in the attention feature matrix is ​​determined as the first point feature of the point cloud point; and the feature composed of the maximum eigenvalue of the first point feature of this partial point cloud point in each feature dimension is determined as the second point feature of the key point.

[0157] In this way, since the first point features of some point cloud points used to determine the second point features of each key point are determined using elements in the attention feature matrix, the first point features integrate point feature information of multiple scales of point cloud points, thereby improving the richness of feature information in the second point features of each key point determined based on this first point feature, and thus improving the richness of feature information in the third point features of each key point determined subsequently, thereby further improving the accuracy of the detection results obtained by subsequent target detection based on the third point features of each key point.

[0158] In the second embodiment, steps 702 and 704 , as well as the following step S4 , may be performed before step 104 .

[0159] S4: Determine a coding feature matrix based on the attention feature matrix and the multiple voxel feature matrices. Detailed description of step S4 can be found in the description of the related embodiment in the aforementioned step 706, which will not be repeated here.

[0160] In this embodiment, the first point features of some of the point cloud points used to determine the second point features of each key point in the key point set in step 702 may be determined based on the elements in the encoding feature matrix in step S4.

[0161] In this way, since the first point features of some point cloud points used to determine the second point features of each key point are determined using elements in the encoding feature matrix, the first point features further integrate the original point feature information of various scales of the point cloud points, thereby further improving the richness of the feature information in the second point features of each key point determined based on this first point feature, and further improving the richness of the feature information in the third point features of each key point determined subsequently, thereby further improving the accuracy of the detection results obtained by subsequent target detection based on the third point features of each key point.

[0162] In a third embodiment, the element at any position in the encoding feature matrix determined in step S1 includes the element at that position in each of the multiple voxel feature matrices and the element at that position in the attention feature matrix. For example, the encoding feature matrix in step S1 can be obtained by concatenating the attention feature matrix and the multiple voxel feature matrices.

[0163] In this embodiment, for example, a preset number of partial point cloud points can be collected in a spherical area with each key point as the center and r as the radius; the position of each point cloud point in the corresponding voxel feature matrix is ​​determined according to the voxel index of each point cloud point in this part of the point cloud points at different scales, and then the position of each point cloud point in the encoding feature matrix is ​​determined; for each point cloud point, the element located at the position of the point cloud point in the encoding feature matrix is ​​determined as the first point feature of the point cloud point; the feature composed of the maximum eigenvalue of the first point feature of this part of the point cloud points in each feature dimension is determined as the second point feature of the key point.

[0164] In this way, the first point feature determined based on the elements in the encoding feature matrix more fully retains the multi-scale original point feature information, thereby further improving the richness of the feature information in the second point feature of each key point determined based on this first point feature, and further improving the richness of the feature information in the third point feature of each key point determined subsequently, thereby further improving the accuracy of the detection results obtained by subsequent target detection based on the third point feature of each key point.

[0165] Figure 11 It is a flowchart of a target detection method according to some other embodiments of the present disclosure.

[0166] and Figure 1 Compared with the embodiment shown, Figure 11The method shown further includes steps 1102 to 1104 .

[0167] In step 1102 , dynamic voxelization is performed on the original point cloud at multiple scales to obtain multiple groups of second voxels in one-to-one correspondence at the multiple scales.

[0168] Here, the voxel feature matrix corresponding to each scale is composed of point features of each point cloud point in a set of second voxels corresponding to the scale.

[0169] The implementation of step 1102 is similar to that of the aforementioned step 702. For detailed instructions, please refer to the description in the relevant embodiment of the aforementioned step 702, which will not be repeated here.

[0170] In step 1104 , a feature pyramid network is used to process the multiple voxel feature matrices corresponding to the multiple scales to obtain multiple BEV feature matrices corresponding to the multiple scales.

[0171] In some embodiments, the feature pyramid network may include multiple sub-networks (also called multiple encoding stages) corresponding one-to-one to multiple scales. Multiple voxel feature matrices corresponding to multiple scales are input one-to-one into the multiple sub-networks, and multiple BEV feature matrices output by the multiple sub-networks can be obtained.

[0172] It should be understood that the element at any position in each BEV feature matrix is ​​obtained by processing the element at any position in the corresponding voxel feature matrix using the corresponding sub-network.

[0173] At step 1106 , a first BEV feature is determined based on the plurality of BEV feature matrices.

[0174] In some embodiments, the position of each point cloud point in the corresponding voxel feature matrix can be determined based on the voxel index of each point cloud point, and further the position of each point cloud point in the corresponding BEV feature matrix can be determined. For each point cloud point, the element located at the position of the point cloud point in the corresponding BEV feature matrix can be determined as the first BEV feature of the point cloud point.

[0175] In the above embodiment, multiple groups of second voxels are obtained by dynamically voxelizing the original point cloud at multiple scales. Multiple voxel feature matrices corresponding to the multiple scales composed of the multiple groups of second voxels are then input into a feature pyramid network to obtain multiple BEV feature matrices corresponding to the multiple scales. Based on these multiple BEV feature matrices, a first BEV feature is determined for determining the second BEV feature of each key point. Thus, the first BEV feature is obtained by processing the elements of the voxel feature matrices at different scales via the feature pyramid network. Therefore, the second BEV feature of each key point determined based on this first BEV feature incorporates BEV feature information from multiple different scales of the point cloud points, thereby increasing the richness of the feature information in the second BEV feature of each key point determined based on this first BEV feature. This in turn increases the richness of the feature information in the subsequently determined third-point feature of each key point, thereby further improving the accuracy of the subsequent target detection results obtained based on the third-point feature of each key point.

[0176] In some embodiments, in addition to including multiple sub-networks corresponding to various scales, the feature pyramid network may also include multiple prediction head networks corresponding to the multiple sub-networks. The multiple BEV feature matrices obtained after processing by the multiple sub-networks can be transmitted one-to-one to the multiple prediction head networks to perform multi-scale object detection, thereby obtaining a first-stage object detection result.

[0177] In some embodiments, the second point feature of each key point can be determined according to the first point feature determined in any one of the first embodiment, the second embodiment, and the third embodiment, and the second point feature of each key point can be determined according to the first point feature determined in any one of the first embodiment, the second embodiment, and the third embodiment. Figure 11 The first BEV feature determined in the manner shown is used to determine the second BEV feature of each key point, and then based on the second point feature and the second BEV feature, the third BEV feature of each key point is determined.

[0178] In this way, since the second point feature integrates point feature information of multiple scales of point cloud points and the second BEV feature integrates BEV feature information of multiple scales of point cloud points, the third point feature of each key point not only integrates point feature information of multiple scales of point cloud points, but also integrates BEV feature information of multiple scales of point cloud points, further improving the richness and diversity of the features of the key points used for target detection, thereby further improving the accuracy of the detection results of the target objects.

[0179] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the device embodiments, since they are essentially similar to the method embodiments, their descriptions are relatively simple. For relevant parts, reference can be made to the descriptions of the method embodiments.

[0180] An embodiment of the present disclosure further provides a target detection device, comprising a module configured to execute the method of any one of the above embodiments.

[0181] Figure 12 is a schematic structural diagram of a target detection device according to some embodiments of the present disclosure.

[0182] like Figure 12 As shown, the target detection device 1200 includes a sampling module 1201 and a determination module 1202 .

[0183] The sampling module 1201 may be configured to sample the original point cloud to obtain a key point set of the target object. Here, the original point cloud includes the point cloud of the target object.

[0184] The determination module 1202 can be configured to determine the second point feature of each key point in the key point set based on the first point feature of some point cloud points in the original point cloud; determine the second BEV feature of each key point in the key point set based on the first BEV feature of some point cloud points in the original point cloud at a bird's-eye view; determine the third point feature of each key point based on the second point feature and the second BEV feature; and determine the detection result of the target object based on the third point feature of each key point.

[0185] In some embodiments, the object detection apparatus 1200 may further include other modules for performing other operations of any of the above embodiments.

[0186] Figure 13 2 is a schematic structural diagram of a target detection device according to other embodiments of the present disclosure.

[0187] like Figure 13 As shown, the target detection device 1300 includes a memory 1301 and a processor 1302 coupled to the memory 1301 . The processor 1302 is configured to execute the method of any one of the aforementioned embodiments based on instructions stored in the memory 1301 .

[0188] The memory 1301 may include, for example, a system memory, a fixed non-volatile storage medium, etc. The system memory may store, for example, an operating system, an application program, a boot loader, and other programs.

[0189] The target detection device 1300 may also include an input / output interface 1303, a network interface 1304, a storage interface 1305, and the like. These interfaces 1303, 1304, and 1305, as well as the memory 1301 and the processor 1302, may be connected, for example, via a bus 1306. The input / output interface 1303 provides a connection interface for input / output devices such as a display, mouse, keyboard, and touch screen. The network interface 1304 provides a connection interface for various networked devices. The storage interface 1305 provides a connection interface for external storage devices such as SD cards and USB flash drives.

[0190] The present disclosure also provides a vehicle, including the target detection device of any one of the above-mentioned embodiments (eg, target detection device 1200 / 1300). For example, the vehicle may be a commercial vehicle.

[0191] An embodiment of the present disclosure further provides a computer-readable storage medium, comprising computer program instructions, which implement the method of any one of the above embodiments when executed by a processor.

[0192] An embodiment of the present disclosure further provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, the method of any one of the above embodiments is implemented.

[0193] Thus far, various embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.

[0194] Those skilled in the art will appreciate that embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0195] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present disclosure. It should be understood that the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate the functions for implementing the functions specified in the flowchart. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0196] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0198] Although some specific embodiments of the present disclosure have been described in detail through examples, those skilled in the art will understand that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that the above embodiments may be modified or some technical features may be replaced with equivalents without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A target detection method, comprising: Sampling an original point cloud to obtain a key point set of a target object, wherein the original point cloud includes a point cloud of the target object; Determining a second point feature of each key point in the key point set according to first point features of some point cloud points in the original point cloud; Determining a second BEV feature of each key point in the key point set according to a first bird's-eye view BEV feature of some point cloud points in the original point cloud at a bird's-eye view perspective; determining a third point feature of each key point based on the second point feature and the second BEV feature; According to the third point feature of each key point, the detection result of the target object is determined. Wherein, determining the third point feature of each key point according to the second point feature and the second BEV feature includes: determining a fourth point feature of each key point based on the second point feature and the second BEV feature; Determine the probability value of the fourth feature of each key point belonging to the foreground feature; The third feature of each key point is determined according to the fourth feature of the key point and the probability value that the fourth feature of the key point belongs to the foreground feature.

2. The method according to claim 1, characterized in that Also includes: Determining a target detection area including a point cloud of the target object from the original point cloud; The key point set is obtained by sampling only the target detection area.

3. The method according to claim 2, characterized in that Also includes: Performing voxel segmentation on the target detection area to obtain a plurality of first voxels; Determining the detection result of the target object according to the third point feature of each key point includes: determining a feature of each first voxel based on a third point feature of each key point in the plurality of first voxels; determining a feature of the target detection area according to features of the plurality of first voxels; A detection result of the target object is determined according to the characteristics of the target detection area.

4. The method according to claim 3, characterized in that Performing voxel segmentation on the target detection area to obtain a plurality of first voxels includes: The target detection area is subjected to voxel segmentation at multiple scales to obtain multiple groups of first voxels corresponding to the multiple scales.

5. The method according to claim 4, characterized in that The first voxels in each group of first voxels have the same size.

6. The method according to claim 5, characterized in that The multiple scales include a first scale and a second scale other than the first scale, the first scale being the smallest scale among the multiple scales, the features of each first voxel in each group of first voxels forming a voxel feature matrix, the voxel feature matrix corresponding to the first scale being a first voxel feature matrix, and the voxel feature matrix corresponding to the second scale being a second voxel feature matrix; Determining the characteristics of the target detection area according to the characteristics of the plurality of first voxels includes: Upsampling the second voxel feature matrix to obtain a third voxel feature matrix, where the size of the third voxel feature matrix is ​​the same as the size of the first voxel feature matrix; determining a fourth voxel feature matrix according to the third voxel feature matrix and the first voxel feature matrix, wherein an element at any position in the fourth voxel feature matrix includes an element at any position in the third voxel feature matrix and an element at any position in the first voxel feature matrix; The characteristics of the target detection area are determined according to the elements in the fourth voxel feature matrix.

7. The method according to claim 6, characterized in that The fourth voxel feature matrix includes n matrices, each matrix is ​​an n×n matrix, and the n matrices are arranged in order from 1 to n in the first direction. Determining the features of the target detection area based on elements in the fourth voxel feature matrix includes: Arrange the elements of each matrix row by row or column by column to obtain a first sequence of each matrix; The n first sequences of the n matrices are arranged in order from 1 to n or from n to 1 to obtain a one-dimensional feature matrix serving as a feature of the target detection area.

8. The method according to any one of claims 1 to 7, characterized in that Also includes: Dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding to the multiple scales, wherein a voxel feature matrix corresponding to each scale is composed of initial point features of each point cloud point belonging to the group of second voxels corresponding to the scale; Inputting a plurality of voxel feature matrices corresponding to the plurality of scales into an attention network to obtain an attention feature matrix, wherein an element at any position in the attention feature matrix is ​​obtained by attention encoding an element obtained by attention encoding the element at the any position in each voxel feature matrix in the plurality of voxel feature matrices; The first point feature is determined based on the elements in the attention feature matrix.

9. The method according to claim 8, characterized in that Also includes: Determining a coding feature matrix according to the attention feature matrix and the plurality of voxel feature matrices; The first point feature is determined based on elements in the encoding feature matrix.

10. The method according to claim 9, characterized in that The elements at any position in the encoding feature matrix include the elements at any position in each voxel feature matrix in the multiple voxel feature matrices and the elements at any position in the attention feature matrix.

11. The method according to any one of claims 1 to 7, characterized in that: Also includes: Dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding to the multiple scales, wherein a voxel feature matrix corresponding to each scale is composed of point features of each point cloud point belonging to the group of second voxels corresponding to the scale; Inputting a plurality of voxel feature matrices corresponding to the plurality of scales into an attention network to obtain an attention feature matrix, wherein an element at any position in the attention feature matrix is ​​obtained by attention encoding an element obtained by attention encoding the element at the any position in each voxel feature matrix in the plurality of voxel feature matrices; Performing a first-stage target detection according to the attention feature matrix to determine a target detection area of ​​the point cloud including the target object from the original point cloud; The key point set is obtained by sampling the target detection area.

12. The method according to claim 11, characterized in that Performing a first-stage target detection according to the attention feature matrix to determine a target detection area of ​​the point cloud including the target object from the original point cloud includes: Determining a coding feature matrix according to the attention feature matrix and the plurality of voxel feature matrices; A first-stage target detection is performed according to the encoding feature matrix to determine a target detection area of ​​the point cloud including the target object from the original point cloud.

13. The method according to any one of claims 1 to 7, characterized in that: Also includes: Dynamically voxelizing the original point cloud at multiple scales to obtain multiple groups of second voxels corresponding to the multiple scales, wherein a voxel feature matrix corresponding to each scale is composed of point features of each point cloud point in the group of second voxels corresponding to the scale; Using a feature pyramid network to process the multiple voxel feature matrices corresponding to the multiple scales to obtain multiple BEV feature matrices corresponding to the multiple scales; A first BEV feature is determined based on the plurality of BEV feature matrices.

14. The method according to any one of claims 1 to 7, characterized in that: Also includes: Determine four pixels closest to a key pixel corresponding to each key point in the key point set from a bird's-eye view, where the key pixel is located within a rectangle formed by the four pixels; Wherein, determining the second BEV feature of each key point in the key point set according to the first BEV feature of some point cloud points in the original point cloud at a bird's-eye view includes: A second BEV feature of each key point in the key point set is determined according to the first BEV features of the four pixel points.

15. A target detection device comprising: A module configured to perform the method according to any one of claims 1 to 14.

16. A target detection device comprising: Memory; as well as A processor coupled to the memory is configured to execute the method according to any one of claims 1 to 14 based on instructions stored in the memory.

17. A vehicle comprising: The target detection device according to claim 15 or 16.

18. A computer-readable storage medium comprising computer program instructions, wherein: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 14 is implemented.

Citation Information

Patent Citations

  • Three-dimensional dynamic target detection method and device based on voxel point cloud fusion

    CN113989797A