An object detection method and related devices

By constructing a multi-scale grid point array in the object detection model for point cloud data sampling and feature processing, the inaccurate object detection problem caused by a single sampling scale in the prior art is solved, and the accuracy of object detection is improved.

CN113989188BActive Publication Date: 2025-07-18HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111131185.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-07-18
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

The existing object detection model has a single sampling scale when sampling point cloud data, which makes it impossible to fully characterize the initial area where the target object is located, thereby affecting the accuracy of the final area.

Method used

By constructing multiple grid point arrays in the object detection model, the grid point array has different sizes, performing multi-scale sampling, obtaining the point cloud data set around the grid point, and performing feature extraction and processing to obtain the final area of the target object.

Benefits of technology

The complete characterization of the area where the target object is located is achieved and the accuracy of object detection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989188B_ABST
    Figure CN113989188B_ABST
Patent Text Reader

Abstract

The present application provides an object detection method and related devices. In the second stage of object detection, the point cloud data sampled by the object detection model can completely represent the initial region where the target object is located, so that the final region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy. The method of the present application includes: processing the point cloud data of the target scene to obtain a first region where the target object is located in the target scene; constructing a plurality of grid point arrays in the first region, and different grid point arrays have different sizes; obtaining a set of point cloud data of the grid points in the plurality of grid point arrays, and the set of point cloud data includes the point cloud data around the grid points; processing the set of point cloud data to obtain a second region where the target object is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence (AI) technology, and in particular, to a method for determining a location based on a vehicle networking and related devices thereof. Background Art

[0002] Three-dimensional object detection is one of the important tasks in computer vision and has important applications in fields such as autonomous driving and industrial vision.

[0003] Currently, point cloud data of a target scene can be obtained through a lidar, so as to determine the area where a target object is located in the target scene. Specifically, after obtaining the point cloud data of the target scene, an object detection model can process this part of the point cloud data to predict the initial area where the target object is located. Then, the object detection model can sample the point cloud data within or near the initial area, and further process the sampled point cloud data to obtain the final area where the target object is located.

[0004] However, when the object detection model samples point cloud data, its sampling scale is usually single. Therefore, the point cloud data sampled by the model cannot completely represent the initial area where the target object is located, resulting in the final area where the target object is located obtained by the model based on these point cloud data being inaccurate enough. Summary of the Invention

[0005] Embodiments of this application provide an object detection method and related devices thereof. In the second stage of object detection, the point cloud data sampled by the object detection model can completely represent the initial area where the target object is located, so that the final area where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0006] The first aspect of the embodiments of this application provides an object detection method, and the method includes:

[0007] When a user needs to perform three-dimensional object detection on a target scene, the point cloud data of the target scene can be obtained first, and the point cloud data of the target scene is input into an object detection model, so that the object detection model processes the point cloud data of the target scene to obtain a first area where a target object is located in the target scene, that is, the initial area where the target object is located in the target scene.

[0008] Next, the object detection model can construct multiple arrays of grid points in the first region. Among these multiple arrays of grid points, different arrays of grid points have different sizes, where the size of some arrays of grid points can be larger than the size of the first region, and the size of some other arrays of grid points can be smaller than the size of the first region. Therefore, these multiple arrays of grid points can be used to sample the point cloud data within or near the first region at different scales.

[0009] Then, the object detection model can obtain the set of point cloud data for all grid points in the multiple arrays of grid points. The set of point cloud data for each grid point contains the point cloud data around that grid point.

[0010] Finally, the object detection model can process the set of point cloud data for all grid points to obtain the second region where the target object is located, that is, the final region where the target object is located in the target scene, which can be used as the object detection result of the target scene. Thus, the three-dimensional object detection of the target scene is completed, and the object detection result of the target scene can be fed back to the user for use.

[0011] As can be seen from the above method: after obtaining the point cloud data of the target scene, the object detection model can process the point cloud data of the target scene to obtain the first region where the target object is located in the target scene. Then, the object detection model can construct multiple arrays of grid points in the first region and obtain the set of point cloud data for all grid points in the multiple arrays of grid points. The set of point cloud data for each grid point contains the point cloud data around that grid point. Finally, the object detection model can process the set of point cloud data for all grid points to obtain the second region where the target object is located. In the foregoing process, since different arrays of grid points in the multiple arrays of grid points have different sizes, the size of some arrays of grid points can be larger than the size of the initial region, and the size of some other arrays of grid points can be smaller than the size of the initial region. Therefore, based on these multiple arrays of grid points, the object detection model can sample the point cloud data within or near the first region at multiple scales. In this way, the point cloud data sampled by the model can completely represent the first region where the target object is located, so that the second region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0012] In a possible implementation, constructing multiple grid point arrays in the first region includes: determining the positions of the target grid points based on the serial numbers of the target grid points in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, where the target grid point array is any one of the multiple grid point arrays, and the target grid point is any one of the grid points in the target grid point array; constructing the target grid point array in the first region based on the positions of all the grid points in the target grid point array; and repeating the above steps for the remaining grid point arrays in the multiple grid point arrays except the target grid point array until multiple grid point arrays are constructed in the first region. In the foregoing implementation, for any one of the multiple grid point arrays, the object detection model can determine the positions of all the grid points in the grid point array based on the serial numbers of all the grid points in the grid point array, the ratio between the size of the grid point array and the size of the first region, the grid point data of the grid point array, and the parameters of the first region, so as to construct the grid point array in the first region. In this way, multiple grid point arrays can be constructed in the first region. Since the ratios between the sizes of different grid point arrays and the size of the first region are different (i.e., different grid point arrays have different sizes), and the grid point data of different grid arrays can be the same or different, the multiple grid point arrays constructed in the first region can be used for the object detection model to perform different-scale sampling on the point cloud data within or near the first region.

[0013] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region. In the foregoing implementation, the first region where the target object obtained by the object detection model is located can be regarded as a 7-dimensional vector, and this vector includes 7 elements: the abscissa of the center of the first region, the ordinate of the center of the first region, the vertical coordinate of the center of the first region, the width of the first region, the height of the first region, the length of the first region, and the yaw angle of the first region.

[0014] In a possible implementation, obtaining the point cloud data set of the grid points in multiple grid point arrays includes: based on the distribution of the point cloud data in the first region, obtaining the sampling radius of the target grid point array, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of the multiple grid point arrays; based on the target grid point and the sampling radius of the target grid point array, determining the sampling range of the target grid point, where the target grid point is any grid point in the target grid point array; obtaining the point cloud data within the sampling range of the target grid point to obtain the point cloud data set of the target grid point; and repeating the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the point cloud data set of all grid points in the multiple grid point arrays. In the foregoing implementation, the object detection model can adjust the sampling radius of each grid point array according to the distribution of the point cloud data in the first region, thereby realizing the dynamic adjustment of the sampling range of the point cloud data, avoiding the situation where the object detection model collects invalid point cloud data, reducing the computational amount of the object detection model in the three-dimensional object detection process, and being beneficial to saving computing resources.

[0015] In a possible implementation, the sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere. In the foregoing implementation, with the target grid point as the center of the sphere, the probability of collecting the point cloud data within the radius of the sphere from the target grid point is generally very small. Therefore, the sampling range of the target grid point can be determined in this way, and the object detection model does not consider the point cloud data outside the sampling range, thereby reducing the computational amount of the object detection model.

[0016] In a possible implementation, among the multiple grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array. In the foregoing implementation, among the multiple grid point arrays, the sampling radius of the grid point array usually has a certain correlation with the size of the grid point array. For example, the sampling radius of the grid point array is positively correlated with the size of the grid point array, that is, the larger the size of the grid point array, the larger the sampling radius of the grid point array. For example, assume there are grid point array 1, grid point array 2, grid point array 3, and grid point array 4, and the size of the grid point arrays is sorted as: grid array 1, grid array 2, grid array 3, and grid array 4. Then, the sampling radius of the grid point arrays is sorted as: grid array 1, grid array 2, grid array 3, and grid array 4.

[0017] In a possible implementation, processing the point cloud data set to obtain the second region where the target object is located includes: performing a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any grid points in the target grid point array, and the target grid point array is any one of multiple grid point arrays; repeating the above steps for the remaining point cloud data in the point cloud data set of the target grid points except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; performing a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; repeating the above steps for the remaining grid points in the multiple grid point arrays except the target grid points to obtain the second features of all the grid points in the multiple grid point arrays; performing a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the target object is located. In the foregoing implementation, for any grid point, the object detection model can calculate the first features of all the point cloud data in the point cloud data set of the grid point, and perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the grid point to obtain the second feature of the grid point. In this way, the second features of all the grid points can be obtained, so the object detection model can accurately obtain the second region where the target object is located by further processing the second features of all the grid points.

[0018] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array. In the foregoing implementation, when the object detection model calculates the features of the target grid point, it can assign a certain weight to each point cloud data in the point cloud data set of the target grid point, taking into account the different effects of the point cloud data at different distances from the grid point on the target grid point. Therefore, the calculated features of the target grid point can contain more information, thereby further improving the accuracy of the second region where the target object is located.

[0019] In a possible implementation, performing a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data includes: performing a linear transformation process on the distance between the target point cloud data and the target grid points to obtain the third feature of the target point cloud data; performing a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; performing a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; and performing a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data. In the foregoing implementation, when the object detection model extracts the features of a certain point cloud data, it uses a processing based on the attention mechanism to replace the pooling process in the traditional method. Therefore, the features of the point cloud data obtained by the processing can include richer information such as the information of the point cloud data itself and the relationship between the point cloud data and the grid points. Then, based on these features of the point cloud data, it is beneficial to obtain better features of the grid points subsequently.

[0020] In a possible implementation, the fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

[0021] The second aspect of the embodiments of the present application provides a model training method, which includes: obtaining the point cloud data of the target scene and the real region where the object to be detected is located in the target scene; inputting the point cloud data of the target scene into the model to be trained to obtain the second region where the object to be detected is located, and the model to be trained is used for: processing the point cloud data of the target scene to obtain the first region where the object to be detected is located in the target scene; constructing a plurality of grid point arrays in the first region, and different grid point arrays have different sizes; obtaining the point cloud data set of the grid points in the plurality of grid point arrays, and the point cloud data set includes the point cloud data around the grid points; processing the point cloud data set to obtain the second region where the object to be detected is located; and training the model to be trained based on the real region and the second region to obtain an object detection model.

[0022] The object detection model obtained by the above method has the function of detecting the region where the target object is located in the target scene. During the three-dimensional object detection process of the object detection model, since different grid point arrays in the plurality of grid point arrays have different sizes, the sizes of some grid point arrays can be larger than the size of the initial region, and the sizes of some other grid point arrays can be smaller than the size of the initial region. Therefore, based on these plurality of grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data within or near the initial region. In this way, the point cloud data sampled by the model can completely represent the initial region where the target object is located, so that the final region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0023] In a possible implementation, the model to be trained is used to: determine the position of a target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, where the target grid point array is any one of multiple grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; and repeat the above steps for the remaining grid point arrays other than the target grid point array among the multiple grid point arrays until multiple grid point arrays are constructed in the first region.

[0024] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

[0025] In a possible implementation, the model to be trained is used to: obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of multiple grid point arrays; determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array, where the target grid point is any one of the grid points in the target grid point array; obtain the point cloud data within the sampling range of the target grid point to obtain the point cloud data set of the target grid point; and repeat the above steps for the remaining grid points other than the target grid point among the multiple grid point arrays to obtain the point cloud data sets of all grid points in the multiple grid point arrays.

[0026] In a possible implementation, the sampling range of the target grid point is a sphere with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0027] In a possible implementation, among the multiple grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array.

[0028] In a possible implementation, the model to be trained is used to: perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any grid points in the target grid point array, and the target grid point array is any one of multiple grid point arrays; repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid points except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the second features of all the grid points in the multiple grid point arrays; perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the object to be detected is located.

[0029] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0030] In a possible implementation, the model to be trained is used to: perform a linear transformation process on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; perform a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; perform a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; perform a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

[0031] In a possible implementation, the fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

[0032] A third aspect of the embodiments of the present application provides an object detection device, which includes: a first processing module for processing the point cloud data of the target scene to obtain the first region where the target object is located in the target scene; a construction module for constructing multiple grid point arrays in the first region, where different grid point arrays have different sizes; an acquisition module for acquiring the point cloud data sets of the grid points in the multiple grid point arrays, where the point cloud data set includes the point cloud data around the grid points; a second processing module for processing the point cloud data set to obtain the second region where the target object is located, and the second region represents the detection result of the target object.

[0033] As can be seen from the above device: after obtaining the point cloud data of the target scene, the object detection model can process the point cloud data of the target scene to obtain the first region where the target object is located in the target scene. Then, the object detection model can construct a plurality of grid point arrays in the first region and obtain the point cloud data sets of all grid points in the plurality of grid point arrays. The point cloud data set of each grid point includes the point cloud data around the grid point. Finally, the object detection model can process the point cloud data sets of all grid points to obtain the second region where the target object is located. In the foregoing process, since different grid point arrays in the plurality of grid point arrays have different sizes, the sizes of some grid point arrays can be larger than the size of the initial region, and the sizes of some other grid point arrays can be smaller than the size of the initial region. Therefore, based on these plurality of grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data in or near the first region. In this way, the point cloud data sampled by the model can completely represent the first region where the target object is located, so that the second region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0034] In a possible implementation, the construction module is configured to: determine the position of the target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region. The target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; repeat the above steps for the remaining grid point arrays other than the target grid point array in the plurality of grid point arrays until a plurality of grid point arrays are constructed in the first region.

[0035] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

[0036] In a possible implementation, the acquisition module is configured to: obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region. The sampling radius of the target grid point array is different from the sampling radii of other grid point arrays. The target grid point array is any one of the plurality of grid point arrays; determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array. The target grid point is any one of the grid points in the target grid point array; obtain the point cloud data in the sampling range of the target grid point to obtain the point cloud data set of the target grid point; repeat the above steps for the remaining grid points other than the target grid point in the plurality of grid point arrays to obtain the point cloud data sets of all grid points in the plurality of grid point arrays.

[0037] In a possible implementation, the sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere, and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0038] In a possible implementation, among multiple grid point arrays, the sampling radius of a grid point array is positively correlated with the size of the grid point array.

[0039] In a possible implementation, the second processing module is configured to: perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any point cloud data in the point cloud data set of the target grid point, the target grid point is any grid point in the target grid point array, and the target grid point array is any one of multiple grid point arrays; repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid point except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid point; perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid point to obtain the second feature of the target grid point; repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the second features of all the grid points in the multiple grid point arrays; perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the target object is located.

[0040] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0041] In a possible implementation, the second processing module is configured to: perform a linear transformation process on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; perform a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; perform a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; perform a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

[0042] In a possible implementation, the fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

[0043] The fourth aspect of the embodiments of the present application provides a model training device, which includes: an acquisition module, configured to acquire the point cloud data of the target scene and the real area where the object to be detected is located in the target scene; a processing module, configured to input the point cloud data of the target scene into the model to be trained, and obtain a second area where the object to be detected is located. The model to be trained is used to: process the point cloud data of the target scene to obtain a first area where the object to be detected is located in the target scene; construct a plurality of grid point arrays in the first area, and different grid point arrays have different sizes; obtain the point cloud data sets of the grid points in the plurality of grid point arrays, and the point cloud data set includes the point cloud data around the grid points; process the point cloud data set to obtain a second area where the object to be detected is located; a training module, configured to train the model to be trained based on the real area and the second area to obtain an object detection model.

[0044] The object detection model obtained by the above device has the function of detecting the area where the target object is located in the target scene. In the process of three-dimensional object detection, since different grid point arrays in the plurality of grid point arrays have different sizes, the sizes of some grid point arrays can be larger than the size of the initial area, and the sizes of some other grid point arrays can be smaller than the size of the initial area. Therefore, based on these plurality of grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data within or near the initial area. In this way, the point cloud data sampled by the model can completely represent the initial area where the target object is located, so that the final area where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0045] In a possible implementation manner, the model to be trained is used to: determine the position of the target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first area, the number of grid points in the target grid point array, and the parameters of the first area. The target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first area based on the positions of all the grid points in the target grid point array; for the remaining grid point arrays other than the target grid point array in the plurality of grid point arrays, repeat the above steps until a plurality of grid point arrays are constructed in the first area.

[0046] In a possible implementation manner, the parameters of the first area include the size of the first area, the central position of the first area, and the yaw angle of the first area.

[0047] In a possible implementation, the model to be trained is used to: obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of multiple grid point arrays; determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array, where the target grid point is any one of the grid points in the target grid point array; obtain the point cloud data within the sampling range of the target grid point to obtain the point cloud data set of the target grid point; and repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the point cloud data sets of all the grid points in the multiple grid point arrays.

[0048] In a possible implementation, the sampling range of the target grid point is a sphere with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0049] In a possible implementation, among the multiple grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array.

[0050] In a possible implementation, the model to be trained is used to: perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid point, the target grid point is any one of the grid points in the target grid point array, and the target grid point array is any one of multiple grid point arrays; repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid point except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid point; perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid point to obtain the second feature of the target grid point; repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the second features of all the grid points in the multiple grid point arrays; and perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the object to be detected is located.

[0051] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0052] In a possible implementation, the model to be trained is used to: perform a linear transformation on the distance between the target point cloud data and the target grid points to obtain a third feature of the target point cloud data; perform feature extraction and linear transformation on the target point cloud data to obtain a fourth feature of the target point cloud data; perform feature extraction on the target point cloud data to obtain a fifth feature of the target point cloud data; and fuse the third feature, the fourth feature, and the fifth feature to obtain a first feature of the target point cloud data. In a possible implementation, the fusion processing includes at least one of addition processing, multiplication processing, and mapping processing.

[0053] The fifth aspect of the embodiments of the present application provides an object detection device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the object detection device executes the method described in the first aspect or any possible implementation of the first aspect.

[0054] The sixth aspect of the embodiments of the present application provides a device, which may be a vehicle, a wearable device, or a mobile device. The device includes the device described in the fifth aspect.

[0055] The seventh aspect of the embodiments of the present application provides a model training device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training device executes the method described in the second aspect or any possible implementation of the second aspect.

[0056] The eighth aspect of the embodiments of the present application provides a circuit system, which includes a processing circuit configured to execute the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0057] The ninth aspect of the embodiments of the present application provides a chip system, which includes a processor for calling a computer program or computer instruction stored in a memory so that the processor executes the method described in the first aspect, any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0058] In a possible implementation, the processor is coupled to the memory through an interface.

[0059] In a possible implementation, the chip system further includes a memory, and a computer program or computer instruction is stored in the memory.

[0060] The tenth aspect of the embodiments of the present application provides a computer storage medium storing a computer program, which when executed by a computer, causes the computer to implement the method described in the first aspect, any possible implementation manner of the first aspect, the second aspect, or any possible implementation manner of the second aspect.

[0061] The eleventh aspect of the embodiments of the present application provides a computer program product storing instructions, which when executed by a computer, causes the computer to implement the method described in the first aspect, any possible implementation manner of the first aspect, the second aspect, or any possible implementation manner of the second aspect.

[0062] In the embodiments of the present application, after obtaining the point cloud data of the target scene, the object detection model can process the point cloud data of the target scene to obtain the initial area where the target object is located in the target scene. Then, the object detection model can construct a plurality of grid point arrays in the initial area and obtain the point cloud data sets of all grid points in the plurality of grid point arrays. The point cloud data set of each grid point includes the point cloud data around the grid point. Finally, the object detection model can process the point cloud data sets of all grid points to obtain the final area where the target object is located. In the foregoing process, since different grid point arrays in the plurality of grid point arrays have different sizes, the sizes of some grid point arrays can be larger than the size of the initial area, and the sizes of some other grid point arrays can be smaller than the size of the initial area. Therefore, based on these plurality of grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data within or near the initial area. In this way, the point cloud data sampled by the model can completely represent the initial area where the target object is located, so that the final area where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy. Description of the Drawings

[0063] Figure 1 It is a schematic structural diagram of a structure of an artificial intelligence main framework;

[0064] Figure 2a It is a schematic structural diagram of an object detection system provided by an embodiment of the present application;

[0065] Figure 2b It is another schematic structural diagram of an object detection system provided by an embodiment of the present application;

[0066] Figure 2c It is a schematic diagram of a related device for object detection provided by an embodiment of the present application;

[0067] Figure 3 It is a schematic diagram of the architecture of system 100 provided by an embodiment of the present application;

[0068] Figure 4 It is a schematic flowchart of an object detection method provided by an embodiment of the present application;

[0069] Figure 5a It is a schematic diagram of a plurality of grid point arrays provided by an embodiment of the present application;

[0070] Figure 5b It is a schematic diagram of a grid point array provided by an embodiment of the present application;

[0071] Figure 6a It is a schematic diagram for determining a sampling radius provided by an embodiment of the present application;

[0072] Figure 6b It is a schematic diagram for determining a sampling radius provided by an embodiment of the present application;

[0073] Figure 6c It is a schematic diagram for determining a sampling radius provided by an embodiment of the present application;

[0074] Figure 6d It is a schematic diagram for determining a sampling radius provided by an embodiment of the present application;

[0075] Figure 7 It is a schematic flowchart of a model training method provided by an embodiment of the present application;

[0076] Figure 8 It is a schematic structural diagram of an object detection device provided by an embodiment of the present application;

[0077] Figure 9 It is a schematic structural diagram of a model training device provided by an embodiment of the present application;

[0078] Figure 10 It is a schematic structural diagram of an execution device provided by an embodiment of the present application;

[0079] Figure 11 It is a schematic structural diagram of a training device provided by an embodiment of the present application;

[0080] Figure 12 It is a schematic structural diagram of a chip provided by an embodiment of the present application. Detailed implementation manners

[0081] The embodiments of the present application provide an object detection method and related devices. In the second stage of object detection, the point cloud data sampled by the object detection model can completely represent the initial region where the target object is located, so that the final region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0082] In the description, claims and the above-mentioned drawings of this application, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0083] 3D object detection is one of the important tasks in computer vision and has important applications in fields such as autonomous driving and industrial vision.

[0084] Currently, point cloud data of a target scene can be obtained through lidar to determine the area where a target object is located in the target scene. Specifically, after obtaining the point cloud data of the target scene, this part of the point cloud data can be input into an object detection model, so that the object detection model performs two-stage processing on this part of the point cloud data. Among them, the first stage is the stage of predicting the initial area where the object is located, and the second stage is to optimize the initial area where the object is located to obtain the final area where the object is located. In the first stage, the object detection model can process the point cloud data of the target scene and predict the initial area where the target object is located. In the second stage, the object detection model can sample the point cloud data within or near the initial area and further process the sampled point cloud data to obtain the final area where the target object is located.

[0085] However, when the object detection model samples point cloud data, its sampling scale is usually single. Therefore, the point cloud data sampled by the model cannot fully represent the initial area where the target object is located, resulting in the final area where the target object is located obtained by the model based on these point cloud data being inaccurate.

[0086] To solve the above problems, the embodiments of this application provide an object detection method, which can be implemented by combining artificial intelligence (AI) technology. AI technology is a technical discipline that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by perceiving the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Using artificial intelligence to solve partial differential equations is a common application method of artificial intelligence.

[0087] First, the overall workflow of the artificial intelligence system is described. Please refer to Figure 1 , Figure 1 which is a schematic diagram of a structure of the artificial intelligence main framework. The above artificial intelligence theme framework is elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0088] (1) Infrastructure

[0089] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips (such as CPU, NPU, GPU, ASIC, FPGA and other hardware acceleration chips); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.

[0090] (2) Data

[0091] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the perception data such as force, displacement, liquid level, temperature, humidity, etc.

[0092] (3) Data Processing

[0093] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making and other methods.

[0094] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0095] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, based on the reasoning control strategy, using formal information for machine thinking and problem-solving. The typical function is search and matching.

[0096] Decision-making refers to the process of making decisions after intelligent information is inferred, usually providing functions such as classification, sorting, prediction, etc.

[0097] (4) General capabilities

[0098] After the data is processed as mentioned above, some general capabilities can be formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0099] (5) Intelligent products and industry applications

[0100] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields. It is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.

[0101] Next, several application scenarios of this application will be introduced.

[0102] Figure 2a FIG. is a schematic structural diagram of an object detection system provided for an embodiment of this application. The object detection system includes a user device and a data processing device. Among them, the user device includes intelligent terminals such as mobile phones, personal computers, or information processing centers. The user device is the initiating end of object detection and is the initiator of the object detection request. Usually, the user initiates a request through the user device.

[0103] The above-mentioned data processing device can be a device or server with data processing functions such as a cloud server, a network server, an application server, and a management server. The data processing device receives an image processing request from the intelligent terminal through an interaction interface, and then performs image processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through a memory for storing data and a processor for data processing. The memory in the data processing device can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.

[0104] In Figure 2aIn the object detection system shown, the user device can receive instructions from the user. For example, the user device can obtain the point cloud data of the target scene input / selected by the user, and then send a request to the data processing device, so that the data processing device performs an object detection application (such as three-dimensional object detection, etc.) on the point cloud data obtained by the user device, thereby obtaining the corresponding processing result for the point cloud data. Exemplarily, the user can collect the point cloud data of the target scene through a lidar, input this part of the point cloud data into the user device, and then send an object detection request to the data processing device, so that the data processing device performs object detection on this part of the point cloud data, thereby obtaining the area where the target object is located in the target scene, that is, obtaining information such as the position and orientation of the target object in the target scene.

[0105] In Figure 2a the data processing device can execute the object detection method of the embodiments of the present application.

[0106] Figure 2b Fig. is another schematic structural diagram of the object detection system provided by the embodiments of the present application. In Figure 2b the user device directly serves as the data processing device. This user device can directly obtain the input from the user and directly process it by the hardware of the user device itself. The specific process is similar to Figure 2a and can refer to the above description, which will not be elaborated here.

[0107] In Figure 2b the object detection system shown, the user device can receive instructions from the user. For example, the user device can obtain the point cloud data of the target scene input by the user in the user device, and then the user device itself performs an object detection application (such as three-dimensional target detection, etc.) on the point cloud data, thereby obtaining the corresponding processing result for the point cloud data.

[0108] In Figure 2b the user device itself can execute the object detection method of the embodiments of the present application.

[0109] Figure 2c Fig. is a schematic diagram of related devices for object detection provided by the embodiments of the present application.

[0110] The above Figure 2a and Figure 2b the user device in can specifically be Figure 2c the local device 301 or the local device 302 in, Figure 2a the data processing device in can specifically be Figure 2c the execution device 210 in. Among them, the data storage system 250 can store the data to be processed by the execution device 210. The data storage system 250 can be integrated on the execution device 210 or set on the cloud or other network servers.

[0111] Figure 2a and Figure 2b The processor in can perform data training / machine learning / deep learning through a neural network model or other models (e.g., a support vector machine-based model), and use the model finally trained or learned from the data to perform an image processing application on the image, so as to obtain a corresponding processing result.

[0112] Figure 3 FIG. is a schematic diagram of the architecture of system 100 provided by an embodiment of the present application. In Figure 3 it, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. A user can input data to the I / O interface 112 through the client device 140. The input data in the embodiment of the present application may include: each task to be scheduled, callable resources, and other parameters.

[0113] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs relevant processing such as computing (such as implementing the functions of the neural network in the present application), the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.

[0114] Finally, the I / O interface 112 returns the processing result to the client device 140 for providing to the user.

[0115] It should be noted that the training device 120 can generate corresponding target models / rules based on different targets or different tasks and different training data. The corresponding target models / rules can be used to achieve the above targets or complete the above tasks, so as to provide the results required by the user. Among them, the training data can be stored in the database 130 and comes from the training samples collected by the data collection device 160.

[0116] In Figure 3In the case shown, the user can manually provide input data, and this manual provision can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 in the client device 140, and the specific presentation form can be specific ways such as display, sound, action, etc. The client device 140 can also be used as a data acquisition end to collect the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data and store them in the database 130. Of course, it is also possible not to collect through the client device 140, but directly use the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as shown in the figure as new sample data and store them in the database 130.

[0117] It should be noted that Figure 3 is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 3 , the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110. As Figure 3 shown, a neural network can be trained according to the training device 120.

[0118] The embodiments of the present application also provide a chip, which includes a neural network processor NPU. The chip can be set in the execution device 110 as shown in Figure 3 to complete the computing work of the computing module 111. The chip can also be set in the training device 120 as shown in Figure 3 to complete the training work of the training device 120 and output the target model / rule.

[0119] The neural network processor NPU is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are assigned by the main CPU. The core part of the NPU is the arithmetic circuit, and the controller controls the arithmetic circuit to extract data from the memory (weight memory or input memory) and perform arithmetic operations.

[0120] In some implementations, the arithmetic circuit includes multiple processing units (process engines, PEs) inside. In some implementations, the arithmetic circuit is a two-dimensional systolic array. The arithmetic circuit can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit is a general matrix processor.

[0121] For example, suppose there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator.

[0122] The vector calculation unit can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, and so on. For example, the vector calculation unit can be used for network calculations in non-convolution / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.

[0123] In some implementations, the vector calculation unit can store the processed output vector into the unified buffer. For example, the vector calculation unit can apply a non-linear function to the output of the arithmetic circuit, such as the vector of the accumulated value, to generate activation values. In some implementations, the vector calculation unit generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit, such as for use in subsequent layers in a neural network.

[0124] The unified memory is used to store input data and output data.

[0125] The weight data directly transfers the input data in the external memory to the input memory and / or the unified memory, stores the weight data in the external memory into the weight memory, and stores the data in the unified memory into the external memory through the direct memory access controller (DMAC).

[0126] The bus interface unit (BIU) is used to interact between the main CPU, DMAC, and the instruction fetch memory through the bus.

[0127] An instruction fetch buffer connected to the controller, which is used to store instructions used by the controller;

[0128] The controller is used to call the instructions cached in the instruction fetch memory to control the working process of the arithmetic accelerator.

[0129] Generally, the unified memory, input memory, weight memory, and instruction fetch memory are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDRSDRAM), a high bandwidth memory (HBM), or other readable and writable memories.

[0130] Since the embodiments of this application involve a large number of neural network applications, for the sake of easy understanding, relevant terms and related concepts such as neural networks involved in the embodiments of this application will be introduced first below.

[0131] (1) Neural network

[0132] A neural network can be composed of neural units. A neural unit can be an arithmetic unit that takes xs and intercept 1 as inputs. The output of this arithmetic unit can be:

[0133]

[0134] where s = 1, 2,... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be a region composed of several neural units.

[0135] The operation of each layer in a neural network can be described by the mathematical expression y = a(Wx + b). Physically, the operation of each layer in a neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / reduction; 2. Enlargement / reduction; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 1, 2, and 3 are completed by Wx, operation 4 is completed by +b, and operation 5 is implemented by a(). The reason for using the word "space" here is that the objects to be classified are not individual things, but a class of things. Space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. The vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training a neural network is to finally obtain the weight matrices of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially a process of learning the way to control space transformation, and more specifically, learning the weight matrix.

[0136] Because it is desired that the output of the neural network is as close as possible to the value that is truly desired to be predicted, the weight vector of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the neural network). For example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the neural network can predict the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.

[0137] (2) Backpropagation algorithm

[0138] The neural network can use the backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial neural network model during the training process, making the reconstruction error loss of the neural network model smaller and smaller. Specifically, forward propagating the input signal until the output generates an error loss, and updating the parameters in the initial neural network model by backpropagating the error loss information, so as to converge the error loss. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0139] The method provided in this application will be described below from the training side of the neural network and the application side of the neural network.

[0140] The model training method provided in the embodiments of this application is related to the processing of images, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data (such as the point cloud data of the target scene in the model training method of the embodiments of this application), and finally obtains a trained neural network (such as the object detection model in the embodiments of this application); moreover, the object detection method provided in the embodiments of this application can use the above-mentioned trained neural network to input input data (such as the point cloud data of the target scene in the object detection method of the embodiments of this application) into the trained neural network to obtain output data (such as the second region where the target object is located in the embodiments of this application, etc.). It should be noted that the model training method and the image processing method provided in the embodiments of this application are inventions generated based on the same concept, and can also be understood as two parts of a system, or two stages of an overall process: such as the model training stage and the model application stage.

[0141] Figure 4 FIG. is a schematic flowchart of an object detection method provided in an embodiment of this application. This method can be implemented through an object detection model, and the object detection model can process the point cloud data of the target scene to determine the region where the target object is located in the target scene, that is, information such as the position and orientation of the target object in the target scene. As Figure 4 shown, this method includes:

[0142] 401. Process the point cloud data of the target scene to obtain the first region where the target object is located in the target scene.

[0143] In this embodiment, when the user needs to perform object detection on a target scene, the point cloud data of the target scene can be obtained and input into an object detection model, so that the object detection model processes the point cloud data of the target scene to determine the region where the target object is located in the target scene. For example, after the user enables the automatic driving function of the vehicle, the lidar of the vehicle can collect the point cloud data of the surrounding environment and send the point cloud data to an in-vehicle device with data processing capabilities in the vehicle (such as a telematics box (T-BOX), etc.). The in-vehicle device has an object detection model built in, so it can process the point cloud data to detect the region where the other vehicles around the vehicle are located, that is, determine information such as the position and orientation of the other vehicles around the vehicle.

[0144] Specifically, after receiving the point cloud data of the target scene, the object detection model can process the point cloud data of the target scene in various ways to determine the first region where the target object is located in the target scene (which can also be called the initial region where the target object is located in the target scene, or the initial detection box surrounding the target object in the target scene, etc.). The following will be introduced separately:

[0145] In a possible implementation manner, the object detection model can obtain the preset number of key points and the preset sampling radius, and select multiple point cloud data as key points in the point cloud data of the target scene based on the farthest point sampling (FPS) algorithm. After determining multiple key points, for any one key point, the object detection model takes this key point as the sampling center, obtains the point cloud data within the sampling radius, and performs feature extraction processing on these point cloud data to obtain the feature of this key point. It should be noted that the remaining key points can also perform the same operations as this key point, so the features of multiple key points can be obtained. After obtaining the features of multiple key points, the object detection model can perform further feature extraction processing on the features of multiple key points to obtain the first region where the target object is located in the target scene.

[0146] In another possible implementation manner, the object detection model can obtain the preset voxel size and equally divide the target scene (the entire detection space) into multiple voxels. It can be understood that each voxel can contain a certain amount of point cloud data. After completing the voxel division, for any one voxel, the object detection model can perform feature extraction processing on the point cloud data in this voxel to obtain the feature of this voxel. It should be noted that the remaining voxels can also perform the same operations as this voxel, so the features of multiple voxels can be obtained. After obtaining the features of multiple voxels, the object detection model can perform further feature extraction processing on the features of multiple voxels to obtain the first region where the target object is located in the target scene.

[0147] After obtaining the first region where the target object is located, the object detection model can preliminarily determine information such as the position and orientation of the target object in the target scene. It can be understood that the first region where the target object is located can be regarded as a cuboid, so it contains multiple parameters. For example, the first region where the target object is located can be represented as a 7-dimensional vector, that is, (x, y, z, w, h, l, θ), where x is the abscissa of the center of the first region, y is the ordinate of the center of the first region, z is the vertical coordinate of the center of the first region, w is the width of the first region, h is the height of the first region, l is the length of the first region, and θ is the yaw angle of the first region.

[0148] It should be understood that in the foregoing first implementation manner, if no point cloud data can be sampled with a certain key point as the sampling center (indicating that there is no point cloud data around the key point), the feature of the key point can be regarded as an empty feature, that is, a zero vector. Similarly, in the foregoing second implementation manner, if a voxel does not contain any point cloud data, the feature of the voxel can be regarded as an empty feature, that is, a zero vector.

[0149] It should also be understood that this embodiment only gives a schematic illustration in the autonomous driving scenario and does not limit the application scenario of the present application. For example, the present application can also be applied to an augmented reality (AR) scenario or a mixed reality (MR) scenario. Correspondingly, the foregoing device with data processing capabilities can be a wearable device, etc. Another example is that the present application can also be used in a smart home scenario. Correspondingly, the foregoing device with data processing capabilities can be a sweeping robot, etc.

[0150] 402. Construct multiple grid point arrays in the first region, and different grid point arrays have different sizes.

[0151] After determining the first region where the target object is located, the object detection model can construct multiple grid point arrays in the first region. In these multiple grid point arrays, each grid point array contains multiple grid points, and each grid point array is a three-dimensional array. Different grid point arrays have different sizes. The sizes of some grid point arrays can be larger than the size of the first region, and the sizes of some other grid point arrays can be smaller than the size of the first region. Therefore, these multiple grid point arrays can be used to sample point cloud data in or near the first region at different scales. As Figure 5a shown ( Figure 5a(A schematic diagram of multiple grid point arrays provided by an embodiment of the present application. In the first region, 4 grid point arrays are constructed, namely grid point array 1, grid point array 2, grid point array 3, and grid point array 4. In the 3 grid point arrays of grid point array 1, grid point array 2, and grid point array 3, each grid point array contains 4×4×4 grid points (that is, each grid point array is an array of 4 rows, 4 columns, and 4 layers). Grid point array 4 contains 6×6×6 grid points (that is, this grid point array is an array of 6 rows, 6 columns, and 6 layers). However, among these 4 grid point arrays, the sizes (i.e., length, width, and height) of different grid point arrays are different. The size of grid point array 1 is the largest, and the size of grid point array 4 is the smallest.)

[0152] To further understand the size of the grid point array, the following will describe the size of the grid point array in conjunction with Figure 5b as shown in Figure 5b ( Figure 5b (A schematic diagram of the grid point array provided by an embodiment of the present application. Suppose only grid point array 1 is constructed in the first region, and this grid point array contains 4×4×4 grid points. Then, Figure 5b the cuboid larger than the first region is the area occupied by grid point array 1, and the size of this area can be regarded as the size of grid point array 1, that is, the length of this area can be regarded as the length of grid point array 1, the width of this area can be regarded as the width of grid point array 1, and the height of this area can be regarded as the height of grid point array 1.)

[0153] Specifically, the object detection model can construct multiple grid point arrays in the first region through the following method:

[0154] In multiple grid point arrays, for any grid point in any grid point array, hereinafter this grid point array will be called the target grid point array, and this grid point in this grid point array will be called the target grid point.)

[0155] First, the object detection model can obtain information such as the serial number of the target grid points in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region. Among them, the serial number of the target grid points can include the serial numbers of the target grid points in the three dimensions of length, width, and height of the target grid point array (i.e., the target grid points are the grid points of a certain row, certain column, and certain vertical column in the target grid point array). The ratio between the size of the target grid point array and the size of the first region can include the ratio between the length of the target grid point array and the length of the first region, the ratio between the width of the target grid point array and the width of the first region, and the ratio between the height of the target grid point array and the height of the first region. The number of grid points in the target grid point array can include the number of grid points in each dimension direction of the target grid point array (i.e., the target grid point array is an array of several rows, several columns, and several vertical columns). The parameters of the first region include the abscissa of the center of the first region, the ordinate of the center of the first region, the vertical coordinate of the center of the first region, the length of the first region, the width of the first region, the height of the first region, and the yaw angle of the first region.

[0156] Then, based on the serial number of the target grid points in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, the object detection model can determine the position of the target grid points (i.e., the three-dimensional coordinates of the target grid points). Still using the above example, assume that M grid point arrays need to be constructed in the first region, and any one of the grid point arrays is used for schematic introduction. Assume that this grid point array is the m-th grid point array, where m = 1,..., M. Then, the coordinates of any grid point in the m-th grid point array can be expressed by the following formula:

[0157]

[0158]

[0159] In the above formula, is the three-dimensional coordinate of the grid point at the i-th row, j-th column, and k-th vertical column in the m-th grid point array, is the number of grid points in the width direction of the m-th grid point array (i.e., the number of rows of the m-th grid point array), is the number of grid points in the length direction of the m-th grid point array (i.e., the number of columns of the m-th grid point array), is the number of grid points in the height direction of the m-th grid point array (i.e., the number of vertical columns of the m-th grid point array), is the ratio between the width of the m-th grid point array and the width of the first region, is the ratio between the length of the m-th grid point array and the length of the first region, is the ratio between the height of the m-th grid point array and the height of the first region, x is the abscissa of the center of the first region, y is the ordinate of the center of the first region, z is the vertical coordinate of the center of the first region, w is the width of the first region, l is the length of the first region, h is the height of the first region, and θ is the yaw angle of the first region.

[0160] It should be noted that, except for the target grid points, the object detection model can also perform operations on the remaining grid points in the target grid point array as it does on the target grid points. Therefore, the positions of all grid points in the target grid point array can be obtained.

[0161] Finally, the object detection model can construct the target grid point array in the first region based on the positions of all grid points in the target grid point array. It should be noted that, except for the target grid point array, the object detection model can also perform operations on the remaining grid point arrays in the multiple grid point arrays as it does on the target grid point array. Therefore, the object detection model can successfully construct multiple grid point arrays in the first region.

[0162] It should be understood that Figure 5a In the example shown, only some grid point arrays containing the same number of grid points are used for illustrative purposes, and it does not limit the number of grid points in the grid point arrays in the embodiments of the present application. In practical applications, all grid point arrays can also contain different numbers of grid points. For example, grid point array 1 contains 4×4×4 grid points, grid point array 2 contains 6×6×6 grid points, grid point array 3 contains 8×8×8 grid points, grid point array 4 contains 10×10×10 grid points, and so on. Currently, all grid point arrays can also contain the same number of grid points. For example, grid point array 1 contains 4×4×4 grid points, grid point array 2 contains 4×4×4 grid points, grid point array 3 contains 4×4×4 grid points, grid point array 4 contains 4×4×4 grid points, and so on.

[0163] 403. Obtain the point cloud data set of the grid points in multiple grid point arrays. The point cloud data set of the grid points contains the point cloud data around the grid points.

[0164] After constructing multiple grid point arrays in the first region, since each grid point array contains multiple grid points, for any grid point, the object detection model can obtain the point cloud data around the grid point as the point cloud data set of the grid point.

[0165] Specifically, the object detection model can obtain the point cloud data set of the grid points in the following way:

[0166] First, the object detection model can obtain the sampling radii of multiple grid point arrays, where the sampling radii of different grid point arrays are different. Among the multiple grid point arrays, there is usually a certain correlation between the sampling radius of a grid point array and the size of the grid point array. For example, the sampling radius of a grid point array is positively correlated with the size of the grid point array, that is, the larger the size of the grid point array, the larger the sampling radius of the grid point array. For example, assume there are grid point array 1, grid point array 2, grid point array 3, and grid point array 4, and the size order of the grid point arrays is: grid array 1, grid array 2, grid array 3, and grid array 4. Then, the sampling radius size order of the grid point arrays is: grid array 1, grid array 2, grid array 3, and grid array 4. Of course, there can also be other mathematical relationships between the sampling radius of a grid point array and the size of the grid point array, which is not limited here.

[0167] It should be noted that for the target grid point array among the multiple grid point arrays, the object detection model can determine the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region (that is, for any grid point array, the object detection model can determine the size of the sampling radius of the grid point array based on the distribution of the point cloud data in the first region). Then, all the grid points in the target grid point array can share the sampling radius of the target grid point array. For ease of understanding, the following will further introduce the determination process of the sampling radius in combination with Figure 6a , Figure 6b , Figure 6c and Figure 6d ( Figure 6a is a schematic diagram for determining the sampling radius in an embodiment of the present application, Figure 6b is a schematic diagram for determining the sampling radius in an embodiment of the present application, Figure 6c is a schematic diagram for determining the sampling radius in an embodiment of the present application, Figure 6d is a schematic diagram for determining the sampling radius in an embodiment of the present application). As Figure 6a shows, the distribution of the point cloud data in the first region is the densest, and the sampling radius of the target grid point array is the smallest. As Figure 6b shows, the distribution of the point cloud data in the first region is relatively dense, and the sampling radius of the target grid point array is relatively small. As Figure 6c shows, the distribution of the point cloud data in the first region is relatively sparse, and the sampling radius of the target grid point array is relatively large. As Figure 6d shows, the distribution of the point cloud data in the first region is the sparsest, and the sampling radius of the target grid point array is the largest. It can be seen that the denser the distribution of the point cloud data in the first region, the smaller the sampling radius of the target grid point array. Similarly, the object detection model can also determine the sampling radii of other grid point arrays based on the distribution of the point cloud data in the first region, and the determination process can refer to the determination process of the sampling radius of the target grid point array, which will not be elaborated here.

[0168] Then, for the target grid points in the target grid point array, the object detection model can determine the sampling range of the target grid points based on the target grid points and the sampling radius of the target grid point array. As Figure 6a shown, the sampling range of the target grid points is usually a sphere, which takes the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere. It should be noted that this parameter is usually a fixed value, and its size can be set according to actual needs and is not limited here.

[0169] Finally, the object detection model obtains the point cloud data in the sampling range of the target grid points to obtain the point cloud data set of the target grid points. Still taking the above example, after constructing M grid point arrays in the first region, assuming the sampling radius of the m-th grid point array is r, for the grid point in the i-th row, j-th column, and k-th vertical of the m-th grid point array, assuming the sampling range of this grid point is U(r), the sampling range U(r) takes this grid point as the center of the sphere and r + 5τ as the radius of the sphere. Then, the point cloud data located within the sampling range U(r) can be regarded as the point cloud data set of this grid point.

[0170] As for why r + 5τ is taken as the sphere radius, it can be analyzed by the probability of the point cloud data in the sampling range U(r) being collected. For example, for the t-th point cloud data in the sampling range U(r), the probability of the t-th point cloud data being collected can be calculated by the following formula:

[0171]

[0172] In the above formula, s(t|r) is the probability of the t-th point cloud data in the sampling range U(r) of this grid point being collected, p t is the coordinate of the t-th point cloud data, is the distance between the t-th point cloud data and this grid point, sigmoid(a) = (1 + e -a ) -1 , and τ is a preset value. Through actual tests, it can be seen that the sampling range U(r) includes the point cloud data within a distance of r + 5τ from this grid point. Taking the t-th point cloud data in the sampling range U(r) for calculation, s(t|r) is greater than 0.001 (that is, the probability of the t-th point cloud data being collected is greater than 0.001). Taking a point cloud data outside the sampling range U(r) for calculation, the probability of this point cloud data being collected is less than or equal to 0.001. It can be seen that the point cloud data located outside the sampling range U(r) can be ignored, which can reduce the amount of calculation required for subsequent calculation of the features of this grid point.

[0173] It should be noted that, in addition to the target grid points, the object detection model can also perform operations on the remaining grid points in the target grid point array similar to those on the target grid points. Therefore, a point cloud data set of all grid points in the target grid point array can be obtained. Further, in addition to the target grid point array, the object detection model can also perform operations on the remaining grid point arrays in the multiple grid point arrays similar to those on the target grid point array. Therefore, the object detection model can obtain a point cloud data set of all grid points in the multiple grid point arrays.

[0174] It should be understood that Figures 6a to 6d In the example shown, only the denser the distribution of the point cloud data in the first region, the smaller the sampling radius of the grid point array is used for illustrative purposes, which does not limit the correlation between the distribution of the point cloud data in the first region and the sampling radius of the grid point array in the embodiments of the present application.

[0175] It should also be understood that in the foregoing example, only r + 5τ is used for illustrative purposes, which does not limit the radius size of the sampling range of the grid points in the present application, and its size can be set according to actual needs.

[0176] 404. Process the point cloud data set to obtain a second region where the target object is located.

[0177] After obtaining the point cloud data set of all grid points in the multiple grid point arrays, the object detection model can process the point cloud data set of all grid points, so as to obtain a second region (which can also be called the final region where the target object is located in the target scene, and can also be called the final detection box surrounding the target object in the target scene, etc.) where the target object is located in the target scene.

[0178] Specifically, the object detection model can obtain the second region where the target object is located through the following methods:

[0179] In the multiple grid point arrays, for any point cloud data in the point cloud data set of the target grid points of the target grid point array, hereinafter, this point cloud data can be referred to as target point cloud data.

[0180] First, the object detection model can perform a first feature extraction process on the target point cloud data to obtain a first feature of the target point cloud data. It should be noted that the first feature extraction process performed by the object detection model can be a process based on an attention mechanism, including:

[0181] (1) The object detection model can perform a linear transformation on the distance between the target point cloud data and the target grid points to obtain the third feature of the target point cloud data, that is, the Q feature in the attention mechanism. Still as in the above example, for the grid point at the i-th row, j-th column, and k-th vertical of the m-th grid point array, the t-th point cloud data is taken from the sampling range U(r) of this grid point. Therefore, the distance between the t-th point cloud data and this grid point can be calculated, and a linear transformation is performed on the distance between the two to obtain the Q feature of the t-th point cloud data

[0182] (2) The object detection model can also perform feature extraction processing and linear transformation processing on the target point cloud data to obtain the fourth feature of the target point cloud data, that is, the K feature in the attention mechanism. Still as in the above example, feature extraction can be performed on the t-th point cloud data to obtain the feature f of the t-th point cloud data t . Then, a linear transformation is performed on the initial feature f of the t-th point cloud data t to obtain the K feature K of the t-th point cloud data t = Linear(f t ).

[0183] (3) The object detection model can also perform feature extraction processing on the target point cloud data to obtain the fifth feature of the target point cloud data, that is, the V feature in the attention mechanism. Still as in the above example, feature extraction can be performed on the t-th point cloud data to obtain the feature f of the t-th point cloud data t . Then, a multi-layer perceptron is used to perform multiple feature extractions on the initial feature f of the t-th point cloud data t to obtain the V feature V of the t-th point cloud data t = MLP(f t ).

[0184] (4) After obtaining the second, third, and fourth features of the target point cloud data, the object detection model performs a fusion process on the third, fourth, and fifth features of the target point cloud data to obtain the first feature of the target point cloud data. Still as in the above example, after obtaining the Q feature Q t , K feature K t , and V feature V t of the t-th point cloud data, these three features can be fused through the following formula to obtain the final feature (i.e., the aforementioned first feature) of the t-th point cloud data:

[0185] R t = W(σ k K t + σ q Q t + σ qk Q t K t)⊙(V t +σ v Q t ) (4)

[0186] In the above formula, R t is the final feature of the t-th point cloud data, and σ k , σ q , σ qk and σ v are weights (the magnitudes of these weights can be set according to the Q feature Q t , K feature K t and V feature V t of the t-th point cloud data, and there is no restriction here), W is a mapping process that can map a vector to a vector or a scalar, ⊙ is a multiplication process that can represent element-wise multiplication, matrix multiplication, or scalar-vector multiplication, etc., depending on the numerical types on both sides of the symbol.

[0187] It should be noted that in addition to the target point cloud data, the object detection model can also perform the same operations on the remaining point cloud data in the point cloud data set of the target grid points as on the target point cloud data, so the first feature of all the point cloud data in the point cloud data set of the target grid points can be obtained.

[0188] Then, the object detection model can perform a weighted sum process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid points. In the above-mentioned weighted sum process, for any point cloud data in the point cloud data set of the target grid points, the weight of this point cloud data is determined based on the distance between this point cloud data and the target grid point and the array sampling radius of the target grid point, that is, the probability that this point cloud data is collected. Still as in the above example, for the grid point in the m-th grid point array, in the i-th row, j-th column, and k-th vertical, the weighted sum process can be performed on the final features of all the point cloud data in the sampling range U(r) of this grid point, so as to obtain the feature of this grid point (i.e., the aforementioned second feature), that is:

[0189]

[0190] In the above formula, is the feature of this grid point, and s(t|r) is the weight of the t-th point cloud data in the sampling range U(r) of this grid point, that is, the probability that the t-th point cloud data is collected.

[0191] It should be noted that, in addition to the target grid points, the object detection model can also perform the same operations on the remaining grid points in the target grid point array as on the target grid points. Therefore, the second features of all grid points in the target grid point array can be obtained. Further, in addition to the target grid point array, the object detection model can also perform the same operations on the remaining grid point arrays in the multiple grid point arrays as on the target grid point array. Therefore, the object detection model can obtain the second features of all grid points in the multiple grid point arrays.

[0192] Finally, the object detection model can perform a second feature extraction process (such as at least one of multiplication processing, addition processing, concatenation processing, concatenated convolution processing, pooling processing, normalization processing, etc.) on the second features of all grid points in the multiple grid point arrays to obtain the second region where the target object is located.

[0193] After obtaining the second region where the target object is located, the object detection model can finally determine information such as the position and orientation of the target object in the target scene. It can be understood that the second region where the target object is located can also be regarded as a cuboid, so it contains multiple parameters. For the introduction of the parameters of the second region, reference can be made to the relevant description part of the parameters of the aforementioned first region, which will not be elaborated here.

[0194] It should be noted that after determining the second region where the target object is located, the object detection model can also perform an optimization process on the second region, so as to output the optimized second region where the target object is located for user use. For example, after the object detection model of the vehicle-mounted device detects the regions where the other vehicles around the vehicle are located, it can use the non-maximum supression (NMS) algorithm to filter the regions where the other vehicles are located, so as to avoid overlapping objects in the region, and return the filtered result to the vehicle-mounted device, so that the vehicle-mounted device can realize the automatic driving function.

[0195] In addition, the embodiments of the present application mainly improve the processing of the second stage of the object detection model. To prove the improvement effect, the object detection model provided by the embodiments of the present application can be compared with the object detection model of the related technology. Whether it is the object detection model provided by the embodiments of the present application or the object detection model of the related technology, the processing of the first stage is the same, but the processing of the second stage is different. In the comparison experiment, the same test data set (including multiple frames of point cloud data) is input into the object detection model provided by the embodiments of the present application and the object detection model of the related technology. Based on the object detection results output by both, the object detection model provided by the embodiments of the present application has the following advantages: (1) It can capture more background information, which helps to identify and predict the position of distant sparse targets; (2) The detection results on different data sets are consistent, that is, it has strong generalization ability, and relatively accurate detection effects can be obtained for different data acquisition conditions and scenarios; (3) It can achieve improvements in various performances, indicating that the model provided by the embodiments of the present application has good adaptability to the backbone network.

[0196] In the embodiments of the present application, after obtaining the point cloud data of the target scene, the object detection model can process the point cloud data of the target scene to obtain the initial region where the target object is located in the target scene. Then, the object detection model can construct multiple grid point arrays in the initial region and obtain the point cloud data set of all grid points in the multiple grid point arrays. The point cloud data set of each grid point includes the point cloud data around the grid point. Finally, the object detection model can process the point cloud data set of all grid points to obtain the final region where the target object is located. In the foregoing process, since different grid point arrays in the multiple grid point arrays have different sizes, the sizes of some grid point arrays can be larger than the size of the initial region, and the sizes of some other grid point arrays can be smaller than the size of the initial region. Therefore, based on these multiple grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data within or near the initial region. In this way, the point cloud data sampled by the model can completely represent the initial region where the target object is located, so that the final region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0197] Further, the object detection model can adjust the sampling radius of each grid point array according to the distribution of the point cloud data in the initial region, so as to realize the dynamic adjustment of the point cloud data sampling range, avoid the situation that the object detection model collects invalid point cloud data, reduce the calculation amount of the model, and is beneficial to saving computing resources.

[0198] Furthermore, when the object detection model calculates the features of a grid point, it can assign certain weights to each point cloud data in the point cloud data set of the grid point, taking into account the different impacts of the point cloud data at different distances from the grid point on the grid point. Therefore, the calculated features of the grid point can contain more information, thereby further improving the accuracy of the final area where the target object is located.

[0199] The above is a detailed description of the object detection method provided by the embodiments of the present application. Next, the model training method provided by the embodiments of the present application will be introduced. Figure 7 It is a schematic flowchart of a model training method provided by an embodiment of the present application. As Figure 7 shown, the method includes:

[0200] 701. Obtain the point cloud data of the target scene and the true area where the object to be detected is located in the target scene.

[0201] When the model to be trained needs to be trained, a batch of training samples, that is, the point cloud data of the target scene for training, can be obtained. It should be noted that the true area where the object to be detected is located in the target scene is known, so the true area where the object to be detected is located can be directly obtained.

[0202] 702. Input the point cloud data of the target scene into the model to be trained, and obtain a second area where the object to be detected is located. The model to be trained is used to: process the point cloud data of the target scene to obtain a first area where the object to be detected is located in the target scene; construct multiple grid point arrays in the first area, and different grid point arrays have different sizes; obtain the point cloud data sets of the grid points in the multiple grid point arrays, and the point cloud data set contains the point cloud data around the grid points; process the point cloud data set to obtain a second area where the object to be detected is located.

[0203] After obtaining the point cloud data of the target scene, the point cloud data of the target scene can be input into the model to be trained to process these point cloud data through the model to be trained and obtain a second area where the object to be detected is located. Among them, the model to be trained can perform the following steps: process the point cloud data of the target scene to obtain a first area where the object to be detected is located in the target scene; construct multiple grid point arrays in the first area, and different grid point arrays have different sizes; obtain the point cloud data sets of the grid points in the multiple grid point arrays, and the point cloud data set contains the point cloud data around the grid points; process the point cloud data set to obtain a second area where the object to be detected is located (that is, the predicted area where the object to be detected is located).

[0204] In a possible implementation, the model to be trained is used to: determine the position of a target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, where the target grid point array is any one of multiple grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first region based on the positions of all the grid points in the target grid point array; and repeat the above steps for the remaining grid point arrays other than the target grid point array among the multiple grid point arrays until multiple grid point arrays are constructed in the first region.

[0205] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

[0206] In a possible implementation, the model to be trained is used to: obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of multiple grid point arrays; determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array, where the target grid point is any one of the grid points in the target grid point array; obtain the point cloud data within the sampling range of the target grid point to obtain the point cloud data set of the target grid point; and repeat the above steps for the remaining grid points other than the target grid point among the multiple grid point arrays to obtain the point cloud data sets of all the grid points in the multiple grid point arrays.

[0207] In a possible implementation, the sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0208] In a possible implementation, among the multiple grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array.

[0209] In a possible implementation, the model to be trained is used to: perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any grid points in the target grid point array, and the target grid point array is any one of multiple grid point arrays; for the remaining point cloud data in the point cloud data set of the target grid points except the target point cloud data, repeat the above steps to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; for the remaining grid points in the multiple grid point arrays except the target grid points, repeat the above steps to obtain the second features of all the grid points in the multiple grid point arrays; perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the object to be detected is located.

[0210] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0211] In a possible implementation, the model to be trained is used to: perform a linear transformation process on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; perform a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; perform a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; perform a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

[0212] In a possible implementation, the fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

[0213] For the descriptions of the respective steps performed by the model to be trained, reference can be made to the relevant description parts of steps 401 to 404 in the foregoing Figure 4 shown in the embodiment, which will not be elaborated here.

[0214] 703. Based on the real region and the second region, train the model to be trained to obtain an object detection model.

[0215] After obtaining the real region where the object to be detected is located and the predicted region where the object to be detected is located, the target loss function can be used to calculate the real region where the object to be detected is located and the predicted region where the object to be detected is located to obtain the target loss, and the target loss is used to indicate the difference between the real region where the object to be detected is located and the predicted region where the object to be detected is located.

[0216] After obtaining the target loss, the model parameters of the model to be trained can be updated based on the target loss, and the model to be trained with the updated parameters can be trained using the next batch of training samples (i.e., re-execute steps 702 to 703) until the model training condition is met (for example, the target loss converges, etc.), and an object detection model can be obtained.

[0217] The object detection model obtained in the embodiments of the present application has the function of determining the region where the target object is located in the target detection scene. During the process of three-dimensional object detection by the object detection model, since different grid point arrays in the multiple grid point arrays have different sizes, the size of some grid point arrays can be larger than the size of the initial region, and the size of some other grid point arrays can be smaller than the size of the initial region. Therefore, based on these multiple grid point arrays, the object detection model can perform multi-scale sampling on the point cloud data within or near the initial region. In this way, the point cloud data sampled by the model can completely represent the initial region where the target object is located, so that the final region where the target object is located obtained by the model based on these point cloud data can have sufficient accuracy.

[0218] The above is a detailed description of the model training method provided in the embodiments of the present application. Next, the object detection device and the model training device provided in the embodiments of the present application will be introduced respectively. Figure 8 It is a schematic structural diagram of the object detection device provided in the embodiments of the present application. As Figure 8 shown, the device includes:

[0219] A first processing module 801, configured to process the point cloud data of the target scene to obtain a first region where the target object is located in the target scene;

[0220] A construction module 802, configured to construct multiple grid point arrays in the first region, and different grid point arrays have different sizes;

[0221] An acquisition module 803, configured to acquire a set of point cloud data of grid points in the multiple grid point arrays, and the set of point cloud data includes the point cloud data around the grid points;

[0222] A second processing module 804, configured to process the set of point cloud data to obtain a second region where the target object is located, and the second region represents the detection result of the target object.

[0223] In a possible implementation, a construction module 802 is configured to: determine the position of a target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, where the target grid point array is any one of multiple grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; and repeat the above steps for the remaining grid point arrays other than the target grid point array among the multiple grid point arrays until multiple grid point arrays are constructed in the first region.

[0224] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

[0225] In a possible implementation, an acquisition module 803 is configured to: obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of multiple grid point arrays; determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array, where the target grid point is any one of the grid points in the target grid point array; obtain the point cloud data within the sampling range of the target grid point to obtain the point cloud data set of the target grid point; and repeat the above steps for the remaining grid points other than the target grid point among the multiple grid point arrays to obtain the point cloud data sets of all grid points in the multiple grid point arrays.

[0226] In a possible implementation, the sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0227] In a possible implementation, among the multiple grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array.

[0228] In a possible implementation, the second processing module 804 is configured to: perform first feature extraction processing on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any grid points in the target grid point array, and the target grid point array is any one of multiple grid point arrays; repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid points except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; perform weighted summation processing on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the second features of all the grid points in the multiple grid point arrays; perform second feature extraction processing on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the target object is located.

[0229] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0230] In a possible implementation, the second processing module 804 is configured to: perform linear transformation processing on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; perform feature extraction processing and linear transformation processing on the target point cloud data to obtain the fourth feature of the target point cloud data; perform feature extraction processing on the target point cloud data to obtain the fifth feature of the target point cloud data; perform fusion processing on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

[0231] In a possible implementation, the fusion processing includes at least one of addition processing, multiplication processing, and mapping processing.

[0232] Figure 9 This is a schematic structural diagram of the model training device provided in the embodiment of the present application. As Figure 9 shown, the device includes:

[0233] An acquisition module 901, configured to acquire the point cloud data of the target scene and the real region where the object to be detected in the target scene is located;

[0234] A processing module 902, configured to input point cloud data of a target scene into a model to be trained, and obtain a second region where an object to be detected is located. The model to be trained is configured to: process the point cloud data of the target scene to obtain a first region where the object to be detected is located in the target scene; construct a plurality of grid point arrays in the first region, where different grid point arrays have different sizes; obtain a set of point cloud data of grid points in the plurality of grid point arrays, where the set of point cloud data includes the point cloud data around the grid points; process the set of point cloud data to obtain the second region where the object to be detected is located.

[0235] A training module 903, configured to train the model to be trained based on a real region and the second region, so as to obtain an object detection model.

[0236] In a possible implementation, the model to be trained is configured to: determine the position of a target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, where the target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; for the remaining grid point arrays other than the target grid point array in the plurality of grid point arrays, repeat the above steps until a plurality of grid point arrays are constructed in the first region.

[0237] In a possible implementation, the parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

[0238] In a possible implementation, the model to be trained is configured to: obtain a sampling radius of the target grid point array based on the distribution of the point cloud data in the first region, where the sampling radius of the target grid point array is different from the sampling radii of other grid point arrays, and the target grid point array is any one of the plurality of grid point arrays; determine a sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array, where the target grid point is any one of the grid points in the target grid point array; obtain the point cloud data in the sampling range of the target grid point to obtain a set of point cloud data of the target grid point; for the remaining grid points other than the target grid point in the plurality of grid point arrays, repeat the above steps to obtain a set of point cloud data of all grid points in the plurality of grid point arrays.

[0239] In a possible implementation, the sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

[0240] In a possible implementation, in a plurality of grid point arrays, the sampling radius of a grid point array is positively correlated with the size of the grid point array.

[0241] In a possible implementation, the model to be trained is used to: perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any grid points in the target grid point array, and the target grid point array is any one of the plurality of grid point arrays; for the remaining point cloud data in the point cloud data set of the target grid points other than the target point cloud data, repeat the above steps to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; for the remaining grid points in the plurality of grid point arrays other than the target grid points, repeat the above steps to obtain the second features of all the grid points in the plurality of grid point arrays; perform a second feature extraction process on the second features of all the grid points in the plurality of grid point arrays to obtain the second region where the object to be detected is located.

[0242] In a possible implementation, the weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

[0243] In a possible implementation, the model to be trained is used to: perform a linear transformation process on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; perform a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; perform a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; perform a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

[0244] In a possible implementation, the fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

[0245] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device modules / units, since they are based on the same concept as the method embodiments of the present application, the technical effects brought by them are the same as those of the method embodiments of the present application. The specific content can be referred to the description in the method embodiments shown in the foregoing of the embodiments of the present application, and will not be elaborated here.

[0246] The embodiments of the present application also relate to an execution device. Figure 10 This is a schematic structural diagram of the execution device provided by the embodiments of the present application. As Figure 10As shown, the execution device 1000 can specifically be embodied as a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a server, etc., which is not limited herein. Among them, an Figure 8 object detection device described in the corresponding embodiment can be deployed on the execution device 1000 to implement Figure 4 the function of object detection in the corresponding embodiment. Specifically, the execution device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003, and a memory 1004 (where the number of processors 1003 in the execution device 1000 can be one or more, Figure 10 and one processor is taken as an example here). Among them, the processor 1003 can include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003, and the memory 1004 can be connected through a bus or other means.

[0247] The memory 1004 can include a read-only memory and a random access memory, and provide instructions and data to the processor 1003. A part of the memory 1004 can also include a non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions can include various operation instructions for implementing various operations.

[0248] The processor 1003 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system. Among them, the bus system can include a power bus, a control bus, a status signal bus, etc. in addition to a data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.

[0249] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1003. The processor 1003 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed through the integrated logic circuit in hardware or instructions in software form in the processor 1003. The above-mentioned processor 1003 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1003 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1004, and the processor 1003 reads the information in the memory 1004 and combines its hardware to complete the steps of the above method.

[0250] The receiver 1001 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1002 can be used to output digital or character information through the first interface; the transmitter 1002 can also be used to send instructions to the disk array through the first interface to modify the data in the disk array; the transmitter 1002 can also include a display device such as a display screen.

[0251] In an embodiment of the present application, in one case, the processor 1003 is used to Figure 4 process the point cloud data of the target scene through the object detection model in the corresponding embodiment.

[0252] The embodiment of the present application also relates to a training device, Figure 11 which is a schematic structural diagram of the training device provided by the embodiment of the present application. As Figure 11As shown, the training device 1100 is implemented by one or more servers. The training device 1100 may vary significantly due to configuration or performance differences and may include one or more central processing units (CPUs) 1114 (e.g., one or more processors) and a memory 1132, and one or more storage media 1130 (e.g., one or more mass storage devices) for storing application programs 1142 or data 1144. Among them, the memory 1132 and the storage media 1130 may be transient storage or persistent storage. The program stored in the storage media 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1114 may be configured to communicate with the storage media 1130 and execute a series of instruction operations in the storage media 1130 on the training device 1100.

[0253] The training device 1100 may further include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, and one or more input / output interfaces 1158; or, one or more operating systems 1141, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0254] Specifically, the training device can execute Figure 7 the model training method in the corresponding embodiment.

[0255] The embodiment of the present application also relates to a computer storage medium. A program for signal processing is stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the steps executed by the foregoing execution device, or causes the computer to execute the steps executed by the foregoing training device.

[0256] The embodiment of the present application also relates to a computer program product. The computer program product stores instructions that, when executed by a computer, cause the computer to execute the steps executed by the foregoing execution device, or cause the computer to execute the steps executed by the foregoing training device.

[0257] The execution device, training device or terminal device provided by the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiments, or to cause the chip in the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0258] Specifically, please refer to Figure 12 , Figure 12 which is a schematic structural diagram of the chip provided by the embodiments of the present application. The chip may be embodied as a neural network processor NPU 1200, and the NPU 1200 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are allocated by the Host CPU. The core part of the NPU is the arithmetic circuit 1203, and the arithmetic circuit 1203 is controlled by the controller 1204 to extract matrix data from the memory and perform multiplication operations.

[0259] In some implementations, the arithmetic circuit 1203 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1203 is a two-dimensional systolic array. The arithmetic circuit 1203 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1203 is a general matrix processor.

[0260] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1202 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1201 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 1208.

[0261] The unified memory 1206 is used to store input data and output data. The weight data directly passes through the Direct Memory Access Controller (DMAC) 1205 and is transferred to the weight memory 1202 by the DMAC. The input data is also transferred to the unified memory 1206 by the DMAC.

[0262] The BIU is the Bus Interface Unit, i.e., the bus interface unit 1213, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1209.

[0263] The bus interface unit 1213 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1209 to obtain instructions from the external memory, and is also used for the storage unit access controller 1205 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0264] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1206, or transfer the weight data to the weight memory 1202, or transfer the input data to the input memory 1201.

[0265] The vector calculation unit 1207 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit 1203, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for neural network non-convolution / full connection layer network calculations, such as Batch Normalization, pixel-level summation, upsampling of the predicted label plane, etc.

[0266] In some implementations, the vector calculation unit 1207 can store the processed output vector in the unified memory 1206. For example, the vector calculation unit 1207 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1203, such as performing linear interpolation on the predicted label plane extracted by the convolutional layer, or, for example, a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 1207 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1203, for example, for use in subsequent layers in the neural network.

[0267] The instruction fetch buffer 1209 connected to the controller 1204 is used to store the instructions used by the controller 1204;

[0268] The unified memory 1206, the input memory 1201, the weight memory 1202, and the fetch memory 1209 are all On-Chip memories. The external memory is private to the NPU hardware architecture.

[0269] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.

[0270] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0271] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of this application.

[0272] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0273] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. An object detection method, characterized in that, The method is implemented by an object detection model, and the method includes: Processing the point cloud data of the target scene to obtain a first region where the target object is located in the target scene; Constructing a plurality of grid point arrays in the first region, where different grid point arrays have different sizes; Obtaining a set of point cloud data of the grid points in the plurality of grid point arrays, where the set of point cloud data includes the point cloud data around the grid points; Processing the set of point cloud data to obtain a second region where the target object is located, and the second region represents the detection result of the target object.

2. The method according to claim 1, characterized in that The constructing a plurality of grid point arrays in the first region includes: Based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region, determining the position of the target grid point, where the target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; Based on the positions of all the grid points in the target grid point array, constructing the target grid point array in the first region; Repeating the above steps for the remaining grid point arrays in the plurality of grid point arrays except the target grid point array until the plurality of grid point arrays are constructed in the first region.

3. The method according to claim 2, wherein The parameters of the first region include the size of the first region, the central position of the first region, and the yaw angle of the first region.

4. The method according to claim 1, wherein The obtaining a set of point cloud data of the grid points in the plurality of grid point arrays includes: Based on the distribution of the point cloud data in the first region, obtaining the sampling radius of the target grid point array, where the sampling radius of the target grid point array is different from that of other grid point arrays, and the target grid point array is any one of the plurality of grid point arrays; Based on the target grid point and the sampling radius of the target grid point array, determining the sampling range of the target grid point, where the target grid point is any one of the grid points in the target grid point array; Obtaining the point cloud data in the sampling range of the target grid point to obtain the set of point cloud data of the target grid point; Repeating the above steps for the remaining grid points in the plurality of grid point arrays except the target grid point to obtain the set of point cloud data of all the grid points in the plurality of grid point arrays.

5. The method according to claim 4, wherein The sampling range of the target grid point is a sphere, with the target grid point as the center of the sphere and the sum of the sampling radius of the target grid point array and a preset parameter as the radius of the sphere.

6. The method according to claim 4 or 5, characterized in that, In the plurality of grid point arrays, the sampling radius of the grid point array is positively correlated with the size of the grid point array.

7. The method according to claim 1, characterized in that, The processing the set of point cloud data to obtain the second region where the target object is located includes: Perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data. The target point cloud data is any one of the point cloud data sets of the target grid points, the target grid points are any one of the grid points in the target grid point array, and the target grid point array is any one of the multiple grid point arrays; Repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid points except the target point cloud data to obtain the first features of all the point cloud data in the point cloud data set of the target grid points; Perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid points to obtain the second feature of the target grid point; Repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid points to obtain the second features of all the grid points in the multiple grid point arrays; Perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays to obtain the second region where the target object is located.

8. The method according to claim 7, wherein The weight of the target point cloud data is determined based on the distance between the target point cloud data and the target grid point and / or the sampling radius of the target grid point array.

9. The method according to claim 7 or 8, characterized in that The performing a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data includes: Perform a linear transformation process on the distance between the target point cloud data and the target grid point to obtain the third feature of the target point cloud data; Perform a feature extraction process and a linear transformation process on the target point cloud data to obtain the fourth feature of the target point cloud data; Perform a feature extraction process on the target point cloud data to obtain the fifth feature of the target point cloud data; Perform a fusion process on the third feature, the fourth feature, and the fifth feature to obtain the first feature of the target point cloud data.

10. The method according to claim 9, wherein The fusion process includes at least one of an addition process, a multiplication process, and a mapping process.

11. A model training method, characterized in that, The method includes: Obtain the point cloud data of the target scene and the real region where the object to be detected is located in the target scene; Input the point cloud data of the target scene into the model to be trained to obtain the second region where the object to be detected is located. The model to be trained is used for: processing the point cloud data of the target scene to obtain the first region where the object to be detected is located in the target scene; constructing multiple grid point arrays in the first region, where different grid point arrays have different sizes; obtaining the point cloud data sets of the grid points in the multiple grid point arrays, and the point cloud data sets include the point cloud data around the grid points; processing the point cloud data sets to obtain the second region where the object to be detected is located; Train the model to be trained based on the real region and the second region to obtain an object detection model.

12. The method according to claim 11, wherein The model to be trained is used for: Determine the position of the target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region. The target grid point array is any one of the multiple grid point arrays, and the target grid point is any one of the grid points in the target grid point array; Construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; Repeat the above steps for the remaining grid point arrays in the multiple grid point arrays except the target grid point array until the multiple grid point arrays are constructed in the first region.

13. The method according to claim 11, wherein The model to be trained is used for: Obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region. The sampling radius of the target grid point array is different from that of other grid point arrays. The target grid point array is any one of the multiple grid point arrays; Determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array. The target grid point is any one of the grid points in the target grid point array; Obtain the point cloud data in the sampling range of the target grid point to obtain the point cloud data set of the target grid point; Repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the point cloud data sets of all grid points in the multiple grid point arrays.

14. The method according to claim 11, wherein The model to be trained is used for: Perform first feature extraction processing on the target point cloud data to obtain the first feature of the target point cloud data. The target point cloud data is any one of the point cloud data sets of the target grid point. The target grid point is any one of the grid points in the target grid point array. The target grid point array is any one of the multiple grid point arrays; Repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid point except the target point cloud data to obtain the first features of all point cloud data in the point cloud data set of the target grid point; Perform weighted summation processing on the first features of all point cloud data in the point cloud data set of the target grid point to obtain the second feature of the target grid point; Repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point to obtain the second features of all grid points in the multiple grid point arrays; Perform second feature extraction processing on the second features of all grid points in the multiple grid point arrays to obtain the second region where the object to be detected is located.

15. An object detection device, characterized in that, The device includes: A first processing module for processing the point cloud data of the target scene to obtain the first region where the target object is located in the target scene; A construction module for constructing multiple grid point arrays in the first region, where different grid point arrays have different sizes; An acquisition module, configured to acquire a point cloud data set of grid points in the plurality of grid point arrays, where the point cloud data set includes point cloud data around the grid points; A second processing module, configured to process the point cloud data set to obtain a second area where the target object is located, and the second area represents the detection result of the target object.

16. The device according to claim 15, characterized in that, The construction module is configured to: Based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first area, the number of grid points in the target grid point array, and the parameters of the first area, determine the position of the target grid point, where the target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; Based on the positions of all grid points in the target grid point array, construct the target grid point array in the first area; Repeat the above steps for the remaining grid point arrays in the plurality of grid point arrays except the target grid point array until the plurality of grid point arrays are constructed in the first area.

17. The device according to claim 15, characterized in that, The acquisition module is configured to: Based on the distribution of the point cloud data in the first area, acquire the sampling radius of the target grid point array, where the sampling radius of the target grid point array is different from the sampling radii of other grid point arrays, and the target grid point array is any one of the plurality of grid point arrays; Based on the target grid point and the sampling radius of the target grid point array, determine the sampling range of the target grid point, where the target grid point is any one of the grid points in the target grid point array; Acquire the point cloud data in the sampling range of the target grid point to obtain the point cloud data set of the target grid point; Repeat the above steps for the remaining grid points in the plurality of grid point arrays except the target grid point to obtain the point cloud data sets of all grid points in the plurality of grid point arrays.

18. The device according to claim 15, characterized in that, The second processing module is configured to: Perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data, where the target point cloud data is any one of the point cloud data sets of the target grid point, the target grid point is any one of the grid points in the target grid point array, and the target grid point array is any one of the plurality of grid point arrays; Repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid point except the target point cloud data to obtain the first features of all point cloud data in the point cloud data set of the target grid point; Perform a weighted summation process on the first features of all point cloud data in the point cloud data set of the target grid point to obtain the second feature of the target grid point; Repeat the above steps for the remaining grid points in the plurality of grid point arrays except the target grid point to obtain the second features of all grid points in the plurality of grid point arrays; Perform a second feature extraction process on the second features of all grid points in the plurality of grid point arrays to obtain the second area where the target object is located.

19. A model training device, characterized in that, The device includes: An acquisition module, configured to acquire point cloud data of a target scene and the real region where the object to be detected is located in the target scene; A processing module, configured to input the point cloud data of the target scene into a model to be trained, and obtain a second region where the object to be detected is located. The model to be trained is used to: process the point cloud data of the target scene to obtain a first region where the object to be detected is located in the target scene; construct a plurality of grid point arrays in the first region, and different grid point arrays have different sizes; obtain a set of point cloud data of the grid points in the plurality of grid point arrays, where the set of point cloud data includes the point cloud data around the grid points; process the set of point cloud data to obtain a second region where the object to be detected is located; A training module, configured to train the model to be trained based on the real region and the second region to obtain an object detection model.

20. The device according to claim 19, characterized in that The model to be trained is used to: Determine the position of a target grid point based on the serial number of the target grid point in the target grid point array, the ratio between the size of the target grid point array and the size of the first region, the number of grid points in the target grid point array, and the parameters of the first region. The target grid point array is any one of the plurality of grid point arrays, and the target grid point is any one of the grid points in the target grid point array; Construct the target grid point array in the first region based on the positions of all grid points in the target grid point array; Repeat the above steps for the remaining grid point arrays in the plurality of grid point arrays except the target grid point array until the plurality of grid point arrays are constructed in the first region.

21. The device according to claim 19, characterized in that, The model to be trained is used to: Obtain the sampling radius of the target grid point array based on the distribution of the point cloud data in the first region. The sampling radius of the target grid point array is different from that of other grid point arrays. The target grid point array is any one of the plurality of grid point arrays; Determine the sampling range of the target grid point based on the target grid point and the sampling radius of the target grid point array. The target grid point is any one of the grid points in the target grid point array; Obtain the point cloud data in the sampling range of the target grid point to obtain a set of point cloud data of the target grid point; Repeat the above steps for the remaining grid points in the plurality of grid point arrays except the target grid point to obtain a set of point cloud data of all grid points in the plurality of grid point arrays.

22. The device according to claim 19, characterized in that, The model to be trained is used to perform a first feature extraction process on the target point cloud data to obtain the first feature of the target point cloud data. The target point cloud data is any one of the set of point cloud data of the grid points in the target grid point array. The target grid point is any one of the grid points in the target grid point array. The target grid point array is any one of the plurality of grid point arrays; Repeat the above steps for the remaining point cloud data in the point cloud data set of the target grid point except the target point cloud data, to obtain the first feature of all the point cloud data in the point cloud data set of the target grid point; Perform a weighted summation process on the first features of all the point cloud data in the point cloud data set of the target grid point, to obtain the second feature of the target grid point; Repeat the above steps for the remaining grid points in the multiple grid point arrays except the target grid point, to obtain the second features of all the grid points in the multiple grid point arrays; Perform a second feature extraction process on the second features of all the grid points in the multiple grid point arrays, to obtain the second region where the object to be detected is located.

23. An object detection device, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, and when the code is executed, the object detection device executes the method according to any one of claims 1 to 14.

24. A device, characterized in that, The device is a vehicle or a wearable device or a mobile terminal, and the device includes the object detection device according to claim 23.

25. A computer storage medium, characterized in that, The computer storage medium stores one or more instructions, and when the instructions are executed by one or more computers, the one or more computers implement the method according to any one of claims 1 to 14.

26. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Object detection method and device, network device and storage medium

    CN110032962A

  • Target detection method and system based on laser radar

    CN111444839A