Point cloud data processing method and device
By screening and detecting point cloud data collected by radar devices based on the effective perception range information of the target scene, the problem of waste and low efficiency of computing resources in point cloud data processing is solved, and more efficient computing resource utilization and more accurate obstacle identification are achieved.
Patent Information
- Application Number
- CN202010713989.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-22
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-07-22
AI Technical Summary
In application scenarios such as autonomous driving, point cloud data processing requires a lot of computing resources, but not all point cloud data is effective for positioning and obstacle identification, resulting in low computing efficiency and low resource utilization.
By obtaining the point cloud data to be processed by the radar device in the target scenario, the target point cloud data is filtered out based on the effective perception range information of the target scenario and detected it, reducing the calculation amount and improving computing efficiency and resource utilization.
By filtering out effective target point cloud data, unnecessary calculations are reduced, computing efficiency and resource utilization are improved, and support for vehicle positioning and obstacle identification is enhanced.
Smart Images

Figure CN113971694B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information processing technology, and in particular to a point cloud data processing method and device. Background Art
[0002] With the development of science and technology, LiDAR has been widely used in the fields of autonomous driving, drone exploration, and mapping due to its precise distance measurement capability. Taking autonomous driving as an example, in the application scenario of autonomous driving, the point cloud data collected by LiDAR is generally processed to achieve vehicle positioning and obstacle identification. When processing point cloud data, more computing resources are generally consumed. However, since the computing resources of electronic devices that process point cloud data are limited, and not all point cloud data are useful for vehicle positioning and obstacle identification, this calculation method has low computational efficiency and low utilization of computing resources. Summary of the invention
[0003] The embodiments of the present disclosure at least provide a point cloud data processing method and device.
[0004] In a first aspect, an embodiment of the present disclosure provides a point cloud data processing method, comprising:
[0005] Obtaining the point cloud data to be processed obtained by scanning the radar device in the target scene;
[0006] Filtering target point cloud data from the point cloud data to be processed according to the effective perception range information corresponding to the target scene;
[0007] The target point cloud data is detected to obtain a detection result.
[0008] Based on the above method, the point cloud data to be processed collected by the radar device in the target scene can be screened based on the effective perception range information corresponding to the target scene. The screened target point cloud data is the target point cloud data corresponding to the target scene. Therefore, based on the screened point cloud data, detection calculation is performed in the target scene, which can reduce the amount of calculation, improve the calculation efficiency, and the utilization rate of computing resources in the target scene.
[0009] In a possible implementation manner, the effective perception range information corresponding to the target scene is determined according to the following method:
[0010] Obtain computing resource information of a processing device;
[0011] Based on the computing resource information, the effective perception range information matching the computing resource information is determined.
[0012] In this way, different effective perception range information can be determined for different electronic devices that process the point cloud data to be processed in the same target scene, so that it can be adapted to different electronic devices.
[0013] In a possible implementation, the target point cloud data is screened out from the point cloud data to be processed according to the effective perception range information corresponding to the target scene, including:
[0014] Determine a valid coordinate range based on the effective sensing range information;
[0015] Based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed, target point cloud data is filtered out from the point cloud data to be processed.
[0016] In a possible implementation manner, determining the effective coordinate range based on the effective perception range information includes:
[0017] Based on the coordinate information of the reference position point within the effective perception range in the target scene and the position information of the reference position point within the effective perception range, the effective coordinate range corresponding to the target scene is determined.
[0018] In a possible implementation manner, the step of filtering out target point cloud data from the point cloud data to be processed based on the valid coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed includes:
[0019] The radar scanning points whose corresponding coordinate information is within the effective coordinate range are used as the radar scanning points in the target point cloud data.
[0020] In a possible implementation manner, the coordinate information of the reference position point is determined according to the following steps:
[0021] Acquire location information of an intelligent driving device on which the radar device is installed;
[0022] Determine the road type of the road where the intelligent driving device is located based on the location information of the intelligent driving device;
[0023] The coordinate information of the reference position point matching the road type is obtained.
[0024] Here, the point cloud data that needs to be processed when the intelligent driving device is located on roads of different road types may be different. Therefore, by obtaining the coordinate information of the reference position point that matches the road type, the effective coordinate range that is adapted to the current road type on which the intelligent driving device is located can be determined, thereby filtering out the point cloud data under the corresponding road type, thereby improving the accuracy of the detection results of the intelligent driving device on different road types.
[0025] In a possible implementation, the detection result includes a position of the object to be identified in the target scene;
[0026] The detecting the target point cloud data to obtain a detection result includes:
[0027] The target point cloud data is rasterized to obtain a grid matrix; the value of each element in the grid matrix is used to indicate whether there is a point cloud point at the corresponding grid;
[0028] Generate a sparse matrix corresponding to the object to be identified according to the grid matrix and size information of the object to be identified in the target scene;
[0029] Based on the generated sparse matrix, the position of the object to be identified in the target scene is determined.
[0030] In a possible implementation manner, generating a sparse matrix corresponding to the object to be identified according to the grid matrix and the size information of the object to be identified in the target scene includes:
[0031] According to the grid matrix and the size information of the object to be identified in the target scene, performing at least one dilation operation or erosion operation on the target element in the grid matrix to generate a sparse matrix corresponding to the object to be identified;
[0032] The target element is an element representing the existence of a point cloud point at the corresponding grid.
[0033] In a possible implementation, according to the grid matrix and the size information of the object to be identified in the target scene, performing at least one expansion operation or corrosion operation on the target element in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including:
[0034] The target elements in the grid matrix are subjected to at least one shift processing and logic operation processing to obtain a sparse matrix corresponding to the object to be identified, wherein the difference between the coordinate range size of the obtained sparse matrix and the size size of the object to be identified in the target scene is within a preset threshold range.
[0035] In a possible implementation, based on the grid matrix and the size information of the object to be identified in the target scene, performing at least one expansion operation on the elements in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including:
[0036] Performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain a grid matrix after the first inversion operation;
[0037] Performing at least one convolution operation on the grid matrix after the first inversion operation based on a first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene;
[0038] A second inversion operation is performed on the elements in the grid matrix with preset sparsity after the at least one convolution operation to obtain the sparse matrix.
[0039] In a possible implementation manner, performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain the grid matrix after the first inversion operation includes:
[0040] Based on the second preset convolution kernel, a convolution operation is performed on the other elements in the grid matrix before the current dilation operation except the target element to obtain a first negated element, and based on the second preset convolution kernel, a convolution operation is performed on the target element in the grid matrix before the current dilation operation to obtain a second negated element;
[0041] Based on the first negated element and the second negated element, a grid matrix after a first negation operation is obtained.
[0042] In a possible implementation, performing at least one convolution operation on the grid matrix after the first inversion operation based on the first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation includes:
[0043] For the first convolution operation, convolution operation is performed on the grid matrix after the first inversion operation and the first preset convolution kernel to obtain a grid matrix after the first convolution operation;
[0044] Determining whether the sparsity of the grid matrix after the first convolution operation reaches a preset sparsity;
[0045] If not, the step of performing a convolution operation on the grid matrix after the previous convolution operation with the first preset convolution kernel to obtain the grid matrix after the current convolution operation is executed in a loop until a grid matrix with a preset sparsity is obtained after at least one convolution operation.
[0046] In a possible implementation, the first preset convolution kernel has a weight matrix and a bias corresponding to the weight matrix; for the first convolution operation, the grid matrix after the first inversion operation is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation, including:
[0047] For the first convolution operation, selecting each grid sub-matrix from the grid matrix after the first inversion operation according to the size of the first preset convolution kernel and the preset step size;
[0048] For each selected grid sub-matrix, multiply the grid sub-matrix with the weight matrix to obtain a first operation result, and add the first operation result to the offset to obtain a second operation result;
[0049] Based on the second operation results corresponding to each of the grid sub-matrices, a grid matrix after the first convolution operation is determined.
[0050] In a possible implementation, according to the grid matrix and the size information of the object to be identified in the target scene, performing at least one corrosion operation on the elements in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including:
[0051] Performing at least one convolution operation on the grid matrix to be processed based on a third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene;
[0052] The grid matrix with preset sparsity after the at least one convolution operation is determined as a sparse matrix corresponding to the object to be identified.
[0053] In a possible implementation manner, the target point cloud data is subjected to rasterization processing to obtain a raster matrix, including:
[0054] Performing rasterization processing on the target point cloud data to obtain a raster matrix and a corresponding relationship between each element in the raster matrix and each point cloud point coordinate range information;
[0055] The step of determining the position range of the object to be identified in the target scene based on the generated sparse matrix includes:
[0056] Based on the correspondence between each element in the grid matrix and each point cloud point coordinate range information, determining the coordinate information corresponding to each target element in the generated sparse matrix;
[0057] The coordinate information corresponding to each of the target elements in the sparse matrix is combined to determine the position of the object to be identified in the target scene.
[0058] In a possible implementation manner, determining the position of the object to be identified in the target scene based on the generated sparse matrix includes:
[0059] Performing at least one convolution process on each target element in the generated sparse matrix based on the trained convolutional neural network to obtain a convolution result;
[0060] Based on the convolution result, a position of the object to be identified in the target scene is determined.
[0061] In a possible implementation manner, after detecting the target point cloud data and obtaining the detection result, the method further includes:
[0062] An intelligent driving device provided with the radar device is controlled based on the detection result.
[0063] In a second aspect, the present disclosure also provides a point cloud data processing device, including:
[0064] An acquisition module is used to acquire the point cloud data to be processed obtained by scanning the radar device in the target scene;
[0065] A screening module, used for screening target point cloud data from the to-be-processed point cloud data according to the effective perception range information corresponding to the target scene;
[0066] The detection module is used to detect the target point cloud data and obtain a detection result.
[0067] In a third aspect, an embodiment of the present disclosure further provides a computer device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the above-mentioned first aspect, or any possible implementation of the first aspect are performed.
[0068] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned first aspect or any possible implementation of the first aspect are executed.
[0069] For a description of the effects of the above-mentioned point cloud data processing device, computer equipment, and computer-readable storage medium, please refer to the description of the above-mentioned point cloud data processing method, which will not be repeated here.
[0070] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can also be obtained based on these drawings without creative work.
[0072] Figure 1 A flow chart of a point cloud data processing method provided by an embodiment of the present disclosure is shown;
[0073] Figure 2 A schematic diagram showing coordinates of various position points of a rectangular parallelepiped provided by an embodiment of the present disclosure is shown;
[0074] Figure 3 A flow chart of a method for determining the coordinate information of the reference position point provided by an embodiment of the present disclosure is shown;
[0075] Figure 4 A flow chart of a method for determining a detection result provided by an embodiment of the present disclosure is shown;
[0076] FIG5( a ) shows a schematic diagram of a grid matrix before encoding provided in Embodiment 1 of the present disclosure;
[0077] FIG5( b ) shows a schematic diagram of a sparse matrix provided in Embodiment 1 of the present disclosure;
[0078] FIG5( c ) shows a schematic diagram of a post-encoding grid matrix provided in Embodiment 1 of the present disclosure;
[0079] FIG6( a ) shows a schematic diagram of a left-shifted grid matrix provided in Embodiment 1 of the present disclosure;
[0080] FIG6( b ) shows a schematic diagram of a logical OR operation provided by the first embodiment of the present disclosure;
[0081] FIG. 7( a ) shows a schematic diagram of a grid matrix after a first inversion operation provided by the first embodiment of the present disclosure;
[0082] FIG7( b ) shows a schematic diagram of a grid matrix after a convolution operation provided in the first embodiment of the present disclosure;
[0083] Figure 8 A schematic diagram of the architecture of a point cloud data processing device provided by an embodiment of the present disclosure is shown;
[0084] Fig. 9A schematic diagram of the structure of a computer device 900 provided in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0085] In order to make the purpose, technical scheme and advantages of the embodiments of the present disclosure clearer, the technical scheme in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present disclosure.
[0086] In the related technology, when processing point cloud data, more computing resources are generally consumed, but not all collected point cloud data are useful for the required calculation results, so some unnecessary point cloud data will be involved in the calculation process, which will lead to a waste of computing resources.
[0087] Based on this, the present disclosure provides a point cloud data processing method and device, which can filter the to-be-processed point cloud data collected by the radar device in the target scene based on the effective perception range information corresponding to the target scene. The filtered target point cloud data is the corresponding valid point cloud data in the target scene. Therefore, based on the filtered target point cloud data, detection calculation is performed in the target scene, which can reduce the amount of calculation, improve the calculation efficiency, and the utilization rate of computing resources in the target scene.
[0088] The defects existing in the above solutions are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the present disclosure for the above problems below should be the contributions made by the inventor to the present disclosure during the disclosure process.
[0089] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0090] To facilitate understanding of the present embodiment, a point cloud data processing method disclosed in the present embodiment is first introduced in detail. The execution subject of the point cloud data processing method provided in the present embodiment is generally a computer device with certain computing capabilities, and the computer device includes, for example: a terminal device or a server or other processing device, and the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a personal digital assistant (Personal Digital Assistant, PDA), a computing device, a vehicle-mounted device, etc. In some possible implementations, the point cloud data processing method can be implemented by a processor calling a computer-readable instruction stored in a memory.
[0091] See also Figure 1 FIG. 1 is a flowchart of a point cloud data processing method provided by an embodiment of the present disclosure, wherein the method includes steps 101 to 103, wherein:
[0092] Step 101: Obtain the point cloud data to be processed obtained by scanning the radar device in the target scene.
[0093] Step 102: Filter out target point cloud data from the point cloud data to be processed according to the effective perception range information corresponding to the target scene.
[0094] Step 103: Detect the target point cloud data to obtain a detection result.
[0095] The following is a detailed description of the above steps 101 to 103.
[0096] The radar device can be deployed on an intelligent driving device. During the driving process of the intelligent driving device, the radar device can scan and obtain point cloud data to be processed.
[0097] The effective perception range information may include coordinate thresholds in each coordinate dimension in a reference coordinate system, and the reference coordinate system is a three-dimensional coordinate system.
[0098] Exemplarily, the effective perception range information may be description information constituting a cuboid, including a maximum value x_max and a minimum value x_min in the x-axis direction, a maximum value y_max and a minimum value y_min in the y-axis direction, and a maximum value z_max and a minimum value z_min in the z-axis direction.
[0099] The coordinates of each position point constituting a cuboid based on the maximum value x_max and the minimum value x_min in the x-axis direction, the maximum value y_max and the minimum value y_min in the y-axis direction, and the maximum value z_max and the minimum value z_min in the z-axis direction can be exemplarily expressed as follows: Figure 2As shown, the coordinate origin can be the lower left vertex of the cuboid, and its coordinate value is (x_min, y_min, z_min).
[0100] In another possible implementation, the effective perception range information may also be description information of a sphere, cube, etc. For example, only the radius of a sphere or the length, width, and height of a cube may be given. The specific effective perception range information may be described according to the actual application scenario, which is not limited in the present disclosure.
[0101] In a specific implementation, since the scanning range of the radar device is limited, for example, the farthest scanning distance is 200 meters, in order to ensure the constraints of the effective perception range on the point cloud data to be processed, the constraints on the effective perception range can be set in advance. For example, the values of x_max, y_max, and z_max can be set to be less than or equal to 200 meters.
[0102] In a possible application scenario, when calculations are performed based on point cloud data, operations are performed based on the spatial voxels corresponding to the point cloud data, such as a hierarchical learning network VoxelNet based on three-dimensional spatial information of point clouds. Therefore, in this application scenario, while limiting the coordinate thresholds of the reference radar scanning points in each coordinate dimension in the reference coordinate system, the number of spatial voxels of the reference radar scanning points in each coordinate dimension can also be limited to not exceed the spatial voxel threshold.
[0103] Exemplarily, the number of spatial voxels in each coordinate dimension can be calculated by the following formula:
[0104] N_x=(x_max–x_min) / x_gridsize;
[0105] N_y=(y_max–y_min) / y_gridsize;
[0106] N_z=(z_max–z_min) / z_gridsize.
[0107] Among them, x_gridsize, y_gridsize, and z_gridsize respectively represent the preset resolutions corresponding to each dimension, N_x represents the number of spatial voxels in the x-axis direction, N_y represents the number of spatial voxels in the y-axis direction, and N_z represents the number of spatial voxels in the z-axis direction.
[0108] In another possible application scenario, when performing calculations based on point cloud data, it may also be an algorithm that performs calculations based on point cloud data within the area of a bird's-eye view, such as a point cloud-based fast target detection framework PointPillars. Therefore, the area of the bird's-eye view voxels may also be limited, for example, the value of N_x*N_y may be limited.
[0109] In a possible implementation, when determining the effective perception range information corresponding to the target scene, the effective perception range information obtained in advance based on experiments can be obtained. The effective perception range information can be used as a preset and fixed value in the target scene, and the limited perception range information also follows the above-mentioned restrictions.
[0110] In another possible implementation, when determining the effective perception range information corresponding to the target scene, the computing resource information of the processing device may be first acquired; and then based on the computing resource information, the effective perception range information matching the computing resource information may be determined.
[0111] The computing resource information includes at least one of the following information:
[0112] The memory of the central processing unit CPU, the video memory of the graphics processing unit GPU, and the computing resources of the field programmable gate array FPGA.
[0113] Specifically, when determining the effective perception range information that matches the computing resource information based on the computing resource information, the correspondence between the computing resource information of each level and the effective perception range information can be set in advance. Then, when the method provided by the present disclosure is applied to different electronic devices, the effective perception range information that matches the computing resource information of the electronic device can be searched based on the comparison relationship. Alternatively, when a change in the computing resource information of the electronic device is detected, the effective perception range information can be dynamically adjusted.
[0114] Taking the computing resource information including the memory of the central processing unit CPU as an example, the corresponding relationship between the computing resource information of each level and the effective sensing range information can be shown in the following Table 1:
[0115] Table 1
[0116]
[0117] The correspondence between the computing resource information of each level and the effective perception range information may be obtained in advance through experimental tests.
[0118] In this way, different effective perception range information can be determined for different electronic devices that process the point cloud data to be processed in the same target scene, so that it can be adapted to different electronic devices.
[0119] In a possible implementation, when filtering out target point cloud data from the point cloud data to be processed based on the effective perception range information corresponding to the target scene, the effective coordinate range can be first determined based on the effective perception range information, and then the target point cloud data can be filtered out from the point cloud data to be processed based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed.
[0120] Here, there may be two situations: one is that the effective perception range information and the effective coordinate range are both fixed, and the other is that the effective coordinate range can be changed according to the change of the effective perception range information.
[0121] For the first case, exemplarily, the effective perception range information can be the description information of a cuboid, including the length, width and height of the cuboid, with the radar device as the intersection point of the diagonals of the cuboid. Therefore, since the position of the intersection point of the diagonals of the cuboid remains unchanged, the cuboid is fixed, and the coordinate range within the cuboid is the effective coordinate range, so the effective coordinate range is also fixed.
[0122] For the second case, when determining the effective coordinate range based on the effective perception range information, the effective coordinate range corresponding to the target scene can be determined based on the coordinate information of the reference position point within the effective perception range in the target scene, and the position information of the reference position point within the effective perception range.
[0123] Exemplarily, the effective perception range information may be description information of a rectangular parallelepiped, and the reference position point may be the intersection of the diagonals of the rectangular parallelepiped. Then, as the reference position point changes, the effective perception range information will also change in different target scenes, and therefore, the corresponding effective coordinate range will also change.
[0124] The coordinate information of the reference position point may be the coordinate information of the reference position point in a radar coordinate system, and the radar coordinate system may be a three-dimensional coordinate system established with the radar device as the left origin.
[0125] If the effective perception range information is description information of a rectangular parallelepiped, the reference position point may be the intersection of the diagonals of the rectangular parallelepiped; if the effective perception range information is description information of a sphere, the reference position point may be the center of the sphere; or, the reference position point may be any reference radar scanning point within the effective perception range information.
[0126] In a specific implementation, when determining the effective coordinate range corresponding to the target scene based on the coordinate information of the reference position point and the position information of the reference position point within the effective perception range, the coordinate thresholds on each coordinate dimension in the effective perception range information in the reference coordinate system can be converted into coordinate thresholds on each coordinate dimension in the laser radar coordinate system based on the coordinate information of the reference position point in the laser radar coordinate system.
[0127] Specifically, the reference position point may correspond to first coordinate information in the reference coordinate system and to second coordinate information in the laser radar coordinate system. Based on the first coordinate information and the second coordinate information of the reference position point, the conversion relationship between the reference coordinate system and the laser radar coordinate system may be determined. Based on the conversion relationship, the coordinate thresholds of the reference radar scanning points in the effective perception range information in each coordinate dimension in the reference coordinate system may be converted into coordinate thresholds in each coordinate dimension in the laser radar coordinate system.
[0128] In another possible implementation, the relative position relationship between the threshold coordinate point corresponding to the coordinate threshold of the reference radar scanning point in the effective perception range information in each coordinate dimension in the reference coordinate system and the reference position point can be first determined, and then based on the relative position relationship, the coordinate threshold of the reference radar scanning point in the effective perception range information in each coordinate dimension in the reference coordinate system and the coordinate threshold of each coordinate dimension in the laser radar coordinate system can be determined.
[0129] Here, when the coordinate information of the reference position point changes, the coordinate thresholds of the reference radar scanning point in the effective perception range information determined based on the information on the left side of the reference position point in each coordinate dimension under the radar coordinate system will also change accordingly, that is, the effective coordinate range corresponding to the target scene will also change. Therefore, the effective coordinate range in different target scenes can be controlled by controlling the coordinate information of the reference position point.
[0130] In a possible implementation, when filtering out target point cloud data from the point cloud data to be processed based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed, the radar scanning point whose corresponding coordinate information is within the effective coordinate range can be used as the radar scanning point in the target point cloud data.
[0131] Specifically, when the radar scanning point is stored, the three-dimensional coordinate information of the radar scanning point may be stored, and then based on the three-dimensional coordinate information of the radar scanning point, it may be determined whether the radar scanning point is within a valid coordinate range.
[0132] Exemplarily, if the three-dimensional coordinate information of a radar scan point is (x, y, z), then when determining whether the radar scan point is a radar scan point in the target point cloud data, it can be determined whether the three-dimensional coordinate information of the radar scan point satisfies the following conditions:
[0133] x_min < x < x_max and y_min < y < y_max and z_min < z < z_max.
[0134] Next, the application of the above point cloud data processing method will be introduced in combination with specific application scenarios. In one possible implementation, the above point cloud data processing method can be applied to an autonomous driving scenario.
[0135] In one possible application scenario, a radar device is provided on an intelligent driving device. When determining the coordinate information of a reference position point, the coordinate information of the reference position point can be determined by the method as Figure 3 described, including the following steps:
[0136] Step 301, obtain the position information of the intelligent driving device provided with the radar device.
[0137] When obtaining the position information of the intelligent driving device, for example, it can be obtained based on the Global Positioning System (GPS). The present disclosure does not limit other ways to obtain the position information of the intelligent driving device.
[0138] Step 302, determine the road type of the road where the intelligent driving device is located based on the position information of the intelligent driving device.
[0139] In specific implementation, the road type of each section of the road within the drivable range of the intelligent driving device can be preset in advance. The road type can include, for example, intersections, T-junctions, highways, parking lots, etc. Based on the position information of the intelligent driving device, the road where the intelligent driving device is located can be determined, and then the road type of the road where the intelligent driving device is located can be determined according to the preset road type of each section of the road within the drivable range of the intelligent driving device.
[0140] Step 303, obtain the coordinate information of the reference position point matching the road type.
[0141] The locations of point cloud data that need to be processed in focus for different road types may be different. For example, if the intelligent driving device is on a highway, the point cloud data that needs to be processed by the intelligent driving device may be the point cloud data in front of the intelligent driving device. If the intelligent driving device is at an intersection, the point cloud data that needs to be processed by the intelligent driving device may be the point cloud data around the intelligent driving device. Therefore, the point cloud data for different road types can be screened by pre-setting the coordinate information of reference position points that match different road types.
[0142] Here, the point cloud data that needs to be processed when the intelligent driving device is located on roads of different road types may be different. Therefore, by obtaining the coordinate information of the reference position point that matches the road type, the effective coordinate range that is adapted to the current road type on which the intelligent driving device is located can be determined, thereby filtering out the point cloud data under the corresponding road type, thereby improving the accuracy of the filtered point cloud data.
[0143] In a possible implementation, after the target point cloud data is screened out from the point cloud data to be processed, the target point cloud data may be detected, and after the detection result is obtained, the intelligent driving device with the radar device is controlled based on the detection result.
[0144] For example, after filtering out the target point cloud data, the objects to be identified (such as obstacles) during the driving process of the intelligent driving device can be detected based on the filtered target point cloud data, and based on the detection results, the driving of the intelligent driving device equipped with a radar device can be controlled.
[0145] Controlling the intelligent driving device to travel may be controlling the intelligent driving device to accelerate, decelerate, turn, brake, etc.
[0146] With respect to step 103, in a possible implementation manner, the detection result includes the position of the object to be identified in the target scene. The process of detecting the target point cloud data will be described in detail below in conjunction with a specific embodiment. Figure 4 FIG. 1 is a flow chart of a method for determining a detection result provided by an embodiment of the present disclosure, comprising the following steps:
[0147] Step 401: rasterize the target point cloud data to obtain a grid matrix; the value of each element in the grid matrix is used to indicate whether there is a point cloud point at the corresponding grid.
[0148] Step 402: Generate a sparse matrix corresponding to the object to be identified according to the grid matrix and the size information of the object to be identified in the target scene.
[0149] Step 403: Determine the position of the object to be identified in the target scene based on the generated sparse matrix.
[0150] In the disclosed embodiment, the target point cloud data may first be subjected to rasterization processing, and then the grid matrix obtained by the rasterization processing may be subjected to sparse processing to generate a sparse matrix. The rasterization processing process here may refer to the process of mapping the spatially distributed target point cloud data containing each point cloud point into a set grid, and performing grid encoding (corresponding to a zero-one matrix) based on the point cloud points corresponding to the grid, and the sparse processing process may refer to the process of performing an expansion processing operation (corresponding to the processing result of increasing the number of elements indicated as 1 in the zero-one matrix) or an erosion processing operation (corresponding to the processing result of reducing the number of elements indicated as 1 in the zero-one matrix) on the above zero-one matrix based on the size information of the object to be identified in the target scene. Next, the above-mentioned rasterization processing process and the sparse processing process are further described.
[0151] In the process of the above-mentioned rasterization processing, the point cloud points distributed in the Cartesian continuous real number coordinate system may be converted into a rasterized discrete coordinate system.
[0152] In order to facilitate the understanding of the above-mentioned rasterization process, an example can be used for specific explanation. The disclosed embodiment has point cloud points such as point A (0.32m, 0.48m), point B (0.6m, 0.4801m) and point C (2.1m, 3.2m), and is rasterized with a grid width of 1m. The range from (0m, 0m) to (1m, 1m) corresponds to the first grid, and the range from (0m, 1m) to (1m, 2m) corresponds to the second grid, and so on. After rasterization, A'(0,0) and B'(0,0) are both in the first row and first column of the grid, and C'(2,3) can be in the second row and third column of the grid, thereby realizing the conversion from the Cartesian continuous real number coordinate system to the discrete coordinate system. Among them, the coordinate information of the point cloud points can be determined with reference to a reference point (such as the location of the radar equipment that collects the point cloud data), which will not be elaborated here.
[0153] In the embodiment of the present disclosure, two-dimensional rasterization can be performed, and three-dimensional rasterization can also be performed. Three-dimensional rasterization adds height information on the basis of two-dimensional rasterization. Next, two-dimensional rasterization can be taken as an example for specific description.
[0154] For two-dimensional rasterization, the limited space can be divided into N*M grids, generally divided at equal intervals, and the interval size can be configured. At this time, the zero-one matrix (that is, the above-mentioned grid matrix) can be used to encode the target point cloud data after rasterization. Each grid can be represented by a unique coordinate consisting of a row number and a column number. If there are one or more point cloud points in the grid, the grid is encoded as 1, otherwise it is 0, so that the encoded zero-one matrix can be obtained.
[0155] After the grid matrix is determined according to the above method, a sparse processing operation may be performed on the elements in the grid matrix according to the size information of the object to be identified in the target scene to generate a corresponding sparse matrix.
[0156] Among them, the size information of the object to be identified can be acquired in advance. Here, the size information of the object to be identified can be determined in combination with the image data synchronously collected by the target point cloud data, or the size information of the object to be identified can be roughly estimated based on the specific application scenario. For example, in the field of autonomous driving, the object in front of the vehicle can be a vehicle, and its general size information can be determined to be 4m×4m. In addition, the embodiment of the present disclosure can also determine the size information of the object to be identified based on other methods, and the embodiment of the present disclosure does not impose specific restrictions on this.
[0157] In the embodiment of the present disclosure, the sparse processing operation may be at least one expansion processing operation on the target element in the grid matrix (i.e., the element representing the existence of point cloud points at the corresponding grid). The expansion processing operation here may be performed when the coordinate range of the grid matrix is smaller than the size of the object to be identified in the target scene. That is, through one or more expansion processing operations, the range of elements representing the existence of point cloud points at the corresponding grid can be gradually expanded, so that the expanded element range can match the object to be identified, thereby achieving position determination; in addition, the sparse processing operation in the embodiment of the present disclosure may also be at least one corrosion processing operation on the target element in the grid matrix. The corrosion processing operation here may be performed when the coordinate range of the grid matrix is larger than the size of the object to be identified in the target scene. That is, through one or more corrosion processing operations, the range of elements representing the existence of point cloud points at the corresponding grid can be gradually reduced, so that the reduced element range can match the object to be identified, thereby achieving position determination.
[0158] In a specific application, whether to perform one expansion processing operation, multiple expansion processing operations, one corrosion processing operation, or multiple corrosion processing operations depends on whether the difference between the coordinate range size of the sparse matrix obtained by performing at least one shift processing and logical operation processing and the size of the object to be identified in the target scene falls within a preset threshold range, that is, the expansion or corrosion processing operation adopted in the present disclosure is performed based on the constraint of the size information of the object to be identified, so that the information represented by the determined sparse matrix is more consistent with the relevant information of the object to be identified.
[0159] It can be understood that the purpose of sparse processing, whether based on the expansion processing operation or the erosion processing operation, is to enable the generated sparse matrix to represent more accurate relevant information of the object to be identified.
[0160] In the disclosed embodiment, the above expansion processing operation can be implemented based on a shift operation and a logical OR operation, or based on convolution after negation, and then convolution and negation. The specific methods used in the two operations are different, but the effect of the sparse matrix generated in the end can be the same.
[0161] In addition, the above-mentioned corrosion processing operation can be implemented based on the shift operation and the logical AND operation, or can be implemented directly based on the convolution operation. Similarly, although the specific methods adopted by the two operations are different, the effect of the sparse matrix finally generated can also be consistent.
[0162] Next, take the expansion operation as an example, combined with Figure 5(a) to 5(b) The specific example diagram of generating a sparse matrix shown in the figure further illustrates the generation process of the above sparse matrix.
[0163] As shown in Figure 5(a), which is a schematic diagram of the grid matrix obtained after rasterization processing (corresponding to before encoding), the corresponding sparse matrix 5(b) can be obtained by performing an eight-neighborhood expansion operation on each target element in the grid matrix (corresponding to the grid with a filling effect). It can be seen that the embodiment of the present disclosure performs an eight-neighborhood expansion operation on the target elements with point cloud points at the corresponding grid in 5(a), so that each target element becomes an element set after expansion, and the grid width corresponding to the element set can match the size of the object to be identified.
[0164] Among them, the expansion operation of the above-mentioned eight-neighborhood can be a process of determining the elements whose absolute value of the difference between the horizontal coordinate or the vertical coordinate of the element does not exceed 1. Except for the elements at the edge of the grid, there are generally eight elements in the neighborhood of an element (corresponding to the above-mentioned element set). The input of the expansion processing result can be the coordinate information of the 6 target elements, and the output can be the coordinate information of the element set in the eight-neighborhood of the target element, as shown in Figure 5(b).
[0165] It should be noted that, in practical applications, in addition to the above-mentioned eight-neighborhood expansion operation, a four-neighborhood expansion operation can also be performed, and the latter other expansion operations are not specifically limited here. In addition, the embodiment of the present disclosure can also perform multiple expansion operations. For example, based on the expansion result shown in FIG5(b), the expansion operation is performed again to obtain a sparse matrix with a larger element set range, which will not be repeated here.
[0166] In the embodiment of the present disclosure, based on the generated sparse matrix, the position of the object to be identified in the target scene can be determined. In the embodiment of the present disclosure, this can be specifically implemented through the following two aspects.
[0167] First aspect: Here, the position range of the object to be identified can be determined based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point. This can be achieved through the following steps:
[0168] Step 1: Based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point, determine the coordinate information corresponding to each target element in the generated sparse matrix;
[0169] Step 2: Combine the coordinate information corresponding to each target element in the sparse matrix to determine the position of the object to be identified in the target scene.
[0170] Here, based on the above description of the rasterization process, it can be known that each target element in the grid matrix can correspond to multiple point cloud points, so that the point cloud point coordinate range information corresponding to the element and the multiple point cloud points can be predetermined. Here, still taking the N*M dimensional grid matrix as an example, the target element with point cloud points can correspond to P point cloud points, and the coordinates of each point are (Xi, Yi), i belongs to 0 to P-1, Xi, Yi represents the position of the point cloud point in the grid matrix, 0<=Xi <N,0<=Yi<M。
[0171] In this way, after generating the sparse matrix, the coordinate information corresponding to each target element in the sparse matrix can be determined based on the predetermined correspondence between the above-mentioned elements and the coordinate range information of each point cloud point, that is, the de-rasterization processing operation is performed.
[0172] It should be noted that since the sparse matrix is obtained by sparsely processing the elements in the grid matrix that represent the point cloud points at the corresponding grid, the target elements in the sparse matrix here can also represent the elements of the point cloud points at the corresponding grid.
[0173] In order to facilitate the understanding of the above-mentioned de-rasterization process, an example can be used for specific explanation. Here, point A'(0,0) indicated by the sparse matrix, point B'(0,0) is in the first row and first column of the grid; point C'(2,3) is in the second row and third column of the grid as an example. In the process of de-rasterization, the first grid (0,0) is mapped back to the Cartesian coordinate system using its center, and (0.5m, 0.5m) can be obtained. The grid (2,3) in the second row and third column is mapped back to the Cartesian coordinate system using its center, and (2.5m, 3.5m) can be obtained. That is, (0.5m, 0.5m) and (2.5m, 3.5m) can be determined as the mapped coordinate information. In this way, the mapped coordinate information can be combined to determine the position of the object to be identified in the target scene.
[0174] The disclosed embodiments can not only determine the position range of the object to be identified based on the approximate relationship between the above-mentioned sparse matrix and the target detection result, but also determine the position range of the object to be identified based on the trained convolutional neural network.
[0175] Second aspect: The disclosed embodiment can firstly perform at least one convolution process on the generated sparse matrix based on the trained convolutional neural network, and then determine the position range of the object to be identified based on the convolution result obtained by the convolution process.
[0176] In the related technology of using convolutional neural networks to realize target detection, it is necessary to traverse all the input data, find the neighborhood points of the input points in turn to perform convolution operations, and finally output the set of all domain points. The method provided by the embodiment of the present disclosure only needs to quickly traverse the target elements in the sparse matrix to find the location of the valid point (that is, the element that is 1 in the zero-one matrix) for convolution operations, thereby greatly speeding up the calculation process of the convolutional neural network and improving the efficiency of determining the position range of the object to be identified.
[0177] Considering the key role of the sparse processing operation in the point cloud data processing method provided by the embodiment of the present disclosure, it can be explained in the following two aspects.
[0178] First aspect: when the sparse processing operation is an expansion processing operation, the embodiment of the present disclosure can be implemented by combining shift processing and logical operation, and can also be implemented based on convolution after negation, and then convolution and then negation.
[0179] First, in the embodiment of the present disclosure, one or more dilation operations may be performed based on at least one shift process and a logical OR operation. In a specific implementation, the specific number of dilation operations may be determined in combination with the size information of the object to be identified in the target scene.
[0180] Here, for the first expansion processing operation, the target elements representing the point cloud points at the corresponding grid can be shifted in multiple preset directions to obtain the corresponding multiple shifted grid matrices, and then the grid matrix and the multiple shifted grid matrices corresponding to the first expansion processing operation can be logically ORed, so as to obtain a sparse matrix after the first expansion processing operation. Here, it can be determined whether the coordinate range size of the obtained sparse matrix is smaller than the size of the object to be identified, and whether the corresponding difference is large enough (such as greater than a preset threshold). If so, the target elements in the sparse matrix after the first expansion processing operation can be shifted in multiple preset directions and logically ORed according to the above method to obtain a sparse matrix after the second expansion processing operation, and so on, until it is determined that the difference between the coordinate range size of the latest sparse matrix and the size of the object to be identified in the target scene falls within the preset threshold range, and the sparse matrix is determined.
[0181] It should be noted that no matter which expansion operation is performed, the sparse matrix obtained is essentially a zero-one matrix. As the number of expansion operations increases, the number of target elements in the sparse matrix that represent the existence of point cloud points at the corresponding grid also increases. Since the grid mapped by the zero-one matrix has width information, the coordinate range size corresponding to each target element in the sparse matrix can be used to verify whether the size of the object to be identified in the target scene is reached, thereby improving the accuracy of subsequent target detection applications.
[0182] The above logical OR operation can be implemented according to the following steps:
[0183] Step 1: selecting a shifted grid matrix from a plurality of shifted grid matrices;
[0184] Step 2: Perform a logical OR operation on the grid matrix before the current dilation operation and the selected grid matrix after the shift to obtain an operation result;
[0185] Step 3: Loop and select the grid matrices that are not involved in the operation from the multiple shifted grid matrices, and perform a logical OR operation on the selected grid matrix and the most recent operation result until all grid matrices are selected to obtain a sparse matrix after the current expansion operation.
[0186] Here, first, a shifted grid matrix can be selected from multiple shifted grid matrices. In this way, the grid matrix before the current expansion processing operation can be logically ORed with the selected shifted grid matrix to obtain the operation result. Here, the grid matrix that does not participate in the operation can be cyclically selected from the multiple shifted grid matrices and participate in the logical OR operation until all the shifted grid matrices are selected, and the sparse matrix after the current expansion processing operation can be obtained.
[0187] The expansion processing operation in the embodiment of the present disclosure can be a four-neighborhood expansion centered on the target element, or an eight-neighborhood expansion centered on the target element, or other field processing operation methods. In specific applications, the corresponding field processing operation method can be selected based on the size information of the object to be identified, and no specific restrictions are made here.
[0188] It should be noted that for different field processing operation modes, the corresponding preset directions of shift processing are not the same. Taking four-field expansion as an example, the grid matrix can be shifted in four preset directions, namely left shift, right shift, up shift and down shift. Taking eight-field expansion as an example, the grid matrix can be shifted in four preset directions, namely left shift, right shift, up shift, down shift, up shift and down shift under the premise of left shift, and up shift and down shift under the premise of right shift. In addition, in order to adapt to subsequent logical or operations, after determining the shifted grid matrix based on multiple shift directions, a logical or operation can be performed first, and then the result of the logical or operation can be shifted in multiple shift directions, and then the next logical or operation can be performed, and so on, until a sparse matrix after expansion processing is obtained.
[0189] To facilitate understanding of the above expansion processing operation, the grid matrix before encoding shown in Figure 5(a) can be converted into the grid matrix after encoding as shown in Figure 5(c), and then the first expansion processing operation is illustrated in combination with 6(a) to 6(b).
[0190] As shown in the grid matrix of FIG5(c), the grid matrix is a zero-one matrix, and all the positions of 1 in the matrix can represent the grid where the target element is located, and all the positions of 0 in the matrix can represent the background.
[0191] In the embodiment of the present disclosure, the neighborhood of all elements whose element values are 1 in the zero-one matrix can first be determined by using matrix shift. Four preset shift processes can be defined here, namely left shift, right shift, upward shift and downward shift. Among them, left shift means that the column coordinates corresponding to all elements whose element values are 1 in the zero-one matrix are reduced by one, as shown in Figure 6(a); right shift means that the column coordinates corresponding to all elements whose element values are 1 in the zero-one matrix are increased by one; upward shift means that the row coordinates corresponding to all elements whose element values are 1 in the zero-one matrix are reduced by one; downward shift means that the row coordinates corresponding to all elements whose element values are 1 in the zero-one matrix are increased by one.
[0192] Secondly, the disclosed embodiment can use a matrix logical OR operation to merge the results of all neighborhoods. Matrix logical OR, that is, when receiving two sets of zero-one matrix inputs of the same size, logical OR operations are performed on the zeros and ones at the same positions of the two sets of matrices in turn, and the results are combined into a new zero-one matrix as output. FIG6( b ) is a specific example of a logical OR operation.
[0193] In the specific process of realizing the logical OR, the grid matrix after the left shift, the grid matrix after the right shift, the grid matrix after the upward shift, and the grid matrix after the downward shift can be selected in turn to participate in the logical OR operation. For example, the grid matrix can be logically ORed with the grid matrix after the left shift, and the obtained operation result can be logically ORed with the grid matrix after the right shift, and the obtained operation result can be logically ORed with the grid matrix after the upward shift, and the obtained operation result can be logically ORed with the grid matrix after the downward shift, so as to obtain a sparse matrix after the first expansion operation.
[0194] It should be noted that the above-mentioned selection order of the grid matrix after translation is only a specific example. In practical applications, it can also be selected in combination with other methods. Considering the symmetry of the translation operation, it can be selected here to pair up and down shifts and perform logical OR operations, and to pair left and right shifts and perform logical operations. The two logical OR operations can be performed simultaneously, which can save calculation time.
[0195] Second, in the embodiment of the present disclosure, the dilation operation can be implemented by combining convolution and two inversion processes, which can be implemented by the following steps:
[0196] Step 1: performing a first inversion operation on the elements in the grid matrix before the current expansion operation to obtain a grid matrix after the first inversion operation;
[0197] Step 2: performing at least one convolution operation on the grid matrix after the first inversion operation based on the first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be identified in the target scene;
[0198] Step three: performing a second inversion operation on the elements in the grid matrix with preset sparsity after at least one convolution operation to obtain a sparse matrix.
[0199] The disclosed embodiment can implement the expansion processing operation by performing convolution after inversion and then convolution and then inversion. The obtained sparse matrix can also represent the relevant information of the object to be identified to a certain extent. In addition, considering that the above-mentioned convolution operation can be automatically combined with the convolutional neural network used in subsequent target detection and other applications, the detection efficiency can be improved to a certain extent.
[0200] In the disclosed embodiment, the negation operation can be implemented based on a convolution operation, or can be implemented based on other negation operation modes. In order to facilitate the cooperation with subsequent application networks (such as convolutional neural networks used for target detection), here, a convolution operation can be used for specific implementation, and the above-mentioned first negation operation is specifically described below.
[0201] Here, a convolution operation can be performed on the elements other than the target element in the grid matrix before the current expansion processing operation based on the second preset convolution kernel to obtain a first negated element. A convolution operation can also be performed on the target element in the grid matrix before the current expansion processing operation based on the second preset convolution kernel to obtain a second negated element. Based on the above first negated element and second negated element, the grid matrix after the first negation operation can be determined.
[0202] The implementation process of the second inversion operation can refer to the implementation process of the first inversion operation mentioned above, which will not be repeated here.
[0203] In the embodiment of the present disclosure, the first preset convolution kernel can be used to perform at least one convolution operation on the grid matrix after the first inversion operation, so as to obtain a grid matrix with a preset sparsity. If the expansion operation can be used as a means to expand the number of target elements in the grid matrix, the above convolution operation can be regarded as a process of reducing the number of target elements in the grid matrix (corresponding to the corrosion operation). Since the convolution operation in the embodiment of the present disclosure is performed on the grid matrix after the first inversion operation, the inversion operation is combined with the corrosion operation, and then the inversion operation is performed again to achieve an equivalent operation equivalent to the above expansion operation.
[0204] Among them, for the first convolution operation, the grid matrix after the first inversion operation is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation. After it is determined that the sparsity of the grid matrix after the first convolution operation does not reach the preset sparsity, the grid matrix after the first convolution operation can be convolved with the first preset convolution kernel again to obtain the grid matrix after the second convolution operation, and so on, until a grid matrix with a preset sparsity can be determined.
[0205] Among them, the above-mentioned sparsity can be determined by the proportion distribution of target elements and non-target elements in the grid matrix. The more the proportion of target elements, the larger the size information of the object to be identified that it represents. Conversely, the smaller the proportion of target elements, the smaller the size information of the object to be identified that it represents. The embodiment of the present disclosure can stop the convolution operation when the proportion distribution reaches a preset sparsity.
[0206] The convolution operation in the embodiment of the present disclosure may be performed once or multiple times. Here, the specific operation process of the first convolution operation is described, including the following steps:
[0207] Step 1: for the first convolution operation, select each grid sub-matrix from the grid matrix after the first inversion operation according to the size of the first preset convolution kernel and the preset step size;
[0208] Step 2: for each selected grid sub-matrix, multiply the grid sub-matrix with the weight matrix to obtain a first operation result, and add the first operation result to the offset to obtain a second operation result;
[0209] Step 3: Based on the second operation results corresponding to each grid sub-matrix, determine the grid matrix after the first convolution operation.
[0210] Here, the grid matrix after the first inversion operation can be traversed in a traversal manner. In this way, for each traversed grid sub-matrix, the grid sub-matrix and the weight matrix can be multiplied to obtain a first operation result, and the first operation result and the offset can be added to obtain a second operation result. In this way, the second operation results corresponding to each grid sub-matrix are combined into the corresponding matrix elements to obtain the grid matrix after the first convolution operation.
[0211] In order to facilitate the understanding of the above expansion processing operation, the encoded grid matrix shown in Figure 5(c) is still taken as an example here, and the expansion processing operation is illustrated in combination with Figures 7(a) to 7(b).
[0212] Here, a 1*1 convolution kernel (i.e., the second preset convolution kernel) can be used to implement the first inversion operation. The weight of the second preset convolution kernel is -1, and the bias is 1. At this time, the weight and bias are substituted into the convolution formula {output = input grid matrix * weight + bias}. If the input is the target element in the grid matrix, its value corresponds to 1, then the output = 1*-1+1=0; if the input is a non-target element in the grid matrix, its value corresponds to 0, then the output = 0*-1+1=1; in this way, after the 1*1 convolution kernel acts on the input, the zero-one matrix can be inverted, the element value 0 becomes 1, and the element value 1 becomes 0, as shown in Figure 7(a).
[0213] For the above-mentioned corrosion processing operation, in a specific application, a 3*3 convolution kernel (i.e., the first preset convolution kernel) and a linear rectifier function (Rectified Linear Unit, ReLU) can be used to implement it. The weights included in the weight matrix of the above-mentioned first preset convolution kernel are all 1, and the offset is 8. In this way, the above-mentioned corrosion processing operation can be implemented using the formula {output = ReLU (the grid matrix after the first inversion operation of the input * weight + offset)}.
[0214] Here, only when all elements in the input 3*3 grid sub-matrix are 1, the output = ReLU(9-8) = 1; otherwise, the output = ReLU(input grid sub-matrix*1-8) = 0, where (input grid sub-matrix*1-8)<0. As shown in Figure 7(b), this is the grid matrix after the convolution operation.
[0215] Here, each nested layer of the convolutional network with the second preset convolution kernel can be superimposed with an erosion operation to obtain a grid matrix with fixed sparsity. The inversion operation is equivalent to an expansion operation, thereby realizing the generation of a sparse matrix.
[0216] Second aspect: When the sparse processing operation is an erosion processing operation, the embodiment of the present disclosure can be implemented by combining shift processing and logical operations, and can also be implemented based on convolution operations.
[0217] First, in the embodiment of the present disclosure, one or more corrosion processing operations can be performed based on at least one shift processing and logical AND operation. In the specific implementation process, the specific number of corrosion processing operations can be determined in combination with the size information of the object to be identified in the target scene.
[0218] Similar to the expansion process based on the shift process and the logical OR operation in the first aspect, in the process of the corrosion process, the grid matrix may be shifted first. Unlike the above-mentioned expansion process, the logical operation here may be a logical AND operation for the shifted grid matrix. For the process of implementing the corrosion process based on the shift process and the logical AND operation, please refer to the above description, which will not be repeated here.
[0219] Similarly, the corrosion processing operation in the embodiment of the present disclosure can be a four-neighborhood corrosion centered on the target element, or an eight-neighborhood corrosion centered on the target element, or other field processing operation methods. In specific applications, the corresponding field processing operation method can be selected based on the size information of the object to be identified, and no specific restrictions are made here.
[0220] Second, in the embodiment of the present disclosure, the corrosion processing operation can be implemented in combination with the convolution processing, which can be implemented specifically through the following steps:
[0221] Step 1: performing at least one convolution operation on the grid matrix based on the third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be identified in the target scene;
[0222] Step 2: Determine the grid matrix with preset sparsity after at least one convolution operation as a sparse matrix corresponding to the object to be identified.
[0223] The above convolution operation can be regarded as a process of reducing the number of target elements in the grid matrix, that is, an erosion process. Among them, for the first convolution operation, the grid matrix is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation. After judging that the sparsity of the grid matrix after the first convolution operation does not reach the preset sparsity, the grid matrix after the first convolution operation can be convolved with the third preset convolution kernel again to obtain the grid matrix after the second convolution operation, and so on, until a grid matrix with a preset sparsity can be determined, that is, a sparse matrix corresponding to the object to be identified is obtained.
[0224] The convolution operation in the embodiment of the present disclosure may be performed once or multiple times. For the specific process of the convolution operation, please refer to the relevant description of the expansion processing based on convolution and negation in the first aspect above, which will not be repeated here.
[0225] It should be noted that in specific applications, convolutional neural networks with different data processing bit widths can be used to realize the generation of sparse matrices. For example, 4 bits can be used to represent the network input, output, and calculation parameters, such as the element value of the grid matrix (0 or 1), weights, biases, etc. In addition, 8 bits can be used for representation to adapt to the network processing bit width and improve computing efficiency.
[0226] Based on the above method, the point cloud data to be processed collected by the radar device in the target scene can be screened based on the effective perception range information corresponding to the target scene. The screened target point cloud data is the corresponding valid point cloud data in the target scene. Therefore, based on the screened target point cloud data, detection calculation is performed in the target scene, which can reduce the amount of calculation, improve the calculation efficiency, and the utilization rate of computing resources in the target scene.
[0227] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0228] Based on the same inventive concept, a point cloud data processing device corresponding to the point cloud data processing method is also provided in the embodiment of the present disclosure. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned point cloud data processing method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0229] Reference Figure 8 , which is a schematic diagram of the architecture of a point cloud data processing device provided by an embodiment of the present disclosure, the device comprises: an acquisition module 801, a screening module 802, and a detection module 803; wherein,
[0230] An acquisition module 801 is used to acquire the point cloud data to be processed obtained by scanning the radar device in the target scene;
[0231] A screening module 802 is used to screen out target point cloud data from the to-be-processed point cloud data according to the effective perception range information corresponding to the target scene;
[0232] The detection module 803 is used to detect the target point cloud data to obtain a detection result.
[0233] In a possible implementation manner, the screening module 802 is further configured to determine the effective perception range information corresponding to the target scene according to the following method:
[0234] Obtain computing resource information of a processing device;
[0235] Based on the computing resource information, the effective perception range information matching the computing resource information is determined.
[0236] In a possible implementation manner, the screening module 802, when screening out target point cloud data from the to-be-processed point cloud data according to the effective perception range information corresponding to the target scene, is configured to:
[0237] Determine a valid coordinate range based on the effective sensing range information;
[0238] Based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed, target point cloud data is filtered out from the point cloud data to be processed.
[0239] In a possible implementation, the screening module 802, when determining the effective coordinate range based on the effective perception range information, is configured to:
[0240] Based on the coordinate information of the reference position point within the effective perception range in the target scene and the position information of the reference position point within the effective perception range, the effective coordinate range corresponding to the target scene is determined.
[0241] In a possible implementation manner, the screening module 802, when screening out target point cloud data from the point cloud data to be processed based on the valid coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed, is used to:
[0242] The radar scanning points whose corresponding coordinate information is within the effective coordinate range are used as the radar scanning points in the target point cloud data.
[0243] In a possible implementation manner, the screening module 802 is further configured to determine the coordinate information of the reference position point according to the following steps:
[0244] Acquire location information of an intelligent driving device on which the radar device is installed;
[0245] Determine the road type of the road where the intelligent driving device is located based on the location information of the intelligent driving device;
[0246] The coordinate information of the reference position point matching the road type is obtained.
[0247] In a possible implementation, the detection result includes a position of the object to be identified in the target scene;
[0248] The detection module 803, when detecting the target point cloud data and obtaining the detection result, is used to:
[0249] The target point cloud data is rasterized to obtain a grid matrix; the value of each element in the grid matrix is used to indicate whether there is a point cloud point at the corresponding grid;
[0250] Generate a sparse matrix corresponding to the object to be identified according to the grid matrix and size information of the object to be identified in the target scene;
[0251] Based on the generated sparse matrix, the position of the object to be identified in the target scene is determined.
[0252] In a possible implementation manner, the detection module 803, when generating a sparse matrix corresponding to the object to be identified according to the grid matrix and the size information of the object to be identified in the target scene, is configured to:
[0253] According to the grid matrix and the size information of the object to be identified in the target scene, performing at least one dilation operation or erosion operation on the target element in the grid matrix to generate a sparse matrix corresponding to the object to be identified;
[0254] The target element is an element representing the existence of a point cloud point at the corresponding grid.
[0255] In a possible implementation manner, the detection module 803, when performing at least one dilation operation or erosion operation on the target element in the grid matrix according to the grid matrix and the size information of the object to be identified in the target scene to generate a sparse matrix corresponding to the object to be identified, is used to:
[0256] The target elements in the grid matrix are subjected to at least one shift processing and logic operation processing to obtain a sparse matrix corresponding to the object to be identified, wherein the difference between the coordinate range size of the obtained sparse matrix and the size size of the object to be identified in the target scene is within a preset threshold range.
[0257] In a possible implementation manner, the detection module 803, when performing at least one expansion operation on the elements in the grid matrix according to the grid matrix and the size information of the object to be identified in the target scene to generate a sparse matrix corresponding to the object to be identified, is used to:
[0258] Performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain a grid matrix after the first inversion operation;
[0259] Performing at least one convolution operation on the grid matrix after the first inversion operation based on a first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene;
[0260] A second inversion operation is performed on the elements in the grid matrix with preset sparsity after the at least one convolution operation to obtain the sparse matrix.
[0261] In a possible implementation manner, the detection module 803, when performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain the grid matrix after the first inversion operation, is configured to:
[0262] Based on the second preset convolution kernel, a convolution operation is performed on the other elements in the grid matrix before the current dilation operation except the target element to obtain a first negated element, and based on the second preset convolution kernel, a convolution operation is performed on the target element in the grid matrix before the current dilation operation to obtain a second negated element;
[0263] Based on the first negated element and the second negated element, a grid matrix after a first negation operation is obtained.
[0264] In a possible implementation manner, the detection module 803, when performing at least one convolution operation on the grid matrix after the first inversion operation based on the first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation, is used to:
[0265] For the first convolution operation, convolution operation is performed on the grid matrix after the first inversion operation and the first preset convolution kernel to obtain a grid matrix after the first convolution operation;
[0266] Determining whether the sparsity of the grid matrix after the first convolution operation reaches a preset sparsity;
[0267] If not, the step of performing a convolution operation on the grid matrix after the previous convolution operation with the first preset convolution kernel to obtain the grid matrix after the current convolution operation is executed in a loop until a grid matrix with a preset sparsity is obtained after at least one convolution operation.
[0268] In a possible implementation manner, the detection module 803 has a weight matrix and a bias corresponding to the weight matrix in a first preset convolution kernel; for a first convolution operation, the grid matrix after the first inversion operation is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation, which is used to:
[0269] For the first convolution operation, selecting each grid sub-matrix from the grid matrix after the first inversion operation according to the size of the first preset convolution kernel and the preset step size;
[0270] For each selected grid sub-matrix, multiply the grid sub-matrix with the weight matrix to obtain a first operation result, and add the first operation result to the offset to obtain a second operation result;
[0271] Based on the second operation results corresponding to each of the grid sub-matrices, a grid matrix after the first convolution operation is determined.
[0272] In a possible implementation manner, the detection module 803, when performing at least one corrosion processing operation on the elements in the grid matrix according to the grid matrix and the size information of the object to be identified in the target scene to generate a sparse matrix corresponding to the object to be identified, is used to:
[0273] Performing at least one convolution operation on the grid matrix to be processed based on a third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene;
[0274] The grid matrix with preset sparsity after the at least one convolution operation is determined as a sparse matrix corresponding to the object to be identified.
[0275] In a possible implementation manner, the detection module 803, when performing rasterization processing on the target point cloud data to obtain a raster matrix, is used to:
[0276] Performing rasterization processing on the target point cloud data to obtain a raster matrix and a corresponding relationship between each element in the raster matrix and each point cloud point coordinate range information;
[0277] The detection module 803, when determining the position range of the object to be identified in the target scene based on the generated sparse matrix, is used to:
[0278] Based on the correspondence between each element in the grid matrix and each point cloud point coordinate range information, determining the coordinate information corresponding to each target element in the generated sparse matrix;
[0279] The coordinate information corresponding to each of the target elements in the sparse matrix is combined to determine the position of the object to be identified in the target scene.
[0280] In a possible implementation manner, when the detection module 803 determines the position of the object to be identified in the target scene based on the generated sparse matrix, it is configured to:
[0281] Performing at least one convolution process on each target element in the generated sparse matrix based on the trained convolutional neural network to obtain a convolution result;
[0282] Based on the convolution result, a position of the object to be identified in the target scene is determined.
[0283] In a possible implementation manner, the device further includes a control module 804, configured to:
[0284] After detecting the target point cloud data and obtaining a detection result, the intelligent driving device of the radar device is controlled based on the detection result.
[0285] Based on the above device, the point cloud data to be processed collected by the radar device in the target scene can be screened based on the effective perception range information corresponding to the target scene. The screened target point cloud data is the target point cloud data corresponding to the target scene. Therefore, based on the screened point cloud data, detection calculations are performed in the target scene, which can reduce the amount of calculations, improve calculation efficiency, and utilization of computing resources in the target scene.
[0286] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.
[0287] Based on the same technical concept, the embodiment of the present disclosure also provides a computer device. Fig. 9 As shown, it is a schematic diagram of the structure of a computer device 900 provided in an embodiment of the present disclosure, including a processor 901, a memory 902, and a bus 903. Among them, the memory 902 is used to store execution instructions, including a memory 9021 and an external memory 9022; the memory 9021 here is also called an internal memory, which is used to temporarily store the operation data in the processor 901, and the data exchanged with the external memory 9022 such as a hard disk. The processor 901 exchanges data with the external memory 9022 through the memory 9021. When the computer device 900 is running, the processor 901 communicates with the memory 902 through the bus 903, so that the processor 901 executes the following instructions:
[0288] Obtaining the point cloud data to be processed obtained by scanning the radar device in the target scene;
[0289] Filtering target point cloud data from the point cloud data to be processed according to the effective perception range information corresponding to the target scene;
[0290] The target point cloud data is detected to obtain a detection result.
[0291] The present disclosure also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the point cloud data processing method described in the above method embodiment are executed. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0292] The computer program product of the point cloud data processing method provided in the embodiments of the present disclosure includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the steps of the point cloud data processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0293] The present disclosure also provides a computer program, which implements any one of the methods of the aforementioned embodiments when executed by a processor. The computer program product can be implemented in hardware, software, or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK), etc.
[0294] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0295] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0296] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0297] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0298] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed in the present disclosure, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be based on the protection scope of the claims.
Claims
1. A point cloud data processing method, It is characterized in that include: Obtaining the point cloud data to be processed obtained by scanning the radar device in the target scene; Acquire the position information of the intelligent driving device on which the radar device is installed, determine the road type of the road where the intelligent driving device is located based on the position information of the intelligent driving device, and acquire the coordinate information of the reference position point matching the road type; Determine a valid coordinate range corresponding to the target scene based on coordinate information of a reference position point within the effective perception range in the target scene and position information of the reference position point within the effective perception range; Based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed, filtering out target point cloud data from the point cloud data to be processed; The target point cloud data is detected to obtain a detection result.
2. The method according to claim 1, It is characterized in that Determine the effective perception range information corresponding to the target scene according to the following method: Obtain computing resource information of processing equipment; Based on the computing resource information, the effective perception range information matching the computing resource information is determined.
3. The method according to claim 1, It is characterized in that The step of filtering out target point cloud data from the point cloud data to be processed based on the valid coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed includes: The radar scanning points whose corresponding coordinate information is within the effective coordinate range are used as the radar scanning points in the target point cloud data.
4. The method according to claim 1, It is characterized in that The detection result includes the position of the object to be identified in the target scene; The detecting the target point cloud data to obtain a detection result includes: The target point cloud data is rasterized to obtain a grid matrix; the value of each element in the grid matrix is used to indicate whether there is a point cloud point at the corresponding grid; Generate a sparse matrix corresponding to the object to be identified according to the grid matrix and size information of the object to be identified in the target scene; Based on the generated sparse matrix, the position of the object to be identified in the target scene is determined.
5. The method according to claim 4, It is characterized in that The step of generating a sparse matrix corresponding to the object to be identified according to the grid matrix and the size information of the object to be identified in the target scene comprises: According to the grid matrix and the size information of the object to be identified in the target scene, performing at least one dilation operation or erosion operation on the target element in the grid matrix to generate a sparse matrix corresponding to the object to be identified; The target element is an element representing the existence of a point cloud point at the corresponding grid.
6. The method according to claim 5, It is characterized in that According to the grid matrix and the size information of the object to be identified in the target scene, performing at least one dilation processing operation or an erosion processing operation on the target element in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including: The target elements in the grid matrix are subjected to at least one shift processing and logic operation processing to obtain a sparse matrix corresponding to the object to be identified, wherein the difference between the coordinate range size of the obtained sparse matrix and the size size of the object to be identified in the target scene is within a preset threshold range.
7. The method according to claim 5, It is characterized in that According to the grid matrix and the size information of the object to be identified in the target scene, performing at least one expansion operation on the elements in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including: Performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain a grid matrix after the first inversion operation; Performing at least one convolution operation on the grid matrix after the first inversion operation based on a first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene; A second inversion operation is performed on the elements in the grid matrix with preset sparsity after the at least one convolution operation to obtain the sparse matrix.
8. The method according to claim 7, It is characterized in that The step of performing a first inversion operation on the elements in the grid matrix before the current dilation operation to obtain the grid matrix after the first inversion operation includes: Based on the second preset convolution kernel, a convolution operation is performed on the other elements in the grid matrix before the current dilation operation except the target element to obtain a first negated element, and based on the second preset convolution kernel, a convolution operation is performed on the target element in the grid matrix before the current dilation operation to obtain a second negated element; Based on the first negated element and the second negated element, a grid matrix after a first negation operation is obtained.
9. The method according to claim 7 or 8, It is characterized in that The method of performing at least one convolution operation on the grid matrix after the first inversion operation based on the first preset convolution kernel to obtain a grid matrix with preset sparsity after at least one convolution operation includes: For the first convolution operation, convolution operation is performed on the grid matrix after the first inversion operation and the first preset convolution kernel to obtain a grid matrix after the first convolution operation; Determining whether the sparsity of the grid matrix after the first convolution operation reaches a preset sparsity; If not, the step of performing a convolution operation on the grid matrix after the previous convolution operation with the first preset convolution kernel to obtain the grid matrix after the current convolution operation is executed in a loop until a grid matrix with a preset sparsity is obtained after at least one convolution operation.
10. The method according to claim 9, It is characterized in that The first preset convolution kernel has a weight matrix and a bias corresponding to the weight matrix; for the first convolution operation, the grid matrix after the first inversion operation is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation, including: For the first convolution operation, selecting each grid sub-matrix from the grid matrix after the first inversion operation according to the size of the first preset convolution kernel and the preset step size; For each selected grid sub-matrix, multiply the grid sub-matrix with the weight matrix to obtain a first operation result, and add the first operation result to the offset to obtain a second operation result; Based on the second operation results corresponding to each of the grid sub-matrices, a grid matrix after the first convolution operation is determined.
11. The method according to claim 5, It is characterized in that According to the grid matrix and the size information of the object to be identified in the target scene, at least one corrosion processing operation is performed on the elements in the grid matrix to generate a sparse matrix corresponding to the object to be identified, including: Performing at least one convolution operation on the grid matrix to be processed based on a third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be identified in the target scene; The grid matrix with preset sparsity after the at least one convolution operation is determined as a sparse matrix corresponding to the object to be identified.
12. The method according to claim 5, It is characterized in that The target point cloud data is rasterized to obtain a raster matrix, including: Performing rasterization processing on the target point cloud data to obtain a raster matrix and a corresponding relationship between each element in the raster matrix and each point cloud point coordinate range information; The step of determining the position range of the object to be identified in the target scene based on the generated sparse matrix includes: Based on the correspondence between each element in the grid matrix and each point cloud point coordinate range information, determining the coordinate information corresponding to each target element in the generated sparse matrix; The coordinate information corresponding to each of the target elements in the sparse matrix is combined to determine the position of the object to be identified in the target scene.
13. The method according to claim 5, It is characterized in that The step of determining the position of the object to be identified in the target scene based on the generated sparse matrix includes: Performing at least one convolution process on each target element in the generated sparse matrix based on the trained convolutional neural network to obtain a convolution result; Based on the convolution result, a position of the object to be identified in the target scene is determined.
14. The method according to claim 1, It is characterized in that After detecting the target point cloud data and obtaining the detection result, the method further includes: An intelligent driving device provided with the radar device is controlled based on the detection result.
15. A point cloud data processing device, It is characterized in that include: An acquisition module is used to acquire the point cloud data to be processed obtained by scanning the radar device in the target scene; A screening module is used to obtain the position information of the intelligent driving device on which the radar device is installed, determine the road type of the road where the intelligent driving device is located based on the position information of the intelligent driving device, and obtain the coordinate information of the reference position point matching the road type; determine the effective coordinate range corresponding to the target scene based on the coordinate information of the reference position point within the effective perception range in the target scene and the position information of the reference position point within the effective perception range; and screen out target point cloud data from the point cloud data to be processed based on the effective coordinate range and the coordinate information of each radar scanning point in the point cloud data to be processed; The detection module is used to detect the target point cloud data and obtain a detection result.
16. A computer device, It is characterized in that include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the point cloud data processing method as described in any one of claims 1 to 14 are performed.
17. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the point cloud data processing method according to any one of claims 1 to 14 are executed.
Citation Information
Patent Citations
Three-dimensional target detection method based on graph convolution attention network
CN110674829A
Target detection and tracking method, related equipment and computer readable storage medium
CN111192295A
Three-dimensional target detection method and device, computer equipment and storage medium
CN111199206A