A method, apparatus, electronic device, and storage medium for processing point cloud data

Through rasterization processing and sparse processing, the problem of customized design of point cloud data processing in the existing technology is solved, and efficient and flexible point cloud data processing is achieved.

CN113971712BActive Publication Date: 2025-05-30SHANGHAI SENSETIME LINGANG INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010712674.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-22
Publication Date
2025-05-30
Estimated Expiration
2040-07-22

AI Technical Summary

Technical Problem

In the prior art, point cloud data processing solutions need to be customized for different application environments, resulting in waste of resources and inefficiency.

Method used

Through rasterization processing and sparse processing under size information constraints, a sparse matrix is ​​automatically generated, thereby realizing scenario application and reducing processing steps and resource consumption.

Benefits of technology

It realizes efficient processing of point cloud data, reduces resource consumption and processing time, and improves processing efficiency and application flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113971712B_ABST
    Figure CN113971712B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, electronic device and storage medium for processing point cloud data. The processing method includes: performing rasterization processing on the point cloud data in the target scene to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster; generating a sparse matrix corresponding to the object to be recognized in the target scene according to the raster matrix and the size information of the object to be recognized in the target scene; and determining the position of the object to be recognized in the target scene based on the generated sparse matrix. The present disclosure realizes the automatic generation of the sparse matrix through rasterization processing and sparse processing under the constraint of size information, and performs object recognition according to the generated sparse matrix, which saves time and effort.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of point cloud data processing. Specifically, it relates to a method, device, electronic device, and storage medium for processing point cloud data. Background Art

[0002] With the continuous development of lidar technology, since the point cloud data collected by lidar includes the accurate position information of the target object, the application of lidar for point cloud data collection is widely used in various fields, such as target detection, three-dimensional target reconstruction, autonomous driving, etc. As a kind of sparse data, point cloud data usually needs to be processed to realize the above applications.

[0003] To facilitate applications, the point cloud processing solutions in the related art need to be customized using different programming languages for different application environments, which will consume a large amount of manpower and material resources. Summary of the Invention

[0004] The embodiments of the present disclosure at least provide a method, device, electronic device, and storage medium for processing point cloud data. Through rasterization processing and sparse processing under size information constraints, the automatic generation of a sparse matrix is realized, and scene applications are realized based on the generated sparse matrix, which saves time and effort.

[0005] It mainly includes the following aspects:

[0006] In a first aspect, the embodiments of the present disclosure provide a method for processing point cloud data, the method includes:

[0007] Obtain point cloud data corresponding to a target scene;

[0008] Perform rasterization processing on the obtained point cloud data to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster.

[0009] Generate a sparse matrix corresponding to the object to be recognized in the target scene according to the raster matrix and the size information of the object to be recognized in the target scene;

[0010] Based on the generated sparse matrix, determine the position of the object to be recognized in the target scene.

[0011] By using the above method for processing point cloud data, when the point cloud data corresponding to the target scene is obtained, the point cloud data can be first rasterized to obtain a raster matrix. The value of the element in the raster matrix can represent whether there is a point cloud point at the corresponding raster. In this way, according to the size information of the object to be recognized in the target scene, the elements in the raster matrix that represent the existence of point cloud points at the corresponding raster can be processed to generate a sparse matrix corresponding to the object to be recognized, so as to determine the position of the object to be recognized in the target scene according to the generated sparse matrix.

[0012] In one implementation manner, the generating a sparse matrix corresponding to the object to be recognized according to the raster matrix and the size information of the object to be recognized in the target scene includes:

[0013] Performing at least one dilation operation or erosion operation on the target elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized;

[0014] Wherein, the target elements are the elements that represent the existence of point cloud points at the corresponding raster.

[0015] In one implementation manner, performing at least one dilation operation or erosion operation on the target elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized, includes:

[0016] Performing at least one shift processing and logical operation processing on the target elements in the raster matrix, to obtain a sparse matrix corresponding to the object to be recognized, wherein the difference between the size of the coordinate range of the obtained sparse matrix and the size of the object to be recognized in the target scene belongs to a preset threshold range.

[0017] In one implementation manner, performing at least one dilation operation on the elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized, includes:

[0018] Performing a first negation operation on the elements in the raster matrix before the current dilation operation, to obtain a raster matrix after the first negation operation;

[0019] Performing at least one convolution operation on the raster matrix after the first negation operation based on a first preset convolution kernel, to obtain a raster matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene;

[0020] Perform a second negation operation on the elements in the grid matrix with a preset sparsity after the at least one convolution operation to obtain the sparse matrix.

[0021] In one implementation, the performing a first negation operation on the elements in the grid matrix before the current dilation processing operation to obtain the grid matrix after the first negation operation includes:

[0022] Performing a convolution operation on other elements in the grid matrix before the current dilation processing operation except the target element based on a second preset convolution kernel to obtain a first negated element, and performing a convolution operation on the target element in the grid matrix before the current dilation processing operation based on the second preset convolution kernel to obtain a second negated element;

[0023] Based on the first negated element and the second negated element, obtain the grid matrix after the first negation operation.

[0024] In one implementation, the performing at least one convolution operation on the grid matrix after the first negation operation based on a first preset convolution kernel to obtain a grid matrix with a preset sparsity after the at least one convolution operation includes:

[0025] For the first convolution operation, perform a convolution operation on the grid matrix after the first negation operation and the first preset convolution kernel to obtain the grid matrix after the first convolution operation;

[0026] Determine whether the sparsity of the grid matrix after the first convolution operation reaches the preset sparsity;

[0027] If not, loop to perform the step of performing a convolution operation on the grid matrix after the previous convolution operation and the first preset convolution kernel to obtain the grid matrix after the current convolution operation until a grid matrix with a preset sparsity after the at least one convolution operation is obtained.

[0028] Here, for the first convolution operation, the grid matrix after the first convolution operation can be determined based on the convolution operation between the grid matrix after the first negation operation and the first preset convolution kernel, and then the grid matrix after the second convolution operation can be determined based on the convolution operation between the grid matrix after the first convolution operation and the first preset convolution kernel, and so on, until a grid matrix with a preset sparsity is obtained.

[0029] In one implementation, the first preset convolution kernel has a weight matrix and a bias corresponding to the weight matrix; for the first convolution operation, performing a convolution operation on the grid matrix after the first negation operation and the first preset convolution kernel to obtain the grid matrix after the first convolution operation includes:

[0030] For the first convolution operation, each grid sub-matrix is selected from the raster matrix after the first negation operation according to the size of the first preset convolution kernel and the preset stride;

[0031] For each of the selected grid sub-matrices, the grid sub-matrix is convolved with the weight matrix to obtain a first operation result, and the first operation result is added to the bias to obtain a second operation result;

[0032] Based on the second operation results corresponding to each of the grid sub-matrices, the raster matrix after the first convolution operation is determined.

[0033] In one implementation, at least one erosion processing operation is performed on the elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene to generate a sparse matrix corresponding to the object to be recognized, including:

[0034] Based on a third preset convolution kernel, at least one convolution operation is performed on the raster matrix to obtain a raster matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene;

[0035] The raster matrix with the preset sparsity after at least one convolution operation is determined as the sparse matrix corresponding to the object to be recognized.

[0036] In one implementation, rasterizing the acquired point cloud data to obtain a raster matrix includes:

[0037] Rasterizing the acquired point cloud data to obtain a raster matrix and the correspondence between each element in the raster matrix and the coordinate range information of each point cloud point;

[0038] The determining the position of the object to be recognized in the target scene based on the generated sparse matrix includes:

[0039] Based on the correspondence between each element in the raster matrix and the coordinate range information of each point cloud point, determine the coordinate information corresponding to each target element in the generated sparse matrix;

[0040] Combine the coordinate information corresponding to each of the target elements in the sparse matrix to determine the position of the object to be recognized in the target scene.

[0041] Here, based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point, the coordinate information of the target element in the generated sparse matrix can be determined. Then, based on the combination of the coordinate information, the coordinate range of the object to be recognized in the sparse matrix can be determined. Subsequently, based on the conversion relationship between the coordinate system where the sparse matrix is located and the physical coordinate system, the position of the object to be recognized in the target scene can be determined.

[0042] In one implementation manner, determining the position of the object to be recognized in the target scene based on the generated sparse matrix includes:

[0043] Performing at least one convolution process on each target element in the generated sparse matrix based on a trained convolutional neural network to obtain a convolution result;

[0044] Based on the convolution result, determining the position of the object to be recognized in the target scene.

[0045] Here, a trained convolutional neural network can be used to perform a convolution process on the generated sparse matrix to determine the position of the object to be recognized in the target scene based on the obtained convolution result. Considering that during the convolution process, convolution operations can be performed only on the target elements where there are point cloud points at the corresponding grids in the sparse matrix, which reduces the convolution calculation amount to a certain extent and improves the efficiency of target detection.

[0046] In a second aspect, an embodiment of the present disclosure further provides a processing device for point cloud data, where the device includes:

[0047] An acquisition module, configured to acquire point cloud data corresponding to a target scene;

[0048] A processing module, configured to perform rasterization processing on the acquired point cloud data to obtain a grid matrix; the value of each element in the grid matrix is used to represent whether there is a point cloud point at the corresponding grid;

[0049] A generation module, configured to generate a sparse matrix corresponding to the object to be recognized according to the grid matrix and the size information of the object to be recognized in the target scene;

[0050] A determination module, configured to determine the position of the object to be recognized in the target scene based on the generated sparse matrix.

[0051] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the method for processing point cloud data according to any one of the first aspect and its various implementation manners are executed.

[0052] In a fourth aspect, embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the method for processing point cloud data as described in any one of the first aspect and its various embodiments.

[0053] For the effect descriptions of the above-mentioned point cloud data processing device, electronic device, and computer-readable storage medium, refer to the description of the above-mentioned method for processing point cloud data, which will not be elaborated here.

[0054] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.

[0056] Figure 1 shows a flowchart of a method for processing point cloud data provided in the first embodiment of the present disclosure;

[0057] FIG. 2(a) shows a schematic diagram of a raster matrix before encoding provided in the first embodiment of the present disclosure;

[0058] FIG. 2(b) shows a schematic diagram of a sparse matrix provided in the first embodiment of the present disclosure;

[0059] FIG. 2(c) shows a schematic diagram of a raster matrix after encoding provided in the first embodiment of the present disclosure;

[0060] FIG. 3(a) shows a schematic diagram of a raster matrix after left shift provided in the first embodiment of the present disclosure;

[0061] FIG. 3(b) shows a schematic diagram of a logical OR operation provided in the first embodiment of the present disclosure;

[0062] FIG. 4(a) shows a schematic diagram of a raster matrix after a first inversion operation provided in the first embodiment of the present disclosure;

[0063] FIG. 4(b) shows a schematic diagram of a raster matrix after a convolution operation provided in the first embodiment of the present disclosure;

[0064] Figure 5 Shows a schematic diagram of a point cloud data processing device provided in the second embodiment of the present disclosure;

[0065] Figure 6 Shows a schematic diagram of an electronic device provided in the third embodiment of the present disclosure. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0067] It has been found through research that the point cloud processing solutions in the related art need to be customized using different programming languages for different application environments, which will consume a large amount of manpower and material resources.

[0068] Based on the above research, the present disclosure provides at least a point cloud data processing solution. Through rasterization processing and sparse processing under size information constraints, automatic generation of a sparse matrix is achieved, and scene applications are realized based on the generated sparse matrix, which saves time and effort.

[0069] All the defects existing in the above solutions are the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the present disclosure for the above problems in the following text should be the contributions made by the inventors to the present disclosure during the process of the present disclosure.

[0070] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0071] For ease of understanding this embodiment, first, a method for processing point cloud data disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the method for processing point cloud data provided in the embodiments of the present disclosure is generally an electronic device with certain computing capabilities. Such an electronic device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the method for processing point cloud data may be implemented by a processor invoking computer-readable instructions stored in a memory.

[0072] The method for processing point cloud data provided in the embodiments of the present disclosure will be described below.

[0073] Embodiment 1

[0074] See Figure 1 As shown in the flowchart of the method for processing point cloud data provided in the embodiments of the present disclosure, the method includes steps S101 to S104, where:

[0075] S101. Obtain point cloud data corresponding to a target scene;

[0076] S102. Perform rasterization processing on the obtained point cloud data to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster.

[0077] S103. Generate a sparse matrix corresponding to the object to be recognized in the target scene according to the raster matrix and the size information of the object to be recognized in the target scene;

[0078] S104. Determine the position of the object to be recognized in the target scene based on the generated sparse matrix.

[0079] Here, for ease of understanding the method for processing point cloud data provided in the embodiments of the present disclosure, the specific application scenarios of this processing method will be described in detail next. The method for processing point cloud data provided in the embodiments of the present disclosure can be mainly applied to fields such as target detection and three-dimensional target reconstruction. Here, target detection is taken as an example for illustration. In the related art, in order to determine information such as the position related to the target, after obtaining data information related to the application scenario (such as point cloud data), target detection can be achieved based on a pre-trained convolutional neural network. Here, considering that in the process of target detection relying on a convolutional neural network, it is necessary to perform convolutional operations on each point cloud point of the point cloud data, which to a certain extent leads to a large amount of convolutional calculations.

[0080] To solve the above problems, the embodiments of the present disclosure provide a solution for generating a sparse matrix based on rasterization processing and sparse processing under size limitation for object detection. On the one hand, since the above sparse matrix is generated by combining the size information of the object to be recognized in the target scene, to a certain extent, the generated sparse matrix can directly represent the relevant information of the object to be recognized. In the case where the accuracy requirement for object detection is not high, it can be directly used as the object detection result. On the other hand, in the process of object detection using the convolutional neural network adopted in the above related technology, since only the elements with point cloud points at the corresponding grids in the generated sparse matrix need to be subjected to convolution operations, to a certain extent, the convolution calculation amount can be reduced and the efficiency of object detection can be improved.

[0081] In the embodiments of the present disclosure, for the acquired point cloud data, first, rasterization processing can be performed, and then the raster matrix obtained by rasterization processing can be subjected to sparse processing to generate a sparse matrix. The process of rasterization processing here can refer to the process of mapping the point cloud data containing each point cloud point distributed in space into a set grid and performing grid encoding (corresponding to a binary matrix) based on the point cloud points corresponding to the grid. The process of sparse processing can refer to the process of performing a dilation processing operation (corresponding to the processing result of increasing the elements indicated as 1 in the binary matrix) or an erosion processing operation (corresponding to the processing result of reducing the elements indicated as 1 in the binary matrix) on the above binary matrix based on the size information of the object to be recognized in the target scene. Next, the above processes of rasterization processing and sparse processing will be further described.

[0082] Among them, in the process of the above rasterization processing, the point cloud points distributed in the Cartesian continuous real coordinate system can be converted to a rasterized discrete coordinate system.

[0083] To facilitate the understanding of the above rasterization processing process, a specific example will be described below. The embodiments of the present disclosure have point cloud points such as point A(0.32m, 0.48m), point B(0.6m, 0.4801m), and point C(2.1m, 3.2m). Rasterization is performed with a grid width of 1m. The range from (0m, 0m) to (1m, 1m) corresponds to the first grid, the range from (0m, 1m) to (1m, 2m) corresponds to the second grid, and so on. After rasterization, A'(0, 0) and B'(0, 0) are both in the grid of the first row and the first column, and C'(2, 3) can be in the grid of the second row and the third column, thus realizing the conversion from the Cartesian continuous real coordinate system to the discrete coordinate system. Among them, the coordinate information of the point cloud points can be determined with reference to a reference point (such as the position of the radar device that collects the point cloud data), which will not be elaborated here.

[0084] In the embodiments of the present disclosure, two-dimensional rasterization can be performed, or three-dimensional rasterization can be performed. Three-dimensional rasterization adds height information compared to two-dimensional rasterization. Next, a specific description will be given taking two-dimensional rasterization as an example.

[0085] For two-dimensional rasterization, a finite space can be divided into N*M grids, generally divided at equal intervals, and the interval size can be configured. At this time, the point cloud data after rasterization can be encoded using a zero-one matrix (i.e., the above-mentioned grid matrix). Each grid can be represented by a coordinate composed of a unique row number and column number. If there is one or more point cloud points in the grid, the grid is encoded as 1, otherwise as 0, so that an encoded zero-one matrix can be obtained.

[0086] After determining the grid matrix according to the above method, the elements in the above grid matrix can be sparsely processed according to the size information of the object to be recognized in the target scene to generate a corresponding sparse matrix.

[0087] Among them, the size information of the object to be recognized can be obtained in advance. Here, the size information of the object to be recognized can be determined by combining the image data synchronously collected with the point cloud data, or can be roughly estimated based on the specific application scenario of the point cloud data processing method provided by the embodiments of the present disclosure. For example, in the field of autonomous driving, the object in front of the vehicle can be the vehicle, and its general size information can be determined as 4m×4m. In addition, the embodiments of the present disclosure can also determine the size information of the object to be recognized based on other methods, and the embodiments of the present disclosure do not make specific limitations on this.

[0088] In the embodiments of the present disclosure, the sparse processing operation can be to perform at least one dilation processing operation on the target elements in the grid matrix (i.e., the elements representing the existence of point cloud points at the corresponding grid). Here, the dilation processing operation can be performed when the size of the coordinate range of the grid matrix is smaller than the size of the object to be recognized in the target scene. That is, through one or more dilation processing operations, the range of elements representing the existence of point cloud points at the corresponding grid can be gradually expanded so that the expanded element range can match the object to be recognized, thereby realizing position determination. In addition, the sparse processing operation in the embodiments of the present disclosure can also be to perform at least one erosion processing operation on the target elements in the grid matrix. Here, the erosion processing operation can be performed when the size of the coordinate range of the grid matrix is larger than the size of the object to be recognized in the target scene. That is, through one or more erosion processing operations, the range of elements representing the existence of point cloud points at the corresponding grid can be gradually reduced so that the reduced element range can match the object to be recognized, thereby realizing position determination.

[0089] In a specific application, whether to perform a single dilation operation, multiple dilation operations, a single erosion operation, or multiple erosion operations depends on whether the difference between the coordinate range of the sparse matrix obtained by performing at least one shift operation and logical operation processing and the size of the object to be recognized in the target scene falls within a preset threshold range. That is to say, the dilation or erosion operation adopted in the present disclosure is carried out based on the constraint of the size information of the object to be recognized, so that the information represented by the determined sparse matrix is more in line with the relevant information of the object to be recognized.

[0090] It can be understood that the purpose of the sparse processing implemented based on either the dilation operation or the erosion operation is to enable the generated sparse matrix to represent more accurately the relevant information of the object to be recognized.

[0091] In the embodiments of the present disclosure, the above-mentioned dilation operation can be implemented based on a shift operation and a logical OR operation, or can also be implemented based on taking the inverse after convolution and then taking the inverse after convolution. Although the specific methods adopted by the two operations are different, the effects of the finally generated sparse matrices can be the same.

[0092] In addition, the above-mentioned erosion operation can be implemented based on a shift operation and a logical AND operation, or can also be directly implemented based on a convolution operation. Similarly, although the specific methods adopted by the two operations are different, the effects of the finally generated sparse matrices can also be the same.

[0093] Next, taking the dilation operation as an example, combined with Figures 2(a) to 2(b) the specific example diagram of generating a sparse matrix shown, the generation process of the above-mentioned sparse matrix will be further described.

[0094] As shown in Fig. 2(a), it is a schematic diagram of the grid matrix obtained after rasterization processing (corresponding to before encoding). By performing a single eight-neighborhood dilation operation on each target element (corresponding to the grid with a filling effect) in this grid matrix, the corresponding sparse matrix as shown in Fig. 2(b) can be obtained. It can be known that in the embodiments of the present disclosure, for the target elements with point cloud points at the corresponding grids in Fig. 2(a), an eight-neighborhood dilation operation is performed, so that each target element becomes an element set after dilation, and the grid width corresponding to this element set can match the size of the object to be recognized.

[0095] Among them, the dilation operation of the above eight-neighborhood can be a process of determining elements whose absolute values of the differences in the abscissa or ordinate with this element do not exceed 1. Except for the elements on the grid edge, generally there are eight elements in the neighborhood of an element (corresponding to the above element set). The input of the dilation processing result can be the coordinate information of 6 target elements, and the output can be the coordinate information of the element set within the eight-neighborhood of the target element, as shown in Fig. 2(b).

[0096] It should be noted that in practical applications, in addition to performing the above dilation operation of the eight-neighborhood, a dilation operation of the four-neighborhood can also be performed, and there are no specific restrictions on other dilation operations for the latter. In addition, the embodiments of the present disclosure can also perform multiple dilation operations. For example, based on the dilation result shown in Fig. 2(b), a dilation operation is performed again to obtain a sparse matrix with a larger element set range, which will not be elaborated here.

[0097] In the embodiments of the present disclosure, based on the generated sparse matrix, the position of the object to be recognized in the target scene can be determined. The embodiments of the present disclosure can be specifically implemented through the following two aspects.

[0098] First aspect: Here, based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point, the position range of the object to be recognized can be determined, and it can be specifically implemented through the following steps:

[0099] Step 1: Based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point, determine the coordinate information corresponding to each target element in the generated sparse matrix;

[0100] Step 2: Combine the coordinate information corresponding to each target element in the sparse matrix to determine the position of the object to be recognized in the target scene.

[0101] Here, based on the above relevant descriptions of the rasterization process, each target element in the grid matrix can correspond to multiple point cloud points. In this way, the point cloud point coordinate range information corresponding to the element and multiple point cloud points can be determined in advance. Here, still taking the grid matrix of N*M dimension as an example, the target element with point cloud points can correspond to P point cloud points, and the coordinate of each point is (Xi, Yi), where i belongs to 0 to P - 1, and Xi, Yi represent the position of the point cloud point in the grid matrix, 0 <= Xi < N, 0 <= Yi < M.

[0102] In this way, after generating the sparse matrix, the coordinate information corresponding to each target element in the sparse matrix can be determined by using the above correspondence between each element and the coordinate range information of each point cloud point determined in advance, that is, an inverse rasterization processing operation is performed.

[0103] It should be noted that since the sparse matrix is obtained by sparsely processing the elements in the grid matrix that represent the presence of point cloud points at the corresponding grids, therefore, the target elements in the sparse matrix here can also represent the elements where there are point cloud points at the corresponding grids.

[0104] To facilitate the understanding of the above inverse rasterization process, a specific example will be given below for illustration. Here, taking the points A'(0,0) and B'(0,0) indicated by the sparse matrix in the first row and first column grid, and the point C'(2,3) in the second row and third column grid as an example. During the inverse rasterization process, for the first grid (0,0), after mapping its center back to the Cartesian coordinate system, (0.5m, 0.5m) can be obtained. For the grid (2,3) in the second row and third column, after mapping its center back to the Cartesian coordinate system, (2.5m, 3.5m) can be obtained. That is, (0.5m, 0.5m) and (2.5m, 3.5m) can be determined as the mapped coordinate information. In this way, by combining the mapped coordinate information, the position of the object to be recognized in the target scene can be determined.

[0105] The embodiments of the present disclosure can not only determine the position range of the object to be recognized based on the approximate relationship between the above sparse matrix and the target detection result, but also determine the position range of the object to be recognized based on the trained convolutional neural network.

[0106] Second aspect: The embodiments of the present disclosure can first perform at least one convolution process on the generated sparse matrix based on the trained convolutional neural network, and then determine the position range of the object to be recognized based on the convolution result obtained from the convolution process.

[0107] In the related technologies that use convolutional neural networks to implement target detection, it is necessary to traverse all the input data, sequentially find the neighborhood points of the input points for convolution operations, and finally output the set of all neighborhood points. However, the method for processing point cloud data provided by the embodiments of the present disclosure only needs to quickly traverse the target elements in the sparse matrix to find the positions of the valid points (i.e., the elements with a value of 1 in the binary matrix) for convolution operations, thereby greatly accelerating the calculation process of the convolutional neural network and improving the efficiency of determining the position range of the object to be recognized.

[0108] Considering the key role of the sparse processing operation in the method for processing point cloud data provided by the embodiments of the present disclosure, the following two aspects will be described separately.

[0109] First aspect: In the case where the sparse processing operation is a dilation processing operation, the embodiments of the present disclosure can be implemented by combining shift processing and logical operations, and can also be implemented based on convolution after inversion and then inversion after convolution.

[0110] First, in the embodiments of the present disclosure, one or more dilation processing operations can be performed based on at least one shift processing and a logical OR operation. In the specific implementation process, the number of specific dilation processing operations can be determined in combination with the size information of the object to be recognized in the target scenario.

[0111] Here, for the first dilation processing operation, the target element representing the presence of a point cloud point at the corresponding grid can be shifted in multiple preset directions to obtain a corresponding plurality of shifted grid matrices. Then, a logical OR operation can be performed on the grid matrix and the plurality of shifted grid matrices corresponding to the first dilation processing operation, so as to obtain a sparse matrix after the first dilation processing operation. Here, it can be determined whether the size of the coordinate range of the obtained sparse matrix is smaller than the size of the object to be recognized, and whether the corresponding difference is large enough (such as greater than a preset threshold). If so, the target element in the sparse matrix after the first dilation processing operation can be shifted in multiple preset directions and a logical OR operation can be performed according to the above method to obtain a sparse matrix after the second dilation processing operation, and so on, until it is determined that the difference between the size of the coordinate range of the newly obtained sparse matrix and the size of the object to be recognized in the target scenario belongs to the preset threshold range, and the sparse matrix is determined.

[0112] It should be noted that, no matter which sparse matrix is obtained after the dilation processing operation, it is essentially a zero-one matrix. As the number of dilation processing operations increases, the number of target elements representing the presence of point cloud points at the corresponding grids in the obtained sparse matrix also increases. And since the grids mapped by the zero-one matrix have width information, here, the size of the coordinate range corresponding to each target element in the sparse matrix can be used to verify whether the size of the object to be recognized in the target scenario is reached, thereby improving the accuracy of subsequent target detection applications.

[0113] Among them, the above logical OR operation can be implemented according to the following steps:

[0114] Step 1: Select a shifted grid matrix from the plurality of shifted grid matrices;

[0115] Step 2: Perform a logical OR operation on the grid matrix before the current dilation processing operation and the selected shifted grid matrix to obtain an operation result;

[0116] Step 3: Loop to select the unprocessed grid matrices from the plurality of shifted grid matrices, and perform a logical OR operation on the selected grid matrix and the result of the most recent operation until all the grid matrices are selected, to obtain the sparse matrix after the current dilation processing operation.

[0117] Here, first, one of the shifted grid matrices can be selected from multiple shifted grid matrices. In this way, the grid matrix before the current dilation processing operation can be logically ORed with the selected shifted grid matrix to obtain an operation result. Here, the unprocessed shifted grid matrices can be cyclically selected from multiple shifted grid matrices and participate in the logical OR operation until all the shifted grid matrices are selected, and then the sparse matrix after the current dilation processing operation can be obtained.

[0118] The dilation processing operation in the embodiments of the present disclosure can be four-neighborhood dilation centered on the target element, can also be eight-neighborhood dilation centered on the target element, or other neighborhood processing operation methods. In specific applications, the corresponding neighborhood processing operation method can be selected based on the size information of the object to be recognized, and no specific limitation is made here.

[0119] It should be noted that for different neighborhood processing operation methods, the preset directions of the corresponding shift processing are different. Taking four-neighborhood dilation as an example, the grid matrix can be shifted in four preset directions, namely left shift, right shift, up shift, and down shift. Taking eight-neighborhood dilation as an example, the grid matrix can be shifted in four preset directions, namely left shift, right shift, up shift, down shift, up shift and down shift on the premise of left shift, and up shift and down shift on the premise of right shift. In addition, in order to adapt to the subsequent logical OR operation, after determining the shifted grid matrix based on multiple shift directions, a logical OR operation can be performed first, and then the result of the logical OR operation is shifted in multiple shift directions, and then the next logical OR operation is performed, and so on, until the sparse matrix after the dilation processing is obtained.

[0120] To facilitate the understanding of the above dilation processing operation, the grid matrix before encoding shown in Fig. 2(a) can be first converted into the grid matrix after encoding shown in Fig. 2(c), and then an example of the first dilation processing operation is described in combination with Figures 3(a) to 3(b) to illustrate the first dilation processing operation.

[0121] As shown in the grid matrix of Fig. 2(c), as a binary matrix, the positions of all 1s in the matrix can represent the grids where the target elements are located, and all 0s in the matrix can represent the background.

[0122] In the embodiments of the present disclosure, first, matrix shifting can be used to determine the neighborhoods of the elements with a value of 1 in the binary matrix. Here, four preset shifting operations in different directions can be defined, namely left shift, right shift, up shift, and down shift. Among them, for the left shift, the column coordinates of the elements with a value of 1 in the binary matrix are decreased by 1, as shown in Fig. 3(a); for the right shift, the column coordinates of the elements with a value of 1 in the binary matrix are increased by 1; for the up shift, the row coordinates of the elements with a value of 1 in the binary matrix are decreased by 1; for the down shift, the row coordinates of the elements with a value of 1 in the binary matrix are increased by 1.

[0123] Secondly, in the embodiments of the present disclosure, matrix logical OR operation can be used to combine the results of all neighborhoods. Matrix logical OR means that when two binary matrices with the same size are received as inputs, the logical OR operation is performed on the binary values at the same positions of the two matrices in sequence, and the resulting values form a new binary matrix as the output. Fig. 3(b) shows a specific example of a logical OR operation.

[0124] In the specific process of implementing the logical OR operation, the raster matrices after left shift, right shift, up shift, and down shift can be sequentially selected to participate in the logical OR operation. For example, the raster matrix can be logically ORed with the raster matrix after left shift first, and the resulting operation result can be logically ORed with the raster matrix after right shift. Then, the resulting operation result can be logically ORed with the raster matrix after up shift, and the resulting operation result can be logically ORed with the raster matrix after down shift, so as to obtain the sparse matrix after the first dilation processing operation.

[0125] It should be noted that the above selection order of the raster matrices after translation is only a specific example. In practical applications, other methods can also be combined for selection. Considering the symmetry of the translation operation, here, the up shift and down shift can be paired for logical OR operation, and the left shift and right shift can be paired for logical operation. The two logical OR operations can be performed synchronously, which can save calculation time.

[0126] Second, in the embodiments of the present disclosure, the dilation processing operation can be implemented by combining convolution and two negation processes. Specifically, it can be realized through the following steps:

[0127] Step 1: Perform a first negation operation on the elements in the raster matrix before the current dilation processing operation to obtain the raster matrix after the first negation operation;

[0128] Step 2: Perform at least one convolution operation on the raster matrix after the first negation operation based on a first preset convolution kernel to obtain a raster matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene;

[0129] Step 3: Perform a second negation operation on the elements in the raster matrix with a preset sparsity after at least one convolution operation to obtain a sparse matrix.

[0130] In the embodiments of the present disclosure, the dilation processing operation can be implemented through the operations of negating after convolution and then negating after convolution. To a certain extent, the obtained sparse matrix can also represent the relevant information of the object to be recognized. In addition, considering that the above convolution operation can be automatically combined with the convolutional neural network used in subsequent applications such as target detection, the detection efficiency can be improved to a certain extent.

[0131] In the embodiments of the present disclosure, the negation operation can be implemented based on the convolution operation or other negation operation methods. For the convenience of cooperating with the subsequent application network (such as the convolutional neural network used for target detection), here, the convolution operation can be specifically used to implement it. Next, the above first negation operation will be specifically described.

[0132] Here, a first negated element can be obtained by performing a convolution operation on other elements except the target element in the raster matrix before the current dilation processing operation based on a second preset convolution kernel, and a second negated element can also be obtained by performing a convolution operation on the target element in the raster matrix before the current dilation processing operation based on the second preset convolution kernel. Based on the above first negated element and second negated element, the raster matrix after the first negation operation can be determined.

[0133] The implementation process of the second negation operation can refer to the implementation process of the above first negation operation, and will not be elaborated here.

[0134] In the embodiments of the present disclosure, at least one convolution operation can be performed on the raster matrix after the first negation operation by using a first preset convolution kernel to obtain a raster matrix with a preset sparsity. If the dilation processing operation can be regarded as a means to increase the number of target elements in the raster matrix, then the above convolution operation can be regarded as a process of reducing the number of target elements in the raster matrix (corresponding to the erosion processing operation). Since the convolution operation in the embodiments of the present disclosure is performed on the raster matrix after the first negation operation, the equivalent operation equivalent to the above dilation processing operation is realized by using the negation operation in combination with the erosion processing operation and then performing the negation operation again.

[0135] Among them, for the first convolution operation, the raster matrix after the first negation operation is convolved with the first preset convolution kernel to obtain the raster matrix after the first convolution operation. After determining that the sparsity of the raster matrix after the first convolution operation does not reach the preset sparsity, the raster matrix after the first convolution operation can be convolved with the first preset convolution kernel again to obtain the raster matrix after the second convolution operation, and so on, until a raster matrix with a preset sparsity can be determined.

[0136] Among them, the above sparsity can be determined by the proportion distribution of target elements and non-target elements in the grid matrix. The more the proportion of target elements, the larger the size information of the object to be recognized represented by it. On the contrary, the less the proportion of target elements, the smaller the size information of the object to be recognized represented by it. In the embodiments of the present disclosure, the convolution operation can be stopped when the proportion distribution reaches the preset sparsity.

[0137] The convolution operation in the embodiments of the present disclosure can be performed once or multiple times. Here, the specific operation process of the first convolution operation will be described, including the following steps:

[0138] Step 1: For the first convolution operation, select each grid sub-matrix from the grid matrix after the first inversion operation according to the size of the first preset convolution kernel and the preset stride.

[0139] Step 2: For each selected grid sub-matrix, perform a multiplication operation on the grid sub-matrix and the weight matrix to obtain a first operation result, and perform an addition operation on the first operation result and the bias to obtain a second operation result.

[0140] Step 3: Based on the second operation results corresponding to each grid sub-matrix, determine the grid matrix after the first convolution operation.

[0141] Here, a traversal method can be used to traverse the grid matrix after the first inversion operation. In this way, for each grid sub-matrix traversed, the grid sub-matrix can be multiplied by the weight matrix to obtain a first operation result, and the first operation result can be added to the bias to obtain a second operation result. In this way, by combining the second operation results corresponding to each grid sub-matrix into the corresponding matrix elements, the grid matrix after the first convolution operation can be obtained.

[0142] To facilitate the understanding of the above dilation processing operation, here, still taking the encoded grid matrix shown in Figure 2(c) as an example, in combination with Figures 4(a) to 4(b) An example of the dilation processing operation will be described.

[0143] Here, a 1*1 convolution kernel (i.e., the second preset convolution kernel) can be used to implement the first inversion operation. The weight of the second preset convolution kernel is -1 and the bias is 1. At this time, substitute the weight and bias into the convolution formula {output = input grid matrix * weight + bias}. If the input is a target element in the grid matrix and its value corresponds to 1, then the output = 1 * -1 + 1 = 0; if the input is a non-target element in the grid matrix and its value corresponds to 0, then the output = 0 * -1 + 1 = 1; in this way, after the 1*1 convolution kernel acts on the input, the binary matrix can be inverted, and the element value 0 becomes 1 and the element value 1 becomes 0, as shown in Figure 4(a).

[0144] For the above corrosion treatment operation, in specific applications, a 3×3 convolutional kernel (i.e., the first preset convolutional kernel) and a rectified linear unit (ReLU) can be used to implement it. Each weight included in the weight matrix of the above first preset convolutional kernel is 1, and the bias is 8. In this way, the above corrosion treatment operation can be implemented using the formula {output = ReLU (the raster matrix after the first inversion operation of the input * weight + bias)}.

[0145] Here, only when all elements in the input 3×3 raster sub-matrix are 1, the output = ReLU(9 - 8) = 1; otherwise, the output = ReLU (the input raster sub-matrix * 1 - 8) = 0, where (the input raster sub-matrix * 1 - 8) < 0. As shown in Figure 4(b), it is the raster matrix after the convolution operation.

[0146] Here, each time a convolutional network with a second preset convolutional kernel is nested, a corrosion operation can be superimposed once, so that a raster matrix with a fixed sparsity can be obtained. Taking the inversion operation again can be equivalent to a dilation processing operation, thus realizing the generation of a sparse matrix.

[0147] Second aspect: When the sparse processing operation is a corrosion treatment operation, the embodiments of the present disclosure can be implemented by combining shift processing and logical operations, and can also be implemented based on convolution operations.

[0148] First, in the embodiments of the present disclosure, one or more corrosion treatment operations can be performed based on at least one shift processing and logical AND operation. In the specific implementation process, the specific number of corrosion treatment operations can be determined in combination with the size information of the object to be recognized in the target scenario.

[0149] Similar to the dilation processing implemented based on shift processing and logical OR operation in the first aspect, during the corrosion treatment operation, the shift processing of the raster matrix can also be performed first. Different from the above dilation processing, the logical operation here can be a logical AND operation on the shifted raster matrix. For the process of implementing the corrosion treatment operation based on shift processing and logical AND operation, please refer to the above description for details and will not be elaborated here.

[0150] Similarly, the corrosion treatment operation in the embodiments of the present disclosure can be a four-neighborhood corrosion centered on the target element, an eight-neighborhood corrosion centered on the target element, or other neighborhood processing operation methods. In specific applications, the corresponding neighborhood processing operation method can be selected based on the size information of the object to be recognized, and no specific limitation is made here.

[0151] Second, in the embodiments of the present disclosure, erosion processing operations can be implemented in combination with convolution processing, and can be specifically implemented through the following steps:

[0152] Step 1: Perform at least one convolution operation on the grid matrix based on a third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scenario;

[0153] Step 2: Determine the grid matrix with a preset sparsity after at least one convolution operation as a sparse matrix corresponding to the object to be recognized.

[0154] The above convolution operation can be regarded as a process of reducing the number of target elements in the grid matrix, that is, the erosion processing process. Among them, for the first convolution operation, the grid matrix is convolved with the first preset convolution kernel to obtain the grid matrix after the first convolution operation. After determining that the sparsity of the grid matrix after the first convolution operation does not reach the preset sparsity, the grid matrix after the first convolution operation can be convolved with the third preset convolution kernel again to obtain the grid matrix after the second convolution operation, and so on, until a grid matrix with a preset sparsity can be determined, that is, a sparse matrix corresponding to the object to be recognized is obtained.

[0155] The convolution operation in the embodiments of the present disclosure can be one or multiple times. For the specific process of the convolution operation, refer to the relevant description of the dilation processing implemented based on convolution and inversion in the above first aspect, which will not be elaborated here.

[0156] It should be noted that in specific applications, convolutional neural networks with different data processing bit widths can be used to generate sparse matrices. For example, 4 bits (bit) can be used to represent the input, output, and parameters used in the calculation of the network, such as the element values (0 or 1) of the grid matrix, weights, bias amounts, etc. In addition, 8 bits can also be used for representation to adapt to the network processing bit width and improve the operation efficiency.

[0157] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0158] Based on the same inventive concept, a processing device for point cloud data corresponding to the processing method of point cloud data is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above processing method of point cloud data in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0159] Embodiment 2

[0160] Reference Figure 5 As shown, it is a schematic structural diagram of a point cloud data processing device provided by an embodiment of the present disclosure. The device includes: an acquisition module 501, a processing module 502, a generation module 503, and a determination module 504; wherein,

[0161] The acquisition module 501 is configured to acquire point cloud data corresponding to a target scene;

[0162] The processing module 502 is configured to perform rasterization processing on the acquired point cloud data to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster;

[0163] The generation module 503 is configured to generate a sparse matrix corresponding to the object to be recognized according to the raster matrix and the size information of the object to be recognized in the target scene;

[0164] The determination module 504 is configured to determine the position of the object to be recognized in the target scene based on the generated sparse matrix.

[0165] By using the above point cloud data processing device, each point cloud point in the point cloud data can be first mapped to the corresponding raster. Some rasters correspond to one or more point cloud points, and some rasters do not correspond to point cloud points. In this way, the raster matrix determined based on the above mapping relationship can be a standardized zero-one matrix, and the zero-one matrix can be used in relevant processing operations to determine the corresponding sparse matrix. Since the above processing operations are performed in combination with the size information of the object to be recognized in the target scene, the elements with a value of 1 in the sparse matrix generated after the processing operations can, to a certain extent, represent the relevant information of the object to be recognized. Here, the position of the object to be recognized in the target scene can be determined.

[0166] In one embodiment, the generation module 503 is configured to generate a sparse matrix corresponding to the object to be recognized according to the raster matrix and the size information of the object to be recognized in the target scene according to the following steps:

[0167] Perform at least one dilation processing operation or erosion processing operation on the target elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene to generate a sparse matrix corresponding to the object to be recognized;

[0168] Wherein, the target element is an element representing that there is a point cloud point at the corresponding raster.

[0169] In one embodiment, the generation module 503 is configured to perform at least one dilation processing operation or erosion processing operation on the target elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene according to the following steps to generate a sparse matrix corresponding to the object to be recognized:

[0170] Perform at least one shift processing and logical operation processing on the target elements in the grid matrix to obtain a sparse matrix corresponding to the object to be recognized, where the difference between the size of the coordinate range of the obtained sparse matrix and the size of the object to be recognized in the target scene belongs to a preset threshold range.

[0171] In one implementation, the generation module 503 is configured to perform at least one dilation processing operation on the elements in the grid matrix according to the following steps based on the grid matrix and the size information of the object to be recognized in the target scene, and generate a sparse matrix corresponding to the object to be recognized:

[0172] Perform a first negation operation on the elements in the grid matrix before the current dilation processing operation to obtain the grid matrix after the first negation operation;

[0173] Perform at least one convolution operation on the grid matrix after the first negation operation based on the first preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene;

[0174] Perform a second negation operation on the elements in the grid matrix with a preset sparsity after at least one convolution operation to obtain a sparse matrix.

[0175] In one implementation, the generation module 503 is configured to perform a first negation operation on the elements in the grid matrix before the current dilation processing operation according to the following steps to obtain the grid matrix after the first negation operation:

[0176] Perform a convolution operation on other elements except the target elements in the grid matrix before the current dilation processing operation based on the second preset convolution kernel to obtain the first negated element, and perform a convolution operation on the target elements in the grid matrix before the current dilation processing operation based on the second preset convolution kernel to obtain the second negated element;

[0177] Obtain the grid matrix after the first negation operation based on the first negated element and the second negated element.

[0178] In one implementation, the generation module 503 is configured to perform at least one convolution operation on the grid matrix after the first negation operation based on the first preset convolution kernel according to the following steps to obtain a grid matrix with a preset sparsity after at least one convolution operation:

[0179] For the first convolution operation, perform a convolution operation on the grid matrix after the first negation operation and the first preset convolution kernel to obtain the grid matrix after the first convolution operation;

[0180] Determine whether the sparsity of the grid matrix after the first convolution operation reaches a preset sparsity;

[0181] If not, loop through the step of convolving the grid matrix after the previous convolution operation with the first preset convolution kernel to obtain the grid matrix after the current convolution operation until a grid matrix with the preset sparsity after at least one convolution operation is obtained.

[0182] In one implementation, the first preset convolution kernel has a weight matrix and a bias corresponding to the weight matrix; the generation module 503 is configured to perform a convolution operation on the grid matrix after the first negation operation with the first preset convolution kernel according to the following steps to obtain the grid matrix after the first convolution operation:

[0183] For the first convolution operation, select each grid sub-matrix from the grid matrix after the first negation operation according to the size of the first preset convolution kernel and the preset stride;

[0184] For each selected grid sub-matrix, perform a convolution operation on the grid sub-matrix with the weight matrix to obtain a first operation result, and perform an addition operation on the first operation result and the bias to obtain a second operation result;

[0185] Based on the second operation results corresponding to each grid sub-matrix, determine the grid matrix after the first convolution operation.

[0186] In one implementation, the generation module 503 is configured to perform at least one erosion processing operation on the elements in the grid matrix according to the following steps based on the grid matrix and the size information of the object to be recognized in the target scene to generate a sparse matrix corresponding to the object to be recognized:

[0187] Perform at least one convolution operation on the grid matrix based on the third preset convolution kernel to obtain a grid matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene;

[0188] Determine the grid matrix with the preset sparsity after at least one convolution operation as the sparse matrix corresponding to the object to be recognized.

[0189] In one implementation, the processing module 502 is configured to perform a rasterization process on the acquired point cloud data according to the following steps to obtain a grid matrix:

[0190] Perform a rasterization process on the acquired point cloud data to obtain a grid matrix and the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point;

[0191] A determination module 504, configured to determine the position of an object to be recognized in a target scene based on the generated sparse matrix according to the following steps:

[0192] Based on the correspondence between each element in the grid matrix and the coordinate range information of each point cloud point, determine the coordinate information corresponding to each target element in the generated sparse matrix;

[0193] Combine the coordinate information corresponding to each target element in the sparse matrix to determine the position of the object to be recognized in the target scene.

[0194] In an implementation, the determination module 504 is configured to determine the position of the object to be recognized in the target scene based on the generated sparse matrix according to the following steps:

[0195] Perform at least one convolution process on each target element in the generated sparse matrix based on a trained convolutional neural network to obtain a convolution result;

[0196] Based on the convolution result, determine the position of the object to be recognized in the target scene.

[0197] For the description of the processing flow of each module in the device and the interaction flow between modules, reference may be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.

[0198] Embodiment III

[0199] This embodiment of the present disclosure further provides an electronic device, as Figure 6 shown, which is a schematic structural diagram of the electronic device provided in this embodiment of the present disclosure, including: a processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions executable by the processor 601 (such as Figure 5 the instructions corresponding to the acquisition module 501, the processing module 502, the generation module 503, and the determination module 504 in the point cloud data processing device shown). When the electronic device runs, the processor 601 communicates with the memory 602 through the bus 603. When the machine-readable instructions are executed by the processor 601, the following processing is performed:

[0200] Obtain point cloud data corresponding to the target scene;

[0201] Perform rasterization processing on the obtained point cloud data to obtain a grid matrix; the value of each element in the grid matrix is used to represent whether there is a point cloud point at the corresponding grid;

[0202] According to the grid matrix and the size information of the object to be recognized in the target scene, generate a sparse matrix corresponding to the object to be recognized;

[0203] Based on the generated sparse matrix, determine the position of the object to be recognized in the target scene.

[0204] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor 601, it executes the steps of the method for processing point cloud data in the foregoing method embodiment. Wherein, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0205] A computer program product for the method for processing point cloud data provided by an embodiment of the present disclosure includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the method for processing point cloud data described in the foregoing method embodiment. For details, please refer to the foregoing method embodiment and will not be elaborated herein.

[0206] An embodiment of the present disclosure also provides a computer program, which implements any one of the foregoing methods when executed by a processor. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0207] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.

[0208] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0209] In addition, in each embodiment of the present disclosure, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0210] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0211] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for processing point cloud data, characterized in that, the method includes: obtaining point cloud data corresponding to a target scene; performing rasterization processing on the obtained point cloud data to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster; performing at least one dilation processing operation or erosion processing operation on target elements in the raster matrix according to the raster matrix and size information of an object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized; wherein, the target elements are elements representing that there are point cloud points at the corresponding rasters; the dilation processing operation includes an operation implemented based on a shift operation and a logical OR operation, or an operation implemented based on taking the inverse after convolution and then taking the inverse after convolution; determining the position of the object to be recognized in the target scene based on the generated sparse matrix.

2. The processing method according to claim 1, characterized in that, performing at least one dilation processing operation or erosion processing operation on target elements in the raster matrix according to the raster matrix and size information of an object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized, includes: performing at least one shift processing and logical operation processing on target elements in the raster matrix, to obtain a sparse matrix corresponding to the object to be recognized, wherein the difference between the size of the coordinate range of the obtained sparse matrix and the size of the object to be recognized in the target scene belongs to a preset threshold range.

3. The processing method according to claim 1, characterized in that, performing at least one dilation processing operation on elements in the raster matrix according to the raster matrix and size information of an object to be recognized in the target scene, to generate a sparse matrix corresponding to the object to be recognized, includes: performing a first inversion operation on elements in the raster matrix before the current dilation processing operation, to obtain a raster matrix after the first inversion operation; performing at least one convolution operation on the raster matrix after the first inversion operation based on a first preset convolution kernel, to obtain a raster matrix with a preset sparsity after at least one convolution operation; the preset sparsity is determined by size information of the object to be recognized in the target scene; performing a second inversion operation on elements in the raster matrix with the preset sparsity after at least one convolution operation, to obtain the sparse matrix.

4. The processing method according to claim 3, characterized in that, the performing a first inversion operation on elements in the raster matrix before the current dilation processing operation, to obtain a raster matrix after the first inversion operation, includes: performing a convolution operation on other elements except the target elements in the raster matrix before the current dilation processing operation based on a second preset convolution kernel, to obtain first inverted elements, and performing a convolution operation on the target elements in the raster matrix before the current dilation processing operation based on a second preset convolution kernel, to obtain second inverted elements; obtaining a raster matrix after the first inversion operation based on the first inverted elements and the second inverted elements.

5. The processing method according to claim 3 or 4, wherein, the at least one convolution operation on the raster matrix after the first inversion operation based on the first preset convolution kernel to obtain a raster matrix with a preset sparsity after the at least one convolution operation includes: for the first convolution operation, performing a convolution operation on the raster matrix after the first inversion operation and the first preset convolution kernel to obtain a raster matrix after the first convolution operation; judging whether the sparsity of the raster matrix after the first convolution operation reaches the preset sparsity; if not, looping to perform the step of performing a convolution operation on the raster matrix after the previous convolution operation and the first preset convolution kernel to obtain a raster matrix after the current convolution operation until a raster matrix with a preset sparsity after the at least one convolution operation is obtained.

6. The processing method according to claim 5, wherein, the first preset convolution kernel has a weight matrix and a bias corresponding to the weight matrix; for the first convolution operation, performing a convolution operation on the raster matrix after the first inversion operation and the first preset convolution kernel to obtain a raster matrix after the first convolution operation includes: for the first convolution operation, selecting each raster sub-matrix from the raster matrix after the first inversion operation according to the size of the first preset convolution kernel and a preset stride; for each selected raster sub-matrix, performing a multiplication operation on the raster sub-matrix and the weight matrix to obtain a first operation result, and performing an addition operation on the first operation result and the bias to obtain a second operation result; determining a raster matrix after the first convolution operation based on the second operation results corresponding to the respective raster sub-matrices.

7. The processing method according to claim 1, wherein, performing at least one erosion processing operation on the elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene to generate a sparse matrix corresponding to the object to be recognized, including: performing at least one convolution operation on the raster matrix based on a third preset convolution kernel to obtain a raster matrix with a preset sparsity after the at least one convolution operation; the preset sparsity is determined by the size information of the object to be recognized in the target scene; determining the raster matrix with the preset sparsity after the at least one convolution operation as a sparse matrix corresponding to the object to be recognized.

8. The processing method according to any one of claims 1 to 7, wherein, performing rasterization processing on the acquired point cloud data to obtain a raster matrix, including: performing rasterization processing on the acquired point cloud data to obtain a raster matrix and the corresponding relationship between each element in the raster matrix and the coordinate range information of each point cloud point; the determining the position of the object to be recognized in the target scene based on the generated sparse matrix includes: determining the coordinate information corresponding to each target element in the generated sparse matrix based on the corresponding relationship between each element in the raster matrix and the coordinate range information of each point cloud point; Combine the coordinate information corresponding to each of the target elements in the sparse matrix to determine the position of the object to be recognized in the target scene.

9. According to the processing method according to any one of claims 1 to 7, characterized in that determining the position of the object to be recognized in the target scene based on the generated sparse matrix includes: Performing at least one convolution process on each target element in the generated sparse matrix based on a trained convolutional neural network to obtain a convolution result; Based on the convolution result, determine the position of the object to be recognized in the target scene.

10. A processing device for point cloud data, characterized in that the device includes: An acquisition module for acquiring point cloud data corresponding to a target scene; A processing module for rasterizing the acquired point cloud data to obtain a raster matrix; the value of each element in the raster matrix is used to represent whether there is a point cloud point at the corresponding raster; A generation module for performing at least one dilation process operation or erosion process operation on the target elements in the raster matrix according to the raster matrix and the size information of the object to be recognized in the target scene, and generating a sparse matrix corresponding to the object to be recognized; wherein, the target element is an element representing that there is a point cloud point at the corresponding raster; the dilation process operation includes an operation implemented based on a shift operation and a logical OR operation, or an operation implemented based on convolution after inversion and then inversion after convolution; A determination module for determining the position of the object to be recognized in the target scene based on the generated sparse matrix.

11. An electronic device, characterized in that it includes: A processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps of the processing method for point cloud data according to any one of claims 1 to 9 are executed.

12. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the processing method for point cloud data according to any one of claims 1 to 9 are executed.

Citation Information

Patent Citations

  • Point cloud classification method, intelligent terminal and storage medium

    CN108399424A

  • Three-dimensional target detection method based on graph convolution attention network

    CN110674829A

  • Three-dimensional target detection method and device, computer equipment and storage medium

    CN111199206A