Point cloud semantic segmentation method and device
By filtering and fusing target point cloud points in the previous frame point cloud data, the problem of the multi-frame point cloud semantic segmentation method decreases the effect of object motion and radar position motion in the environment, achieving a more efficient and stable semantic segmentation effect.
Patent Information
- Application Number
- CN202311521948.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
When using multi-frame point cloud semantic segmentation methods, existing point cloud semantic segmentation methods are susceptible to the movement of objects in the environment and the radar's own positional motion, resulting in a decrease in semantic segmentation effect.
By obtaining the semantic segmentation results of the previous frame point cloud data, the target point cloud points that meet the preset conditions are selected and fused into the current frame point cloud data, and then semantic segmentation is performed to obtain the second semantic segmentation result.
Effectively using multi-frame point clouds for semantic segmentation, reducing the impact of object motion and radar position motion in the environment on the segmentation effect, and improving the accuracy and stability of semantic segmentation.
Smart Images

Figure CN120014253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a point cloud semantic segmentation method and device. Background Art
[0002] Compared with a single-frame point cloud, a multi-frame point cloud formed by the fusion of multiple single-frame point clouds that are adjacent in time series can provide denser point cloud data. Therefore, multi-frame point clouds have been widely used in various point cloud semantic segmentation tasks.
[0003] However, existing point cloud semantic segmentation methods are often easily affected by the movement of objects in the environment and the movement of the radar's own position when using multi-frame point clouds for semantic segmentation, resulting in a decrease in the semantic segmentation effect. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a point cloud semantic segmentation method and device, so that the fused multi-frame point cloud can be effectively applied to the semantic segmentation task, thereby avoiding the influence of the movement of objects in the environment and the movement of the radar's own position, and improving the semantic segmentation effect.
[0005] In a first aspect, an embodiment of the present invention provides a point cloud semantic segmentation method, the method comprising:
[0006] Obtain the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data;
[0007] Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result;
[0008] Fusion of the target point cloud into the current frame point cloud data;
[0009] The fused point cloud data of the current frame is semantically segmented to determine a second semantic segmentation result.
[0010] In a second aspect, an embodiment of the present invention provides a point cloud semantic segmentation device, the device comprising:
[0011] An acquisition unit, used to acquire a first semantic segmentation result of a previous frame of point cloud data and a current frame of point cloud data;
[0012] A screening unit, configured to determine, in the previous frame of point cloud data, target point cloud points that meet a preset screening condition according to the first semantic segmentation result;
[0013] A fusion unit, used for fusing the target point cloud points into the current frame point cloud data;
[0014] The segmentation unit is used to perform semantic segmentation on the fused point cloud data of the current frame to determine a second semantic segmentation result.
[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method as described in any one of the first aspects is implemented.
[0016] In a fourth aspect, an embodiment of the present invention provides an electronic device, the device comprising:
[0017] a memory for storing one or more computer program instructions;
[0018] A processor, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of the first aspects.
[0019] In a fifth aspect, an embodiment of the present invention provides a computer program product, which, when executed on a computer, enables the computer to execute the method as described in any one of the first aspects.
[0020] The embodiment of the present invention obtains the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data, and determines the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result, and then fuses the target point cloud points into the current frame of point cloud data, and then performs semantic segmentation on the fused current frame of point cloud data to determine the second semantic segmentation result. Therefore, by fusing the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data into the current frame of point cloud data, this embodiment can enable the fused multi-frame point cloud to be effectively applied to the semantic segmentation task, thereby avoiding the influence of the movement of objects in the environment and the movement of the radar's own position, and improving the semantic segmentation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0022] Figure 1 Flow chart of a point cloud semantic segmentation method according to an embodiment of the present invention;
[0023] Figure 2 is a flowchart of a semantic segmentation method according to an embodiment of the present invention;
[0024] Figure 3 Schematic diagram of a point cloud semantic segmentation process according to an embodiment of the present invention;
[0025] Figure 4 Schematic diagram of a point cloud semantic segmentation device according to an embodiment of the present invention;
[0026] Figure 5Schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] The present application is described below based on embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, some specific details are described in detail. It is possible for those skilled in the art to fully understand the present application without the description of these details. In order to avoid confusing the essence of the present application, known methods, processes, flows, components and circuits are not described in detail.
[0028] In addition, persons of ordinary skill in the art will appreciate that the drawings provided herein are for illustration purposes and are not necessarily drawn to scale.
[0029] Unless the context clearly requires otherwise, the words "include", "comprising" and similar words throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, the meaning is "including but not limited to".
[0030] In the description of this application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "plurality" is two or more.
[0031] The solutions described in this specification and in the examples, if they involve the processing of personal information, will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.
[0032] When existing point cloud semantic segmentation methods use multi-frame point clouds for semantic segmentation, they are often easily affected by the movement of objects in the environment and the movement of the radar's own position, resulting in a decrease in the semantic segmentation effect. Specifically, when performing semantic segmentation on multi-frame point clouds, the object category to which each point cloud point in the multi-frame point cloud belongs needs to be accurately predicted. However, due to the influence of the movement of objects in the environment and the movement of the radar's own position, there will inevitably be position deviations in the multi-frame point cloud formed by the fusion of multiple single-frame point clouds that are adjacent in the time series. Among them, some objects with large displacements between the previous and next frames (such as raindrops) will form noise points in the multi-frame point cloud. These noise points will have a greater impact on the prediction accuracy of the object category to which the point cloud points belong, and thus lead to a decrease in the semantic segmentation effect.
[0033] In this regard, an embodiment of the present invention provides a point cloud semantic segmentation method and device, so that the fused multi-frame point cloud can be effectively applied to the semantic segmentation task, thereby avoiding the influence of the movement of objects in the environment and the movement of the radar's own position, and improving the semantic segmentation effect.
[0034] It should be understood that the point cloud data in each embodiment of the present invention may be a data set composed of a large number of point cloud points, wherein the point cloud points may specifically be points in a three-dimensional space, which may be used to characterize the position of the surface points of the relevant object in the three-dimensional space.
[0035] Furthermore, the point cloud data can be collected and acquired by a corresponding radar sensor. Specifically, the radar sensor may include a transmitter and a receiver, the transmitter can emit electromagnetic waves, and the electromagnetic waves will be reflected when they hit the surface of an object, and the receiver can receive the reflected electromagnetic waves and determine the distance between the surface point of the object and the radar sensor by calculating the time difference from the emission of the electromagnetic wave to the reception of the electromagnetic wave, and then determine the position of the surface point of the object in three-dimensional space based on the distance. Thus, the radar sensor can respectively determine the position of the surface points of the surrounding objects in three-dimensional space to form corresponding point cloud data.
[0036] Optionally, the radar sensor may be a laser radar sensor. It should be understood that in some embodiments, the laser radar sensor may also be replaced by other related types of radar sensors, which is not limited in the present application.
[0037] Figure 1 Flow chart of the point cloud semantic segmentation method according to an embodiment of the present invention. Figure 1 As shown, the point cloud semantic segmentation method may specifically include the following steps:
[0038] It should be understood that the execution subject of the point cloud semantic segmentation method can be any general data processing device, such as a tablet computer or a computer, etc. Optionally, in some embodiments, the execution subject of the point cloud semantic segmentation method can also be a data processing device mounted on a transportation device, such as a vehicle-mounted terminal, etc.
[0039] S100, obtaining the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data.
[0040] Specifically, the device can obtain the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data.
[0041] It should be understood that in this step, if the previous frame of point cloud data is the first frame of point cloud data collected by the radar sensor, then the previous frame of point cloud data may only include the point cloud data collected by the radar sensor in the previous frame. If the previous frame of point cloud data is not the first frame of point cloud data collected by the radar sensor, then the previous frame of point cloud data, in addition to the point cloud data collected by the radar sensor in the previous frame, should also include the corresponding target point cloud points selected from the point cloud data of the previous frame that is adjacent to the previous frame in time sequence. For example: if the previous frame of point cloud data is the third frame of point cloud data collected by the radar sensor, then the previous frame of point cloud data, in addition to the point cloud data collected by the radar sensor in the third frame, should also include the corresponding target point cloud points selected from the second frame of point cloud data. Furthermore, the second frame of point cloud data will also include the corresponding target point cloud points selected from the first frame of point cloud data. The current frame of point cloud data may be the point cloud data collected by the radar sensor in the current frame.
[0042] It should be understood that the process of the device determining the first semantic segmentation result of the previous frame point cloud data is the same as the process of determining the second semantic segmentation result of the current frame point cloud data. The details can be referred to the subsequent processing process of the current frame point cloud data, which will not be repeated here.
[0043] S200. Determine, according to the first semantic segmentation result, target point cloud points that meet preset screening conditions in the previous frame of point cloud data.
[0044] Specifically, after obtaining the first semantic segmentation result of the previous frame of point cloud data, the device can determine the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result. It should be understood that the target point cloud points in this embodiment can specifically refer to valid points in the point cloud data, that is, point cloud points that will not interfere with subsequent semantic segmentation and will help improve the semantic segmentation effect. The target point cloud points can be screened by preset screening conditions.
[0045] In an optional implementation, the device may filter the target point cloud points from each previous frame of point cloud points according to the object category to which the point cloud points belong, wherein the previous frame of point cloud points may refer to the point cloud points in the previous frame of point cloud data.
[0046] Specifically, the first semantic segmentation result includes the point cloud category information of each of the previous frame point cloud points. The device can determine the previous frame point cloud points belonging to the preset foreground category in the previous frame point cloud data as the target point cloud points according to the point cloud category information. Among them, the preset foreground category can be the object category that needs to be paid attention to in the semantic segmentation task, which can be specifically determined by relevant operators according to the field of application of the current semantic segmentation task. For example, when the semantic segmentation task is applied in the field of intelligent driving, the preset foreground category can be vehicles, pedestrians, and other road obstacles. Therefore, this embodiment can selectively filter out target point cloud points that can help improve the semantic segmentation effect from the point cloud data according to the object category to which the point cloud points belong.
[0047] It should be understood that in some embodiments, the device can also determine the point cloud points in the previous frame of point cloud data that do not belong to the preset background category as the target point cloud points based on the point cloud category information. Among them, in contrast to the preset foreground category, the preset background category can be an object category that does not need to be paid attention to in the semantic segmentation task and will interfere with the subsequent semantic segmentation. The preset background category can also be determined by relevant operators according to the field in which the current semantic segmentation task is applied. For example, when the semantic segmentation task is applied in the field of intelligent driving, the preset background category can be the road surface, etc. Therefore, this embodiment can specifically filter out noise points in the point cloud data that will interfere with subsequent semantic segmentation according to the object category to which the point cloud points belong, and then retain the corresponding target point cloud points.
[0048] Furthermore, the point cloud category information of each of the previous frame point cloud points can be determined by the device according to the voxel category information of the voxel to which each of the previous frame point cloud points belongs, and the voxel category information can be determined by the device by inputting the voxelized previous frame point cloud data into the semantic segmentation model. Specifically, when semantically segmenting the previous frame point cloud data, the device can first voxelize the previous frame point cloud data to obtain the voxelized previous frame point cloud data, and then input the voxelized previous frame point cloud data into the semantic segmentation model to determine the voxel category information of each voxel in the voxelized previous frame point cloud data, and then determine the voxel category information of the voxel to which each of the previous frame point cloud points belongs as the point cloud category information of each of the previous frame point cloud points. Among them, voxel is the abbreviation of volume pixel, which is the smallest unit that can be segmented after segmenting the three-dimensional space. The voxelization of the previous frame point cloud data can specifically refer to assigning each of the previous frame point cloud points of the previous frame point cloud data to the corresponding voxel.
[0049] Furthermore, the semantic segmentation model used in this embodiment can be a deep learning model, and the learning method of the deep learning model can be fully supervised or semi-supervised, etc. It should be understood that the deep learning model can specifically be a neural network model including any one of the above neural networks or combinations of neural networks, such as deep convolutional neural networks (DCNN), recurrent neural networks (RNN), deep neural networks (DNN), convolutional neural networks (CNN) or residual networks.
[0050] In another optional implementation, the device can filter the target point cloud points in each previous frame of point cloud points based on the average point cloud height information of each voxel in the previous frame of point cloud data after voxelization. Wherein, the average point cloud height information may refer to the average height of each point cloud point belonging to the same voxel. Specifically, the first semantic segmentation result may include the average point cloud height information of each voxel, and the device can determine the voxel whose average point cloud height is greater than or equal to the preset height threshold as the target voxel based on the average point cloud height information, and then determine the previous frame of point cloud points belonging to the target voxel as the target point cloud point. It should be understood that the preset height threshold can be set and adjusted by relevant operating personnel according to actual needs. Therefore, this embodiment can filter the target point cloud points according to the height of the point cloud points.
[0051] In another optional implementation, the device can also filter the target point cloud points in each previous frame of point cloud points according to the voxel thickness information of each voxel in the previous frame of point cloud data after voxelization. Among them, the voxel thickness information may refer to the difference between the maximum height and the minimum height of the point cloud points belonging to the same voxel. Specifically, the first semantic segmentation result may include the voxel thickness information of each voxel, and the device can determine the voxel whose voxel thickness is greater than or equal to the preset thickness threshold as the target voxel based on the voxel thickness information, and then determine the previous frame of point cloud points belonging to the target voxel as the target point cloud point. It should be understood that the preset thickness threshold can be set and adjusted by relevant operating personnel according to actual needs. Thus, this embodiment can filter the target point cloud points according to the thickness of the point cloud points within the voxel.
[0052] In another optional implementation, the device can also filter the target point cloud points in each previous frame of point cloud points according to the speed information of each previous frame of point cloud points. Specifically, the first semantic segmentation result can include the speed information of each previous frame of point cloud points, and the device can determine the previous frame of point cloud points whose speed is greater than or equal to the preset speed threshold as the target point cloud point according to the speed information. It should be understood that the preset speed threshold can be set and adjusted by the relevant operating personnel according to actual needs. Thus, this embodiment can filter the target point cloud points according to the speed of the point cloud points.
[0053] It is desired to explain that the device may also combine the above-mentioned methods to filter the target point cloud points, or the device may also use other related methods to filter the target point cloud points. The present application does not limit the specific method of filtering the target point cloud points.
[0054] S300: Fusing the target point cloud points into the current frame point cloud data.
[0055] Specifically, after determining the target point cloud point, the device can fuse the target point cloud point into the current frame point cloud data.
[0056] Therefore, by fusing the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data into the current frame of point cloud data, the device can remove noise points while retaining the valid points in the previous frame of point cloud data, so that the fused multi-frame point cloud can be effectively applied to semantic segmentation tasks to avoid being affected by the movement of objects in the environment and the movement of the radar's own position, thereby improving the semantic segmentation effect.
[0057] S400 , performing semantic segmentation on the fused point cloud data of the current frame to determine a second semantic segmentation result.
[0058] Specifically, after fusing the target point cloud points into the current frame point cloud data, the device may perform semantic segmentation on the fused current frame point cloud data to determine a second semantic segmentation result.
[0059] Figure 2 FIG. 1 is a flowchart of a semantic segmentation method according to an embodiment of the present invention. Figure 2 As shown, the semantic segmentation method may specifically include the following steps:
[0060] It should be understood that the semantic segmentation method can be specifically used to implement the above step S400.
[0061] S410 , voxelize the fused current frame point cloud data to determine the voxelized current frame point cloud data.
[0062] Specifically, the device may voxelize the fused current frame point cloud data to determine the voxelized current frame point cloud data. Voxel is the abbreviation of volume pixel, which is the smallest unit that can be divided after segmenting the three-dimensional space. The voxelization of the current frame point cloud data may specifically refer to assigning each point cloud point of the current frame point cloud data to a corresponding voxel.
[0063] S420, inputting the voxelized current frame point cloud data into a semantic segmentation model to determine voxel category information of each voxel in the voxelized current frame point cloud data.
[0064] Specifically, the device may input the voxelized current frame point cloud data into a semantic segmentation model. Through the semantic segmentation model, the device may determine the voxel category information of each voxel in the voxelized current frame point cloud data.
[0065] S430 , determining point cloud category information of each point cloud point in the fused point cloud data of the current frame according to the voxel category information.
[0066] Specifically, for each point cloud point in the current frame point cloud data, the device can determine the voxel category information of the voxel to which the point cloud point belongs as the point cloud category information of the point cloud point. Thus, the device can respectively determine the point cloud category information of each point cloud point in the current frame point cloud data.
[0067] It should be understood that when performing semantic segmentation on the next frame of point cloud data, the device will also determine the target point cloud points that meet the preset filtering conditions from the fused current frame point cloud data based on the second semantic segmentation result, and then fuse the target point cloud points into the next frame of point cloud data to perform subsequent semantic segmentation operations.
[0068] Figure 3 FIG. 1 is a schematic diagram of a point cloud semantic segmentation process according to an embodiment of the present invention. Figure 3 As shown, Figure 3 The point cloud semantic segmentation process of the device at time T0, T1 and T2 is shown respectively.
[0069] Specifically, at time T0, the device can determine the point cloud data collected by the radar sensor at time T0 as the first frame point cloud data 311, and voxelize the first frame point cloud data 311 to obtain voxelized first frame point cloud data 312, and then input the voxelized first frame point cloud data 312 into the semantic segmentation model to obtain the first voxel category information 313 of each voxel in the voxelized first frame point cloud data 312, and then determine the first point cloud category information 314 of each point cloud point in the first frame point cloud data 311 according to the first voxel category information 313.
[0070] At time T1, the device can select corresponding target point cloud points 315 from the first frame point cloud data 311 according to the first semantic segmentation result of the first frame point cloud data 311, and fuse the target point cloud points 315 into the point cloud data collected by the radar sensor at time T1 to form the second frame point cloud data 321. After forming the second frame point cloud data 321, the device can voxelize the second frame point cloud data 321 to obtain voxelized second frame point cloud data 322, and then input the voxelized second frame point cloud data 322 into the semantic segmentation model to obtain the second voxel category information 323 of each voxel in the voxelized second frame point cloud data 322, and then determine the second point cloud category information 324 of each point cloud point in the second frame point cloud data 321 according to the second voxel category information 323.
[0071] At time T2, the device can select corresponding target point cloud points 325 from the second frame point cloud data 321 according to the second semantic segmentation result of the second frame point cloud data 321, and fuse the target point cloud points 325 into the point cloud data collected by the radar sensor at time T2 to form the third frame point cloud data 332. After forming the third frame point cloud data 331, the device can voxelize the third frame point cloud data 331 to obtain voxelized third frame point cloud data 332, and then input the voxelized third frame point cloud data 332 into the semantic segmentation model to obtain the third voxel category information 333 of each voxel in the voxelized third frame point cloud data 332, and then determine the third point cloud category information 334 of each point cloud point in the third frame point cloud data 331 according to the third voxel category information 333.
[0072] It should be understood that since the point cloud semantic segmentation process of the device at the subsequent time is the same as the point cloud semantic segmentation process at time T1 and T2, Figure 3 The point cloud semantic segmentation process after time T2 is omitted.
[0073] The embodiment of the present invention obtains the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data, and determines the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result, and then fuses the target point cloud points into the current frame of point cloud data, and then performs semantic segmentation on the fused current frame of point cloud data to determine the second semantic segmentation result. Therefore, by fusing the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data into the current frame of point cloud data, this embodiment can enable the fused multi-frame point cloud to be effectively applied to the semantic segmentation task, thereby avoiding the influence of the movement of objects in the environment and the movement of the radar's own position, and improving the semantic segmentation effect.
[0074] Figure 4 Schematic diagram of a point cloud semantic segmentation device according to an embodiment of the present invention. Figure 4As shown, the point cloud semantic segmentation device of the embodiment of the present invention includes an acquisition unit 41, a screening unit 42, a fusion unit 43 and a segmentation unit 44.
[0075] Specifically, the acquisition unit 41 is used to acquire the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data;
[0076] The screening unit 42 is used to determine the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result;
[0077] The fusion unit 43 is used to fuse the target point cloud points into the current frame point cloud data;
[0078] The segmentation unit 44 is used to perform semantic segmentation on the fused current frame point cloud data to determine a second semantic segmentation result.
[0079] The embodiment of the present invention obtains the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data, and determines the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result, and then fuses the target point cloud points into the current frame of point cloud data, and then performs semantic segmentation on the fused current frame of point cloud data to determine the second semantic segmentation result. Therefore, by fusing the target point cloud points that meet the preset screening conditions in the previous frame of point cloud data into the current frame of point cloud data, this embodiment can enable the fused multi-frame point cloud to be effectively applied to the semantic segmentation task, thereby avoiding the influence of the movement of objects in the environment and the movement of the radar's own position, and improving the semantic segmentation effect.
[0080] Figure 5 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. Figure 5 As shown, Figure 5The electronic device shown is a data processing device in the above embodiment, which includes a general computer hardware structure, which at least includes a processor 51 and a memory 52. The processor 51 and the memory 52 are connected via a bus 53. The memory 52 is suitable for storing instructions or programs executable by the processor 51. The processor 51 can be an independent microprocessor or a collection of one or more microprocessors. Thus, the processor 51 executes the instructions stored in the memory 52, thereby executing the method flow of the embodiment of the present invention as described above to realize the processing of data and the control of other devices. The bus 53 connects the above multiple components together, and at the same time connects the above components to the display controller 54 and the display device and the input / output (I / O) device 55. The input / output (I / O) device 55 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a somatosensory input device, a printer, and other devices known in the art. Typically, the input / output device 55 is connected to the system via an input / output (I / O) controller 56.
[0081] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices (equipment) or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may adopt a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0082] The present application is described with reference to flowcharts of methods, apparatuses (devices) and computer program products according to embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.
[0083] These computer program instructions may be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device that implements the process Figure 1 A function specified in a process or multiple processes.
[0084] These computer program instructions may also be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce the instructions for implementing the process Figure 1 A device that specifies functions in a process or multiple processes.
[0085] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used for a computer to execute part or all of the above method embodiments.
[0086] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by specifying relevant hardware through a program, and the program is stored in a storage medium, including several instructions for a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0087] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A point cloud semantic segmentation method, characterized in that: The method comprises: Obtain the first semantic segmentation result of the previous frame of point cloud data and the current frame of point cloud data; Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result; Fusion of the target point cloud into the current frame point cloud data; The fused point cloud data of the current frame is semantically segmented to determine a second semantic segmentation result.
2. The method according to claim 1, characterized in that The first semantic segmentation result includes point cloud category information of each previous frame point cloud point, and the previous frame point cloud point is the point cloud point in the previous frame point cloud data.
3. The method according to claim 2, characterized in that Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result includes: Determine, according to the point cloud category information, a point cloud point of the previous frame belonging to a preset foreground category as the target point cloud point; or According to the point cloud category information, the point cloud points of the previous frame that do not belong to the preset background category are determined as the target point cloud points.
4. The method according to claim 2, characterized in that: The method further comprises: voxelize the previous frame of point cloud data to determine the voxelized previous frame of point cloud data; Inputting the voxelized last frame of point cloud data into a semantic segmentation model to determine voxel category information of each voxel in the voxelized last frame of point cloud data; The point cloud category information of each point cloud point of the previous frame is determined according to the voxel category information.
5. The method according to claim 1, characterized in that The first semantic segmentation result includes average point cloud height information of each voxel in the previous frame of point cloud data after voxelization.
6. The method according to claim 5, characterized in that Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result includes: Determine, according to the average point cloud height information, voxels whose average point cloud height is greater than or equal to a preset height threshold as target voxels; The last frame point cloud point belonging to the target voxel is determined as the target point cloud point, and the last frame point cloud point is the point cloud point in the last frame point cloud data.
7. The method according to claim 1, characterized in that The first semantic segmentation result includes voxel thickness information of each voxel in the previous frame of point cloud data after voxelization.
8. The method according to claim 7, characterized in that Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result includes: Determining, according to the voxel thickness information, a voxel having a voxel thickness greater than or equal to a preset thickness threshold as a target voxel; The last frame point cloud point belonging to the target voxel is determined as the target point cloud point, and the last frame point cloud point is the point cloud point in the last frame point cloud data.
9. The method according to claim 1, characterized in that: The first semantic segmentation result includes speed information of each previous frame point cloud point, and the previous frame point cloud point is the point cloud point in the previous frame point cloud data.
10. The method according to claim 9, characterized in that Determining target point cloud points that meet preset screening conditions in the previous frame of point cloud data according to the first semantic segmentation result includes: According to the speed information, a point cloud point of a previous frame whose speed is greater than or equal to a preset speed threshold is determined as the target point cloud point.
11. The method according to claim 1, characterized in that: Performing semantic segmentation on the fused current frame point cloud data to determine a second semantic segmentation result includes: voxelize the fused current frame point cloud data to determine the voxelized current frame point cloud data; Inputting the voxelized current frame point cloud data into a semantic segmentation model to determine voxel category information of each voxel in the voxelized current frame point cloud data; The point cloud category information of each point cloud point in the fused current frame point cloud data is determined according to the voxel category information.
12. A point cloud semantic segmentation device, characterized in that: The device comprises: An acquisition unit, used to acquire a first semantic segmentation result of a previous frame of point cloud data and a current frame of point cloud data; A screening unit, configured to determine, in the previous frame of point cloud data, target point cloud points that meet a preset screening condition according to the first semantic segmentation result; A fusion unit, used for fusing the target point cloud points into the current frame point cloud data; The segmentation unit is used to perform semantic segmentation on the fused point cloud data of the current frame to determine a second semantic segmentation result.
13. A computer-readable storage medium storing computer program instructions, characterized in that: The computer program instructions implement the method according to any one of claims 1 to 11 when executed by a processor.
14. An electronic device, characterized in that: The device comprises: a memory for storing one or more computer program instructions; A processor, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 11.
15. A computer program product, characterized in that When the computer program product is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 11.