Sparse marker-based sequence point cloud semantic segmentation method, device and equipment

By employing sparse labeling and scene flow propagation, the problem of training data acquisition efficiency and effectiveness in sequential point cloud semantic segmentation is solved, achieving efficient training data acquisition and accurate label information correction, thereby improving the performance of the point cloud semantic segmentation network.

CN115496899BActive Publication Date: 2025-12-19XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211040503.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-12-19
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

Existing semantic segmentation methods for sequential point clouds are inadequate in terms of efficiency and effectiveness in acquiring training data, especially due to the large number of point cloud frames, which leads to excessive resource consumption.

Method used

The sparse labeling method is adopted. By acquiring the pre-label information of the labeled frames with set intervals, the label information is transmitted to the adjacent point cloud frames using scene flow. The label prediction network is then used to predict and correct the labels to obtain the corrected label information, which is used as training data.

Benefits of technology

This improved the efficiency of training data acquisition, ensured the effectiveness of training data, and enhanced the training effect of point cloud semantic segmentation networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496899B_ABST
    Figure CN115496899B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a sparse label-based sequence point cloud semantic segmentation method, device and equipment. The method comprises: acquiring sequence point cloud data; based on scene flow, sequentially transferring first label information corresponding to each point in the label frame to other point cloud frames adjacent to the label frame except the label frame to obtain second label information corresponding to each point in the other point cloud frames; using a label prediction network to perform label prediction on the other point cloud frames to obtain label prediction results corresponding to each point in the other point cloud frames; performing correction processing on the second label information to obtain corrected second label information; and taking the label frame with the first label information and the other point cloud frames with the corrected second label information as target training data. The technical solution of the embodiments of the present application improves the acquisition efficiency of the training data of the point cloud semantic segmentation network and ensures the effectiveness of the training data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a sequence point cloud semantic segmentation method and device based on sparse markers and equipment. BACKGROUND

[0002] With the development of computer technology, many applications need to perform 3D scene understanding, especially robots, autonomous driving and virtual reality, etc. The sequence point cloud semantic segmentation task is a very important task of three-dimensional perception. In the current technical solution, there are some sequence point cloud semantic segmentation methods, which are usually based on full supervision method. Due to the huge number of points in point cloud, when training the point cloud semantic segmentation network, it is necessary to consume huge resources to annotate each frame of large-scale continuous point cloud. Therefore, how to improve the acquisition efficiency of training data of point cloud semantic segmentation network and ensure the effectiveness of training data has become a technical problem to be solved. SUMMARY

[0003] The embodiments of the present application provide a sequence point cloud semantic segmentation method and device based on sparse markers and equipment, which can at least improve the acquisition efficiency of training data of point cloud semantic segmentation network and ensure the effectiveness of training data to some extent.

[0004] Other characteristics and advantages of the present application will become apparent from the following detailed description, or will be learned by practice of the present application.

[0005] According to one aspect of the embodiments of the present application, a sequence point cloud semantic segmentation method based on sparse markers is provided, comprising:

[0006] Obtaining sequence point cloud data, the sequence point cloud data comprising a plurality of interval set marker frames, each point in the marker frame having pre-labeled first label information;

[0007] Based on the scene flow, the first label information corresponding to each point in the marker frame is sequentially transmitted to other point cloud frames adjacent to the marker frame except the marker frame, to obtain the second label information corresponding to each point in the other point cloud frames;

[0008] Using a label prediction network to perform label prediction on the other point cloud frames to obtain label prediction results corresponding to each point in the other point cloud frames, the label prediction results comprising at least one predicted label and corresponding confidence;

[0009] According to the confidence corresponding to the same predicted label in the label prediction results corresponding to the same point in the other point cloud frames and the second label information, the second label information is corrected to obtain the corrected second label information;

[0010] The marked frame with the first label information and the other point cloud frame with the corrected second label information are taken as target training data.

[0011] According to an aspect of an embodiment of the present application, a sparse marker based sequential point cloud semantic segmentation device is provided, comprising:

[0012] An acquisition module is configured to acquire sequential point cloud data, wherein the sequential point cloud data comprises a plurality of marked frames arranged at intervals, and each point in the marked frame has pre-calibrated first label information.

[0013] A label transmission module is configured to sequentially transmit the first label information corresponding to each point in the marked frame to other point cloud frames adjacent to the marked frame and excluding the marked frame based on scene flow, to obtain second label information corresponding to each point in the other point cloud frames.

[0014] A prediction module is configured to perform label prediction on the other point cloud frames by using a label prediction network, to obtain label prediction results corresponding to each point in the other point cloud frames, wherein the label prediction results comprise at least one predicted label and corresponding confidence.

[0015] A correction module is configured to perform correction processing on the second label information according to the confidence corresponding to the predicted label same as the second label information in the label prediction results corresponding to the same point in the other point cloud frames, to obtain corrected second label information.

[0016] A processing module is configured to take the marked frame with the first label information and the other point cloud frame with the corrected second label information as target training data.

[0017] According to an aspect of an embodiment of the present application, a computer readable medium having a computer program stored thereon is provided, wherein the computer program is executed by a processor to implement the sparse marker based sequential point cloud semantic segmentation method as described in the above embodiments.

[0018] According to an aspect of an embodiment of the present application, an electronic device is provided, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the sparse marker based sequential point cloud semantic segmentation method as described in the above embodiments.

[0019] According to an aspect of some embodiments of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the sparse marker based sequential point cloud semantic segmentation method provided in the above embodiments.

[0020] In the technical solutions provided in some embodiments of the present application, by obtaining sequential point cloud data, the sequential point cloud data includes a plurality of interval arranged marker frames, each point in the marker frame has pre-labeled first label information, the first label information corresponding to each point in the marker frame is sequentially transferred to other point cloud frames adjacent to the marker frame except the marker frame based on scene flow to obtain second label information corresponding to each point in the other point cloud frames, a label prediction network is used to perform label prediction on the other point cloud frames to obtain label prediction results corresponding to each point in the other point cloud frames, the label prediction results include at least one predicted label and corresponding confidence, the second label information is corrected according to the confidence corresponding to the predicted label same as the second label information in the label prediction results corresponding to the same point in the other point cloud frames to obtain corrected second label information, and the marker frame with the first label information and the other point cloud frames with the corrected second label information are used as target training data. Thus, the label transfer is performed based on the scene flow, the point cloud frames with label information can be obtained without labeling each point cloud frame, the acquisition efficiency of the training data is improved, at the same time, the transferred label information is corrected based on the label prediction results corresponding to the label prediction network, the accuracy of the label information is improved, and the effectiveness of the training data is ensured.

[0021] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present application. BRIEF DESCRIPTION OF DRAWINGS

[0022] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. It is clear that the accompanying drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings. In the drawings:

[0023] Figure 1 Fig. 1 shows a flowchart of a sparse marker based sequential point cloud semantic segmentation method according to an embodiment of the present application;

[0024] Figure 2A block diagram of a sparse marker based sequential point cloud semantic segmentation apparatus is shown according to an embodiment of the present application;

[0025] Figure 3 A structural schematic diagram of a computer system of an electronic device suitable for implementing embodiments of the present application is shown. DETAILED DESCRIPTION

[0026] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these example implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art.

[0027] Moreover, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the application. One skilled in the relevant art will recognize, however, that the

[0028] The block diagrams in the drawings show only the functional entities and do not necessarily imply a physical structure for the implementation. That is, the functional entities can be implemented in software, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0029] The flow diagrams shown in the drawings are merely examples and do not necessarily include all of the steps and / or operations, nor do the steps and / or operations necessarily need to be performed in the order shown. For example, some operations can be performed in parallel or in a different order, and some operations can be omitted or combined, so the actual order can vary from what is described.

[0030] Figure 1 A flow diagram of a sparse marker based sequential point cloud semantic segmentation method is shown according to an embodiment of the present application. The method can be applied in a terminal or a server, where the terminal can include, but is not limited to, one or more of a smartphone, a tablet computer, a portable computer, or a desktop computer; the server can be a physical server or a cloud server. It is noted that the present application does not limit the number of terminals or servers, for example, the server can be a single server, or a server cluster composed of multiple servers, etc.

[0031] Reference Figure 1As shown, the sparse marker-based sequence point cloud semantic segmentation method at least includes steps S110 to S150, which are described in detail as follows:

[0032] In step S110, sequence point cloud data is acquired, wherein the sequence point cloud data includes a plurality of marker frames arranged at intervals, and each point in the marker frame has pre-labeled first label information.

[0033] In this embodiment, the sequence point cloud data can be pre-acquired continuous point cloud data, which can include a plurality of continuous point cloud frames. In the plurality of continuous point cloud frames, a marker frame can be determined at a predetermined interval, for example, every five point cloud frames, or a marker frame, etc. For the marker frame, a manual labeling method can be used to label the marker frame to obtain the first label information corresponding to each point in the marker frame. In other embodiments, other full-supervised labeling methods can also be used to ensure the accuracy of the first label information.

[0034] In step S120, the first label information corresponding to each point in the marker frame is sequentially transferred to other point cloud frames adjacent to the marker frame except the marker frame based on a scene flow, to obtain second label information corresponding to each point in the other point cloud frames.

[0035] In this embodiment, the scene flow can track the 3D motion between the same point in adjacent point cloud frames, which can provide a powerful representation across frames by describing the correlation of the motion state of each point. Thus, based on the scene flow, the first label information corresponding to each point in the marker frame can be sequentially transferred to other point cloud frames adjacent to the marker frame except the marker frame, to obtain the second label information corresponding to each point in the other point cloud frames. It should be noted that the sequential transfer described in this application refers to the transfer between adjacent point cloud frames, for example, there are continuous point cloud frames 1, 2, 3, 4 and 5, wherein the point cloud frame 1 is the marker frame, and the label information of the point cloud frame 1 can be transferred to the point cloud frame 2 based on the scene flow. After determining the label information of the point cloud frame 2, the label information of the point cloud frame 2 can be transferred to the point cloud frame 3 based on the scene flow, and so on.

[0036] In step S130, a label prediction network is used to predict the labels of the other point cloud frames to obtain label prediction results corresponding to each point in the other point cloud frames, wherein the label prediction results include at least one predicted label and a corresponding confidence.

[0037] In this embodiment, the label prediction network can be a pre-trained point cloud semantic segmentation network, which can be trained by a person skilled in the art according to existing manually annotated training data, or a point cloud semantic segmentation network provided by other third parties. Through the label prediction network, label prediction can be performed on other point cloud frames in the sequence point cloud data except the marker frame, so as to obtain the label prediction result corresponding to each point in the other point cloud frames. The label prediction result includes at least one predicted label and the corresponding confidence. It should be understood that when performing point cloud semantic segmentation, the label classification can be multiple, for example, for point cloud semantic segmentation of a table, the label can be pre-divided into table top, table leg, etc. Therefore, by using the label prediction network to identify the point cloud frame, the confidence of each predicted label corresponding to each point in the point cloud frame can be correspondingly output. For example, the confidence of a point A corresponding to the table top (predicted label) is 70%, the confidence corresponding to the table leg is 45%, etc. Therefore, each point can correspond to multiple predicted labels and confidences.

[0038] In step S140, the second label information is corrected according to the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frames as the second label information, to obtain corrected second label information.

[0039] In this embodiment, the label prediction result of each point in the other point cloud frames can be compared with the second label information, the same predicted label in the label prediction result as the second label information is determined, and the confidence of the predicted label is determined. The second label information is corrected based on the confidence of the same predicted label as the second label information, to obtain corrected second label information. Thus, by comparing the label prediction result with the second label information, the second label information obtained based on the scene flow transmission can be corrected, thereby improving the accuracy of the second label information.

[0040] In step S150, the marker frame with the first label information and the other point cloud frames with the corrected second label information are used as target training data.

[0041] In this embodiment, the marker frame with the first label information and the other point cloud frames with the corrected second label information are used as target training data, only a small amount of marker frames need to be labeled to obtain target training data with label information, which can improve the efficiency of obtaining training data. At the same time, the second label information is corrected based on the label prediction result, which can ensure the accuracy of the second label information obtained based on the scene flow transmission, and ensure the effectiveness of the target training data.

[0042] Based on Figure 1In the embodiment shown, in one embodiment of the present application, the second label information is corrected according to the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information, to obtain corrected second label information, including:

[0043] If the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information is less than a predetermined threshold, the second label information is removed, otherwise, the second label information is retained.

[0044] In this embodiment, if the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information is less than a predetermined threshold, the second label information is removed, and if the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information is greater than or equal to a predetermined threshold, the second label information is retained. In this way, the second label information with a higher error probability can be removed, and the second label information with a higher correct probability can be retained. It should be noted that the predetermined threshold can be a threshold information set by a person skilled in the art according to prior experience, and the predetermined threshold can be used as a basis for determining whether the second label information obtained based on the scene stream is correct.

[0045] Based on the above embodiment, in one embodiment of the present application, before the second label information is corrected according to the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information, to obtain corrected second label information, the method further includes:

[0046] Determining a predetermined threshold corresponding to the second label information according to the number of points in the other point cloud frame corresponding to the same second label information.

[0047] In this embodiment, a person skilled in the art can set different predetermined thresholds according to the number of different points, for example, different predetermined thresholds can be determined according to the numerical interval in which the number of points is located, or a threshold value can be set for points less than a certain number, and another threshold value can be set for points greater than the number, and the like.

[0048] It should be understood that the fewer the number of points corresponding to a certain label information, the more likely it is to become missing or sparse when transmitting based on the scene stream, therefore, the label information of the points with a smaller number should be improved in terms of preservation rate, therefore, the fewer the number of points, the lower the corresponding predetermined threshold, and vice versa, the larger the predetermined threshold, thereby ensuring the accuracy of the point cloud semantic segmentation network obtained through subsequent training.

[0049] Based on Figure 1In the embodiment shown, in one embodiment of the present application, the first label information corresponding to each point in the mark frame is sequentially transmitted to other point cloud frames adjacent to the mark frame and other than the mark frame based on the scene flow, to obtain second label information corresponding to each point in the other point cloud frames.

[0050] The scene flow estimation network is used to determine the predicted positions of the points in the previous point cloud frame in the subsequent point cloud frame, the previous point cloud frame being the mark frame or the other point cloud frame for which the second label information has been determined.

[0051] The second label information of the points in the subsequent point cloud frame is determined according to the first label information or the second label information corresponding to the points in the previous point cloud frame, the predicted positions in the subsequent point cloud frame, and the actual positions of the points in the subsequent point cloud frame.

[0052] In this embodiment, the scene flow estimation network can be a network model for determining the motion relationship of the points in the adjacent point cloud frames. Through the scene flow estimation network, the predicted positions of the points in the previous point cloud frame in the subsequent point cloud frame can be determined. It should be understood that the transmission is between adjacent point cloud frames, and therefore the previous point cloud frame in the adjacent point cloud frames is the mark frame or the other point cloud frame for which the second label information has been determined.

[0053] In an example, considering the computational complexity and the accuracy of scene flow estimation, a PointPWC network (i.e., a scene flow estimation network) can be used for self-supervised scene flow estimation. The PointPWC network uses a self-supervised learning objective function to learn the scene flow in 3D point clouds without labels. Its loss function includes charmer distance, smoothness constraint, and Laplacian regularization. The input of the network is a pair of 3D point clouds S t and S t+1 , and the output is an n x 3 motion vector V t representing the points in S t . Thus, by adding the coordinates of the points in frame S t and the motion vector V t , S′ t is obtained, i.e., the estimated positions of the points in frame S t at time t+1. Thus, according to the predicted positions and the actual positions of the points in S t+1 , the corresponding points have the same label information, so as to perform label transmission, for example, the second label information corresponding to each point in the subsequent point cloud frame is the same as the label information of the point in the previous point cloud frame closest to the predicted position of the point, and so on. In other embodiments, other self-supervised scene flow estimation networks can also be used, which are not particularly limited in the present application.

[0054] Thus, based on the predicted positions of the points in the previous point cloud frame in the subsequent point cloud frame and the actual positions of the points in the subsequent point cloud frame, the correspondence between the points can be determined to ensure the correctness of the label transmission.

[0055] In an embodiment of the present application, the second label information of the points in the subsequent point cloud frame is determined according to the first label information or the second label information corresponding to the points in the previous point cloud frame, the predicted positions of the points in the subsequent point cloud frame, and the actual positions of the points in the subsequent point cloud frame.

[0056] The predetermined number of points in the previous point cloud frame closest to the predicted positions and the actual positions of the points in the subsequent point cloud frame are determined as candidate points.

[0057] The first label information or the second label information with the highest repetition frequency among the candidate points corresponding to the points in the subsequent point cloud frame is determined as the second label information corresponding to the points in the subsequent point cloud frame.

[0058] In this embodiment, to improve the correctness of the label transmission, the predetermined number of points in the previous point cloud frame closest to the predicted positions and the actual positions of the points in the subsequent point cloud frame are determined as candidate points. The predetermined number can be pre-set by a person skilled in the art according to prior experience, for example, the predetermined number can be 8 or 10.

[0059] Then, the first label information or the second label information of the candidate point with the highest repetition frequency of the label information is determined as the second label information corresponding to the point in the subsequent point cloud frame, thereby improving the correctness of the label transmission.

[0060] Based on the foregoing embodiments, in an embodiment of the present application, the first label information corresponding to the points in the marker frame is sequentially transmitted to other point cloud frames adjacent to the marker frame except the marker frame based on the scene flow, to obtain the second label information corresponding to the points in the other point cloud frames.

[0061] The first label information corresponding to the points in the adjacent marker frames is sequentially transmitted from both ends to the middle based on the scene flow, to determine the second label information corresponding to the points in the other point cloud frames between the adjacent two marker frames.

[0062] In this embodiment, it should be understood that one-way propagation is the iteration of transferring label information from the first frame to other frames. However, the propagated label becomes sparse and unreliable as the iteration proceeds, and thus, to reduce the number of iterations, the first label information of each point in adjacent marked frames can be sequentially transferred from both ends to the middle, thereby reducing the number of iterations and ensuring the effectiveness of label information transmission. For example, the consecutive point cloud frames are 1, 2, 3, 4, 5 and 6, at this time, the point cloud frame 1 and the point cloud frame 6 are marked frames, then the label information of the point cloud frame 1 can be sequentially transferred to the point cloud frame 2 and the point cloud frame 3, and the label information of the point cloud frame 6 is sequentially transferred to the point cloud frame 5 and the point cloud frame 4, so as to reduce the number of transmissions and ensure the effectiveness of label transmission.

[0063] Based on the foregoing embodiments, in an embodiment of the present application, after the target training data of the marked frame with the first label information and the other point cloud frame with the corrected second label information are obtained, the method further comprises:

[0064] Training a pre-constructed point cloud semantic segmentation network according to the target training data to obtain a target point cloud semantic segmentation network.

[0065] In this embodiment, any pre-constructed point cloud semantic segmentation network can be trained according to the target training data to obtain a target point cloud semantic segmentation network, thereby improving the efficiency of obtaining the target training data and ensuring the effectiveness of the target training data, and further ensuring the semantic segmentation accuracy of the point cloud semantic segmentation network obtained by training.

[0066] The following describes an apparatus embodiment of the present application, which can be used to execute the sparse marker-based sequential point cloud semantic segmentation method in the above embodiments of the present application. For details not disclosed in the apparatus embodiment of the present application, please refer to the above embodiments of the sparse marker-based sequential point cloud semantic segmentation method of the present application.

[0067] Figure 2 A block diagram of a sparse marker-based sequential point cloud semantic segmentation apparatus according to an embodiment of the present application is shown.

[0068] Referring to Figure 2 According to an embodiment of the present application, the sparse marker-based sequential point cloud semantic segmentation apparatus comprises:

[0069] The acquisition module 210 is configured to acquire sequential point cloud data, wherein the sequential point cloud data comprises a plurality of marked frames arranged at intervals, and each point in the marked frame has pre-labeled first label information.

[0070] a label transmission module 220 configured to transmit first label information corresponding to each point in the mark frame to other point cloud frames adjacent to the mark frame in sequence based on a scene flow, to obtain second label information corresponding to each point in the other point cloud frames;

[0071] a prediction module 230 configured to perform label prediction on the other point cloud frames by using a label prediction network to obtain label prediction results corresponding to each point in the other point cloud frames, the label prediction results including at least one predicted label and a corresponding confidence;

[0072] a correction module 240 configured to correct the second label information according to a confidence corresponding to a predicted label same as the second label information in the label prediction results corresponding to the same point in the other point cloud frames, to obtain corrected second label information;

[0073] a processing module 250 configured to use the mark frame with the first label information and the other point cloud frames with the corrected second label information as target training data.

[0074] In an embodiment of the present application, the correction module 240 is configured to remove the second label information if a confidence corresponding to a predicted label same as the second label information in the label prediction results corresponding to the same point in the other point cloud frames is less than a predetermined threshold, and otherwise, retain the second label information.

[0075] In an embodiment of the present application, the correction module 240 is further configured to determine the predetermined threshold corresponding to the second label information according to a number of points in the other point cloud frames corresponding to the same second label information.

[0076] In an embodiment of the present application, the label transmission module 220 is configured to determine predicted positions of points in a previous point cloud frame in a subsequent point cloud frame by using a scene flow estimation network, the previous point cloud frame being the mark frame or the other point cloud frame with the determined second label information; and determine the second label information of the points in the subsequent point cloud frame according to the first label information or the second label information of the points in the previous point cloud frame, the predicted positions of the points in the subsequent point cloud frame, and actual positions of the points in the subsequent point cloud frame.

[0077] In an embodiment of the present application, the label transmission module 220 is configured to determine, according to the predicted positions of the points in the previous point cloud frame in the subsequent point cloud frame and the actual positions of the points in the subsequent point cloud frame, a predetermined number of points closest to the predicted positions of the points in the previous point cloud frame and the points in the subsequent point cloud frame as candidate points respectively; and determine the second label information of the points in the subsequent point cloud frame according to a first label information or a second label information with the most repetitions in the candidate points corresponding to the points in the subsequent point cloud frame.

[0078] In an embodiment of the present application, the label passing module 220 is configured to pass the first label information corresponding to each point in the adjacent mark frames from both ends to the middle in sequence to determine second label information corresponding to each point in other point cloud frames between the adjacent two mark frames based on the scene flow.

[0079] In an embodiment of the present application, the processing module 250 is further configured to train the pre-constructed point cloud semantic segmentation network according to the target training data to obtain a target point cloud semantic segmentation network.

[0080] Figure 3 A structural schematic diagram of a computer system of an electronic device suitable for implementing embodiments of the present application is shown.

[0081] It should be noted that, Figure 3 The computer system of the electronic device shown is only an example and should not impose any limitation on the functions and use range of embodiments of the present application.

[0082] As Figure 3 shown, the computer system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 302 or loaded from a storage portion 308 into a random access memory (RAM) 303, such as performing the methods described in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0083] The following components are connected to the I / O interface 305: an input section 306 including input devices such as a keyboard and mouse; an output section 307 including output devices such as a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), and a speaker; a storage section 308 including a hard disk; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, and the like. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as necessary. A removable medium 311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 310 as necessary, so that a computer program read therefrom is installed into the storage section 308 as necessary.

[0084] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present application. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing a computer program for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the system of the present application are executed.

[0085] It should be noted that the computer-readable medium in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (Compact Disk Read-Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present application, the computer-readable signal medium can include a data signal carrying computer-readable computer programs in a baseband or as a part of a carrier wave. Such a propagated data signal can take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate or transmit programs for use by or in conjunction with an instruction execution system, device or apparatus. The computer programs contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, or the like, or any suitable combination thereof.

[0086] The flowcharts and block diagrams in the drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In the flowcharts or block diagrams, each block can represent a module, a program segment or a part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than that shown in the drawings. For example, two blocks represented in succession can actually be executed substantially in parallel, and sometimes in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams or flowcharts, and the combination of blocks in the block diagrams or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0087] The units described in the embodiments of the present application can be implemented by software, or by hardware, or by a combination of software and hardware. The units described can also be located in a single processor. In some cases, the names of the units do not limit the units themselves.

[0088] As another aspect, the present application also provides a computer readable medium, which can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device. The computer readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to implement the method described in the above embodiments.

[0089] It should be noted that although several modules or units for performing actions are mentioned in the above detailed description, the division into the modules or units is not mandatory. In fact, according to the embodiments of the present application, features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, features and functions of one module or unit described above can be further divided into a plurality of modules or units.

[0090] From the above description of the embodiments, those skilled in the art will readily appreciate that the example embodiments described herein can be implemented by software and / or by hardware coupled with software. Accordingly, the technical solutions of the embodiments of the present application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, or the like) or on a network, and includes a number of instructions for causing a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to perform the methods according to the embodiments of the present application.

[0091] Other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the application following, in general, the principles of the application and including such

[0092] It should be understood that the present application is not limited to the precise construction that has been described above and illustrated in the accompanying drawings, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present application. The scope of the present application is limited only by the appended claims.

Claims

1. A sparse marker based sequential point cloud semantic segmentation method, characterized in that, The method comprises the following steps: obtaining sequence point cloud data, wherein the sequence point cloud data comprises a plurality of interval setting mark frames, and each point in the mark frame has pre-labeled first label information; based on the scene flow, the first label information corresponding to each point in the mark frame is sequentially transmitted to other point cloud frames adjacent to the mark frame except the mark frame, to obtain second label information corresponding to each point in the other point cloud frames; using a label prediction network to perform label prediction on the other point cloud frames, to obtain label prediction results corresponding to each point in the other point cloud frames, wherein the label prediction results comprise at least one predicted label and corresponding confidence; according to the confidence of the same predicted label in the label prediction results corresponding to the same point in the other point cloud frames and the second label information, the second label information is corrected to obtain corrected second label information; the mark frame with the first label information and the other point cloud frames with the corrected second label information are used as target training data; wherein, according to the confidence of the same predicted label in the label prediction results corresponding to the same point in the other point cloud frames and the second label information, the second label information is corrected to obtain corrected second label information, comprising: if the confidence of the same predicted label in the label prediction results corresponding to the same point in the other point cloud frames and the second label information is less than a predetermined threshold, the second label information is removed, otherwise, the second label information is retained; wherein, based on the scene flow, the first label information corresponding to each point in the mark frame is sequentially transmitted to other point cloud frames adjacent to the mark frame except the mark frame, to obtain second label information corresponding to each point in the other point cloud frames, comprising: using a scene flow estimation network to determine the predicted position of each point in the previous point cloud frame in the subsequent point cloud frame, wherein the previous point cloud frame is a mark frame or a point cloud frame with determined second label information; according to the first label information or the second label information corresponding to each point in the previous point cloud frame, the predicted position in the subsequent point cloud frame and the actual position of each point in the subsequent point cloud frame, the second label information of each point in the subsequent point cloud frame is determined; wherein, after the mark frame with the first label information and the other point cloud frames with the corrected second label information are used as target training data, the method further comprises: training a pre-constructed point cloud semantic segmentation network according to the target training data to obtain a target point cloud semantic segmentation network.

2. The method of claim 1, wherein, Before the second label information is corrected according to the confidence of the same predicted label in the label prediction results corresponding to the same point in the other point cloud frames and the second label information, the method further comprises: determining a predetermined threshold corresponding to the second label information according to the number of points corresponding to the same second label information in the other point cloud frames.

3. The method of claim 2, wherein, According to the first label information or the second label information corresponding to each point in the previous point cloud frame, the predicted position in the subsequent point cloud frame, and the actual position of each point in the subsequent point cloud frame, the second label information of each point in the subsequent point cloud frame is determined, comprising: According to the predicted position of each point in the previous point cloud frame in the subsequent point cloud frame and the actual position of each point in the subsequent point cloud frame, a predetermined number of points closest to the predicted position in the previous point cloud frame and each point in the subsequent point cloud frame are determined as candidate points respectively; According to the first label information or the second label information with the most repetitions in the candidate points corresponding to each point in the subsequent point cloud frame as the second label information corresponding to each point in the subsequent point cloud frame.

4. The method of claim 1, wherein, Based on the scene flow, the first label information corresponding to each point in the mark frame is sequentially transmitted to other point cloud frames adjacent to the mark frame except the mark frame, and the second label information corresponding to each point in the other point cloud frame is obtained, comprising: Based on the scene flow, the first label information corresponding to each point in the adjacent mark frame is sequentially transmitted from both ends to the middle to determine the second label information corresponding to each point in the other point cloud frame between the two adjacent mark frames.

5. A semantic segmentation device for sequence point clouds based on sparse labels, characterized in that, Comprising: The acquisition module is used for acquiring sequence point cloud data, and the sequence point cloud data includes a plurality of mark frames arranged at intervals, and each point in the mark frame has pre-calibrated first label information; The label transmission module is used for sequentially transmitting the first label information corresponding to each point in the mark frame to other point cloud frames adjacent to the mark frame except the mark frame based on the scene flow, and obtaining the second label information corresponding to each point in the other point cloud frame; The prediction module is used for performing label prediction on the other point cloud frame by using a label prediction network to obtain a label prediction result corresponding to each point in the other point cloud frame, and the label prediction result includes at least one predicted label and a corresponding confidence; The correction module is used for correcting the second label information according to the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information to obtain corrected second label information; The processing module is used for taking the mark frame with the first label information and the other point cloud frame with the corrected second label information as target training data; Wherein, the second label information is corrected according to the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information to obtain corrected second label information, comprising: If the confidence of the same predicted label in the label prediction result corresponding to the same point in the other point cloud frame as the second label information is less than a predetermined threshold, the second label information is removed, otherwise, the second label information is retained; Wherein, based on the scene flow, the first label information corresponding to each point in the mark frame is sequentially transmitted to other point cloud frames adjacent to the mark frame except the mark frame, and the second label information corresponding to each point in the other point cloud frame is obtained, comprising: Based on the scene flow, the first label information corresponding to each point in the adjacent mark frame is sequentially transmitted from both ends to the middle to determine the second label information corresponding to each point in the other point cloud frame between the two adjacent mark frames. determining, by a scene flow estimation network, a predicted position of each point in a previous point cloud frame in a subsequent point cloud frame, the previous point cloud frame being a labeled frame or another point cloud frame for which second label information has been determined; determining, according to the first label information or the second label information corresponding to each point in the previous point cloud frame, the predicted position in the subsequent point cloud frame, and an actual position of each point in the subsequent point cloud frame, second label information of each point in the subsequent point cloud frame; wherein after the labeled frame with the first label information and the other point cloud frame with the corrected second label information are used as target training data, the method further comprises: training a previously constructed point cloud semantic segmentation network according to the target training data to obtain a target point cloud semantic segmentation network.

6. A computer readable medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the sparse label based sequential point cloud semantic segmentation method according to any one of claims 1 to 4.

7. An electronic device, comprising: comprise: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the sparse label based sequential point cloud semantic segmentation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • End-to-end semantic instant-positioning and graph building method based deep learning

    CN108665496A

  • A monocular image depth estimation method and apparatus fusing sparse known tags

    CN109461178A