Occupy tag generation method, apparatus, electronic device and storage medium

By splicing and semantically processing dynamic and static objects in multi-frame point cloud data, high-density and high-accuracy occupancy labels are generated, solving the problem of insufficient accuracy of occupancy labels in existing technologies and improving the performance of visual perception networks in autonomous driving environments.

CN117315260BActive Publication Date: 2026-04-17IFLYTEK CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2023-10-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies generate occupancy labels with low accuracy, especially when dealing with dynamic objects. The insufficient density and accuracy of point cloud data lead to poor recognition performance of occupancy grid networks.

Method used

By identifying key and non-key frames in multi-frame point cloud data, the point cloud data associations of dynamic objects are determined and stitched together to generate a third point cloud data with high density. This data is then combined with the point cloud data of static objects and finally, accurate occupancy labels are generated based on semantic information.

Benefits of technology

It improves the accuracy of occupancy labels, enabling them to more accurately reflect the occupancy status of objects in a scene and supporting decision-making modules for applications such as autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315260B_ABST
    Figure CN117315260B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, electronic device, and storage medium for generating occupancy tags. The method includes: for each non-key frame point cloud data in multi-frame point cloud data, determining the target key frame point cloud data whose acquisition time is closest to that of the non-key frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object; based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determining the second point cloud data in the non-key frame point cloud data corresponding to the first point cloud data; concatenating the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object; and based on the third point cloud data, determining the occupancy tags corresponding to the multi-frame point cloud data. Based on this, the gaps between point cloud data can be reduced, resulting in third point cloud data with high density and accuracy, thereby generating highly accurate occupancy tags.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for generating occupancy tags. Background Technology

[0002] Occupancy labels can be ground truth values ​​for occupancy grids, and the accuracy of occupancy labels is one of the factors affecting the training of occupancy grid network algorithms. For example, in the field of autonomous driving, algorithms based on occupancy grids, such as visual perception networks, are used. The higher the accuracy of the occupancy labels generated from data collected in the autonomous driving environment, the better it is for improving the accuracy of the trained visual perception network.

[0003] The data obtained by scanning the surface of an object can be point cloud data, which represents that object. Point cloud data includes all the points that make up the point cloud. When generating occupancy labels based on point cloud data, the higher the density of the point cloud data, the more details it reflects about each object in the scene, thus more accurately representing the occupancy status within the scene. Furthermore, the higher the accuracy of the point cloud data reflecting the same object, the more accurately it reflects the object's occupancy status in the scene. Therefore, occupancy labels generated from point cloud data with both high density and high accuracy will have higher accuracy.

[0004] In related technologies, to improve the density of point cloud data, a common approach is to directly convert multi-frame LiDAR point sequences into a unified coordinate system and then voxelize the connected dense points to obtain point cloud data with high density. However, due to the differences in the motion states of objects in a scene, the relative positions of dynamic objects are not the same in each frame. Therefore, directly overlaying multi-frame point sequences using this processing method results in point cloud data with low accuracy. Furthermore, when directly overlaying and stitching, the point sequences of dynamic objects are sparse and dispersed, and the density of the point cloud data remains low. Consequently, the accuracy of occupancy labels generated from point cloud data obtained in this way is also low. Summary of the Invention

[0005] This invention provides a method, apparatus, electronic device, and storage medium for generating occupancy tags, in order to solve the problem of low accuracy of occupancy tags generated in the prior art and to improve the accuracy of occupancy tags.

[0006] This invention provides a method for generating occupancy tags, comprising:

[0007] For each non-key frame point cloud data in multi-frame point cloud data, determine the target key frame point cloud data that is closest to the acquisition time of the non-key frame point cloud data from all key frame point cloud data in multi-frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object.

[0008] Based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determine the second point cloud data in the non-key frame point cloud data that corresponds to the first point cloud data;

[0009] The first point cloud data and the second point cloud data are concatenated to obtain the third point cloud data corresponding to the dynamic object;

[0010] Based on the third point cloud data, determine the occupancy labels corresponding to multiple frames of point cloud data.

[0011] According to the occupancy tag generation method provided by the present invention, based on third point cloud data, the occupancy tags corresponding to multiple frames of point cloud data are determined, including:

[0012] For each key frame point cloud data in the multi-frame point cloud data, the fourth point cloud data corresponding to the static object in the key frame point cloud data is determined based on the point cloud data in each bounding box in the key frame point cloud data.

[0013] Based on the second point cloud data in the non-keyframe point cloud data, determine the fifth point cloud data corresponding to the static object in the non-keyframe point cloud data.

[0014] The fourth and fifth cloud data points are concatenated to obtain the sixth cloud data point corresponding to the static object.

[0015] Based on the third and sixth point cloud data, the occupancy labels corresponding to the multi-frame point cloud data are determined.

[0016] According to the present invention, an occupancy tag generation method determines occupancy tags corresponding to multiple frames of point cloud data based on third and sixth point cloud data, including:

[0017] For each target point cloud data without semantic meaning in the sixth point cloud data, determine the semantic point cloud data that is closest to the target point cloud data in the sixth point cloud data;

[0018] The semantics of the semantic point cloud data are determined as the semantics of the target point cloud data;

[0019] Based on the semantics of the third point cloud data, the semantics of the semantic point cloud data, and the semantics of the target point cloud data, the occupancy labels corresponding to multiple frames of point cloud data are determined.

[0020] According to a method for generating occupancy tags provided by the present invention, fourth point cloud data and fifth point cloud data are concatenated to obtain sixth point cloud data corresponding to a static object, including:

[0021] The fourth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the seventh point cloud data in the world coordinate system;

[0022] The fifth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the eighth point cloud data in the world coordinate system;

[0023] By concatenating the cloud data at the seventh and eighth points, we obtain the cloud data at the sixth point corresponding to the static object.

[0024] According to the occupancy tag generation method provided by the present invention, based on the first point cloud data in the first target bounding box in the target keyframe point cloud data, the method determines the second point cloud data corresponding to the first point cloud data in the non-keyframe point cloud data, including:

[0025] Based on the position of the first target bounding box, a second target bounding box corresponding to the position of the first target bounding box is determined in the non-keyframe point cloud data, and the area of ​​the second target bounding box is greater than the area of ​​the first target bounding box.

[0026] For each first point cloud data, determine the distance between the first point cloud data and each point cloud data in the second target bounding box, and determine the point cloud data whose minimum distance is less than the preset distance as the second point cloud data.

[0027] According to the present invention, a method for generating an occupation tag further includes:

[0028] Obtain the image corresponding to the keyframe point cloud data;

[0029] Determine the third target bounding box corresponding to the target dynamic object in the image, and determine the first semantic segmentation mask of the target dynamic object in the third target bounding box based on the semantics of the ninth point cloud data corresponding to the target dynamic object in the key frame point cloud data.

[0030] The third point cloud data is projected onto the first semantic segmentation mask. If there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask, the mismatched point cloud data is deleted.

[0031] According to the present invention, a method for generating an occupation tag further includes:

[0032] Obtain the image corresponding to the keyframe point cloud data;

[0033] The fourth point cloud data is projected onto the image, and the second semantic segmentation mask corresponding to the static object in the image is determined based on the semantics of the fourth point cloud data in the key frame point cloud data.

[0034] The sixth point cloud data is projected onto the second semantic segmentation mask. If there are point cloud data in the sixth point cloud data that do not semantically match the second semantic segmentation mask, the mismatched point cloud data is deleted.

[0035] The present invention also provides an occupation tag generation device, comprising:

[0036] The first determining module is used to determine, from all the key frame point cloud data in the multi-frame point cloud data, the target key frame point cloud data that is closest to the acquisition time of the non-key frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object.

[0037] The second determining module is used to determine the second point cloud data in the non-key frame point cloud data that corresponds to the first point cloud data based on the first point cloud data in the first target bounding box in the target key frame point cloud data.

[0038] The splicing module is used to splice the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object;

[0039] The third determination module is used to determine the occupancy labels corresponding to multiple frames of point cloud data based on the third point cloud data.

[0040] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-described methods for generating occupancy tags.

[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the above-described methods for generating occupancy tags.

[0042] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the above-described methods for generating occupancy tags.

[0043] This invention provides a method, apparatus, electronic device, and storage medium for generating occupancy tags. In this method, the bounding boxes in the keyframe point cloud data distinguish the point cloud data corresponding to each dynamic object, ensuring that point cloud data reflecting the same dynamic object are within the same bounding box region. Furthermore, the keyframe point cloud data with the closest acquisition time to the non-keyframe point cloud data is determined as the target keyframe data corresponding to that non-keyframe point cloud data. This allows for the establishment of a correlation between non-keyframes and target keyframes with minimal changes in the dynamic object's position, reducing the difficulty of determining the second point cloud data corresponding to the first point cloud data and enabling the determined second point cloud data to more accurately reflect the dynamic object corresponding to the first point cloud data. By stitching together the second point cloud data, which accurately reflects the dynamic object, and the first point cloud data, the gaps between point clouds reflecting the same dynamic object can be reduced, resulting in a denser third point cloud data that accurately reflects the dynamic object, thus increasing the density of the point cloud data corresponding to the dynamic object. Based on this third point cloud data, highly accurate occupancy tags corresponding to multiple frames of point cloud data can be determined, improving the accuracy of the generated occupancy tags. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 This is a flowchart illustrating the occupation tag generation method provided in an embodiment of the present invention;

[0046] Figure 2 This is a schematic block diagram illustrating the generation of occupation tags provided in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of splicing static object point cloud data provided in an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram illustrating the determination of second point cloud data provided in an embodiment of the present invention;

[0049] Figure 5 This is a schematic diagram of the occupying tag provided in an embodiment of the present invention;

[0050] Figure 6 This is a schematic diagram of the structure of the occupation tag generation device provided in an embodiment of the present invention;

[0051] Figure 7 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. It should be noted that the serial numbers assigned to the objects described in this invention, such as "first," "second," etc., are only used to distinguish the described objects and have no sequential or technical meaning.

[0053] In practical applications, occupancy labels can be generated based on semantically meaningful point cloud data. Point cloud data itself does not possess semantics; semantic annotation is typically performed using sparse radar semantics.

[0054] Currently, mature annotation platforms primarily rely on annotation experience and tools to accurately select and annotate obstacles such as vehicles, pedestrians, and road signs, including using consecutive frames to jointly annotate multiple obstacle categories. Simultaneously, they perform semantic annotation of point clouds based on a 3D global view and a 2D camera view, labeling the object category of each point cloud element to ensure accuracy and reliability. However, the accuracy rate of current annotation data remains low.

[0055] Traditional annotation platforms and tools can only generate bounding boxes of the target objects to be identified in the current frame's point cloud data, along with relatively sparse point cloud semantics for that frame. Sparse point cloud data and semantics cannot meet the needs of occupancy grid training. Training with sparse point cloud data results in a model that outputs relatively sparse results, i.e., occupancy labels with low accuracy. Low-accuracy occupancy labels cannot adequately support the operation of existing decision-making modules.

[0056] To improve the density of point cloud data, multi-frame LiDAR point sequences can be directly converted into a unified coordinate system, and then the connected dense points are voxelized. This method ignores moving objects in the scene. When dynamic objects exist in the scene, the point cloud data of these dynamic objects in the aliased multi-frame point cloud data are chaotically spliced ​​together. The aliased point cloud cannot accurately represent each dynamic object in the scene, which severely affects the recognition effect of the occupancy grid network. It is difficult to accurately identify the occupancy status within the scene, resulting in low accuracy of the generated occupancy labels.

[0057] To address the aforementioned problems, this invention provides an occupancy tag generation method. This method targets keyframe and non-keyframe point cloud data from multi-frame point cloud data. It determines a second point cloud data point in the non-keyframe point cloud data that corresponds to a first point cloud data point in the target keyframe, ensuring that the determined second point cloud data represents the same dynamic object. Furthermore, by stitching together the first and second point cloud data, the gaps between point cloud data points of the same dynamic object are reduced, resulting in a third point cloud data point with higher density. Occupancy tags are then generated based on this highly accurate and dense third point cloud data point, improving the accuracy of the generated occupancy tags. The following describes the method in conjunction with... Figures 1 to 5 The method for generating occupancy tags provided in the embodiments of the present invention will be described.

[0058] Figure 1 This is a flowchart illustrating the occupancy tag generation method provided in this embodiment of the invention. The occupancy tag generation method provided in this embodiment is applicable to occupancy tag generation scenarios based on point cloud data. The executing entity of this method can be an electronic device such as a computer, server, server cluster, or a specially designed occupancy tag generation device, or it can be an occupancy tag generation device installed in such an electronic device. This occupancy tag generation device can be implemented through software, hardware, or a combination of both. Figure 1 As shown, the occupation tag generation method includes steps 110 to 140.

[0059] Step 110: For each non-key frame point cloud data in the multi-frame point cloud data, determine the target key frame point cloud data that is closest to the acquisition time of the non-key frame point cloud data from all the key frame point cloud data in the multi-frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object.

[0060] Specifically, multi-frame point cloud data can be at least two frames of point cloud data acquired based on LiDAR technology. For example, multi-frame point cloud data can be obtained from a database containing point cloud data, or it can be acquired by a data acquisition vehicle, or it can be obtained by other methods.

[0061] For example, for at least one frame of point cloud data in a multi-frame point cloud dataset, bounding box annotation can be used to annotate the dynamic objects in the point cloud data, with each annotated bounding box containing one dynamic object. Keyframe point cloud data can be understood as point cloud data frames annotated with bounding boxes. When multiple dynamic objects exist in the keyframe point cloud data, each dynamic object can be annotated within its corresponding bounding box. Dynamic objects can be understood as movable objects, such as people or vehicles in a scene. Bounding boxes (Bboxes) can be used for annotation, for example.

[0062] For example, in multi-frame point cloud data, non-keyframe point cloud data can be understood as point cloud data frames without bounding boxes. Multi-frame point cloud data may include one or more frames of non-keyframe point cloud data. Based on the acquisition time sequence of each frame of the multi-frame point cloud data, the keyframe point cloud data whose acquisition time is closest to that of the non-keyframe point cloud data is determined as the target keyframe point cloud data corresponding to that non-keyframe point cloud data. The acquisition time sequence is the temporal order in which the multi-frame point cloud data is acquired. It can be understood that the closer the acquisition times, the smaller the positional change of the dynamic object in the non-keyframe point cloud data and the target keyframe point cloud data. Based on this, a correlation and correspondence can be established between keyframe point cloud data and non-keyframe point cloud data, and this is beneficial for determining the point cloud data corresponding to the same dynamic object in non-keyframes.

[0063] For example, the acquisition vehicle collects point cloud data at a frame rate of 10 fps (frames per second), meaning it collects 10 frames of point cloud data every second. The keyframe annotation frame rate is 2 fps. This means that of the 10 frames of point cloud data collected per second, 2 frames can be dynamically annotated using bounding boxes (BBoxes), and these 2 frames are the keyframe point cloud data. The other 8 frames are non-keyframe point cloud data. Based on the acquisition timing, the target keyframe point cloud data corresponding to each of the 8 non-keyframe point cloud data is determined. For example, in a 1-second acquisition sequence, the acquisition times of the two keyframe point cloud data are 0.3 seconds and 0.8 seconds, respectively; the non-keyframe point cloud data acquired at 0.1 seconds, 0.2 seconds, 0.4 seconds, and 0.5 seconds are the closest to the acquisition time of the keyframe point cloud data acquired at 0.3 seconds, so the target keyframe point cloud data corresponding to these four frames of point cloud data is the keyframe point cloud data acquired at 0.3 seconds; similarly, the target keyframe point cloud data corresponding to the non-keyframe point cloud data acquired at 0.6 seconds, 0.7 seconds, 0.9 seconds, and 1 second is the keyframe point cloud data acquired at 0.8 seconds.

[0064] Step 120: Based on the first point cloud data in the first target bounding box of the target keyframe point cloud data, determine the second point cloud data in the non-keyframe point cloud data that corresponds to the first point cloud data.

[0065] Specifically, the first target bounding box can be a bounding box in the target keyframe point cloud data. The point cloud data in the first target bounding box includes first point cloud data, which can be point cloud data of the same dynamic object in the first target bounding box.

[0066] Based on the first point cloud data in the target keyframe point cloud data, the second point cloud data in each non-keyframe point cloud data corresponding to the target keyframe point cloud data can be determined. The determined second point cloud data corresponds to the first point cloud data, and the second point cloud data can also be used to represent the same dynamic object represented by the first point cloud data within the first target bounding box. It can be understood that the corresponding first and second point cloud data are point cloud data in different point cloud data frames used to represent the same dynamic object.

[0067] When determining the second point cloud data corresponding to the first point cloud data, it can be based on the temporal and spatial relationships between the first and second point cloud data, as well as the dynamic data information of the dynamic object. For example, based on the acquisition timing of multiple frames of point cloud data, the acquisition time difference between each frame of point cloud data can be determined; based on the acquisition time difference and the speed and direction of movement of the same dynamic object in space, the position of the same dynamic object in the non-key frame point cloud data can be determined; based on the determined position, the second point cloud data corresponding to the first point cloud data in the non-key frame point cloud data can be determined.

[0068] Step 130: Concatenate the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object.

[0069] Specifically, the first point cloud data and the corresponding second point cloud data both represent the same dynamic object. The first point cloud data and the second point cloud data can be spliced ​​together to improve the density of the point cloud data of the dynamic object.

[0070] For example, stitching together the first and second point cloud data can be understood as superimposing and combining point cloud data. By superimposing and combining corresponding first and second point cloud data from different frames, the density of point cloud data reflecting the same dynamic object can be increased. The resulting point cloud data is the third point cloud data corresponding to that dynamic object. The third point cloud data has a higher density than the first point cloud data, meaning it has more points and a larger data volume. By determining the second point cloud data, the point cloud data in non-key frames can be distinguished, establishing a corresponding relationship between the point cloud data of the same dynamic object in each frame. This avoids the point cloud data chaos caused by directly mixing all point cloud data in each frame, which would prevent an accurate reflection of the occupancy of each dynamic object in the scene space. Here, scene space can be understood as the scene space corresponding to the collected multi-frame point cloud data, such as the scene space of a street during collection.

[0071] Optionally, when annotating bounding boxes in keyframe point cloud data, regular-shaped bounding boxes, such as cube or cubic bounding boxes, can be used; or irregular-shaped bounding boxes, such as bounding boxes of arbitrary shapes that annotate the point cloud data corresponding to the same dynamic object within a minimum volume. It can be understood that if all point data included within a bounding box are point data of the same dynamic object, the point cloud data composed of all point data within that bounding box can more accurately reflect the position information or occupancy of the dynamic object in space; or it can be understood that the less noise data within a bounding box, the more accurately the first point cloud data within that bounding box will reflect the information of the same dynamic object.

[0072] Step 140: Based on the third point cloud data, determine the occupancy labels corresponding to the multi-frame point cloud data.

[0073] Specifically, in multi-frame point cloud data, the semantics of each point cloud in the keyframe point cloud data can be annotated. After obtaining the third point cloud data corresponding to the dynamic object, the semantics of the annotated first point cloud data can be synchronously assigned to its corresponding third point cloud data, so that its corresponding third point cloud data also has semantics. Based on the semantically meaningful third point cloud data, an occupancy label can be generated.

[0074] For example, sparse LiDAR point cloud semantic annotation can be performed on each point cloud in the keyframe point cloud data, assigning semantic meaning to each point cloud in the keyframe point cloud data. This includes the first point cloud data, which will then have the semantic annotation. Semantic annotation can be understood as annotating the object category to which the point cloud data belongs, such as a person, animal, or vehicle. After concatenating the first and second point cloud data, the semantic meaning of the first point cloud data can be assigned to the concatenated third point cloud data, giving the third point cloud data semantic meaning. This allows the object category to be determined for the dynamic object represented by the third point cloud data. Based on the semantically meaningful third point cloud data, corresponding occupancy labels can be generated in the occupancy grid. Since the third point cloud data has relatively high density and accuracy, it can more accurately reflect the occupancy of the dynamic object it represents in the scene space. Therefore, the generated occupancy labels can more accurately reflect the ground truth of the occupancy grid.

[0075] The occupancy tag generation method provided in this invention involves using bounding boxes in keyframe point cloud data to distinguish point cloud data corresponding to different dynamic objects, ensuring that point cloud data reflecting the same dynamic object are within the same bounding box region. Furthermore, the keyframe point cloud data with the closest acquisition time to the non-keyframe point cloud data is identified as the target keyframe data corresponding to that non-keyframe point cloud data. This allows for the establishment of a correlation between non-keyframes and target keyframes with minimal changes in the dynamic object's position, reducing the difficulty of determining the second point cloud data corresponding to the first point cloud data and enabling the determined second point cloud data to more accurately reflect the dynamic object corresponding to the first point cloud data. By stitching together the second point cloud data, which accurately reflects the dynamic object, and the first point cloud data, the gaps between point clouds reflecting the same dynamic object can be reduced, resulting in a denser and more accurate third point cloud data that accurately reflects the dynamic object, thus increasing the density of the point cloud data corresponding to the dynamic object. Based on this third point cloud data, highly accurate occupancy tags corresponding to multiple frames of point cloud data can be determined, improving the accuracy of the generated occupancy tags.

[0076] In practical applications, scene space typically includes not only dynamic objects but also static objects. Static objects can be understood as immobile objects, such as roads, road signs, buildings, or grass. To obtain more accurate occupancy labels, the fourth point cloud data of the same static object in the keyframe point cloud data and the fifth point cloud data in the non-keyframe point cloud data can be stitched together to obtain a sixth point cloud data with higher density representing the static object. This sixth point cloud data can more accurately reflect the occupancy of the static object in the scene space. Furthermore, more accurate occupancy labels can be generated based on the third and sixth point cloud data.

[0077] In one embodiment, the occupancy labels corresponding to multiple frames of point cloud data are determined based on the third point cloud data. Specifically, this can be achieved as follows: For each key frame point cloud data in the multiple frame point cloud data, based on the point cloud data in each bounding box of the key frame point cloud data, the fourth point cloud data corresponding to the static object in the key frame point cloud data is determined; based on the second point cloud data in the non-key frame point cloud data, the fifth point cloud data corresponding to the static object in the non-key frame point cloud data is determined; the fourth point cloud data and the fifth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object; based on the third point cloud data and the sixth point cloud data, the occupancy labels corresponding to the multiple frames of point cloud data are determined.

[0078] Specifically, the point cloud data in each bounding box of the keyframe point cloud data includes the first point cloud data representing dynamic objects, and the point cloud data outside each bounding box of the keyframe point cloud data can be the fourth point cloud data representing static objects. The fourth point cloud data can be understood as the static object point cloud or the static scene point cloud in the keyframe point cloud data, and the object corresponding to the fourth point cloud data can be understood as a static scene object or a background object, etc.

[0079] For example, in keyframe point cloud data, bounding boxes can be used to separate the point clouds corresponding to dynamic objects from those corresponding to static objects, i.e., separating the static scene point cloud from the dynamic object point cloud. For a frame of keyframe point cloud data, a bounding box is used to label all the point clouds corresponding to dynamic objects in the keyframe point cloud data, for example, using a rectangular bounding box. For each bounding box, the coordinate ranges of each point cloud in the x-axis, y-axis, and z-axis directions are obtained from the coordinates of the eight corner points of the bounding box, represented as [x_min, x_max], [y_min, y_max], and [z_min, z_max], respectively, where x_min represents the minimum value of the coordinate in the x-axis direction, x_max represents the maximum value of the coordinate in the x-axis direction, and so on. The coordinates of all points in the keyframe point cloud data are traversed, and it is determined whether the coordinates of each point can simultaneously satisfy the condition that the x, y, and z coordinates all fall within the coordinate range range of the corresponding direction. If a point is within the coordinate range of the corresponding direction, it belongs to the dynamic object within that bounding box; if it is not within the coordinate range of the corresponding direction, it does not belong to the dynamic object within that bounding box. Thus, after calculating and judging all bounding boxes in the keyframe point cloud data, the points that do not belong to any dynamic object within any bounding box are the points of the static scene point cloud, i.e., the points corresponding to static objects. The point cloud composed of the points corresponding to static objects is the fourth point cloud data corresponding to the static objects in the keyframe point cloud data. Based on this, the goal of determining the fourth point cloud data based on the point cloud data in each bounding box of the keyframe point cloud data is achieved, that is, the separation of the static scene point cloud and the dynamic object point cloud is realized.

[0080] For example, after determining the second point cloud data corresponding to the first point cloud data, other point cloud data besides the second point cloud data can be determined as the fifth point cloud data corresponding to the static object in the non-keyframe point cloud data. The fifth point cloud data can be understood as the static point cloud in the non-keyframe point cloud data. The fourth point cloud data and its corresponding fifth point cloud data both represent the same static object.

[0081] By concatenating the fourth and fifth point cloud data, the fourth point cloud data can be superimposed and combined with the fifth point cloud data corresponding to the fourth point cloud data to obtain the sixth point cloud data representing the same static object. The density of the sixth point cloud data is higher than that of the fourth point cloud data. The sixth point cloud data can represent the occupancy of the static object in the scene space in a more detailed and accurate manner.

[0082] Figure 2 This is a schematic block diagram of generating an occupation tag provided in an embodiment of the present invention, such as... Figure 2 As shown, for the target keyframe point cloud data, the point cloud data corresponding to dynamic objects and static objects can be separated using the first target bounding box to obtain the dynamic object point cloud and static object point cloud of the target keyframe point cloud data, i.e., the first point cloud data and the fourth point cloud data. After determining the second point cloud data, the point cloud data corresponding to dynamic objects and static objects in the non-keyframe point cloud data can be separated based on the second point cloud data, i.e., the second point cloud data and the fifth point cloud data. The point cloud data corresponding to dynamic objects separated from the non-keyframe point cloud data is concatenated with the point cloud data corresponding to dynamic objects separated from the target keyframe point cloud data to obtain the concatenated dynamic object point cloud data, i.e., the third point cloud data; the point cloud data corresponding to static objects separated from the non-keyframe point cloud data is concatenated with the point cloud data corresponding to static objects separated from the target keyframe point cloud data to obtain the concatenated static object point cloud data, i.e., the sixth point cloud data. After processing the stitched dynamic object point cloud data and the stitched static object point cloud data through filtering, denoising, surface reconstruction, and semantic assignment, the stitched dynamic object point cloud data and the stitched static object point cloud data are stitched together to obtain complete point cloud data corresponding to multiple frames of point cloud data. This complete point cloud data has a high density, and each point cloud data in it can more accurately reflect the occupancy of each object in the scene space, thus making the occupancy label corresponding to the multi-frame point cloud data generated based on this complete point cloud data more accurate.

[0083] In this embodiment, before generating occupancy tags, the point cloud data corresponding to the spliced ​​dynamic object and the point cloud data corresponding to the spliced ​​static object are spliced ​​together to obtain complete point cloud data of the scene space corresponding to multiple frames of point cloud data. That is, the third point cloud data and the sixth point cloud data are superimposed and combined to obtain complete point cloud data representing all objects in the scene space. Based on this complete point cloud data, the occupancy status of all objects in the scene space can be fully reflected, thereby generating more complete and comprehensive occupancy tags.

[0084] For example, even using sparse LiDAR point cloud semantics to semantically annotate all point cloud data in multi-frame point cloud data requires significant cost and time. Typically, semantic annotation is only performed on point clouds in keyframe point cloud data, while the semantics of point clouds in non-keyframe point cloud data are not annotated. Therefore, the fifth point cloud data determined based on non-keyframe point cloud data is semantically empty. The sixth point cloud data, obtained by concatenating the fourth and fifth point cloud data, will include semantically empty point cloud data; this semantically empty point cloud data is the target point cloud data. The lack of semantic meaning in the target point cloud data in the sixth point cloud data will reduce the accuracy of the generated occupancy labels. To improve the accuracy of the generated occupancy labels, appropriate semantics can be assigned to the target point cloud data to complete the semantic information of the point cloud.

[0085] In one embodiment, when determining the occupancy labels corresponding to multiple frames of point cloud data based on the third point cloud data and the sixth point cloud data, for each target point cloud data without semantics in the sixth point cloud data, the point cloud data with semantics that is closest to the target point cloud data in the sixth point cloud data can be determined; the semantics of the point cloud data with semantics is determined as the semantics of the target point cloud data; and the occupancy labels corresponding to multiple frames of point cloud data are determined based on the semantics of the third point cloud data, the semantics of the point cloud data with semantics, and the semantics of the target point cloud data.

[0086] Specifically, the sixth point cloud data is obtained by stitching together the fourth and fifth point cloud data. The fourth point cloud data is obtained based on keyframe point cloud data, and each point cloud data in the keyframe point cloud data is semantically labeled. Therefore, the fourth point cloud data is semantically meaningful. The fifth point cloud data is obtained based on non-keyframe point cloud data; therefore, the fifth point cloud data is not semantically labeled.

[0087] In the sixth point cloud data obtained by concatenating the fourth and fifth point cloud data, the semantics of the point cloud data with the closest semantic proximity to the target point cloud data can be determined as the semantics of the target point cloud data. This is achieved using a nearest-neighbor strategy to determine the semantics of the target point cloud data. Both the fourth and fifth point cloud data represent static objects. Since the positions of static objects are relatively stationary, the closest point cloud data in the concatenated sixth point cloud data are highly likely to represent the same static object. Based on this, the semantics of the point cloud data with the closest semantic proximity to the target point cloud data can be determined as the semantics of the target point cloud data. This semantics has a high probability of being correct, thus allowing the target point cloud data without semantics to be assigned a suitable semantic. It is important to note that when determining the semantics of the target point cloud data, a strategy of separating static and dynamic objects is used. That is, when determining the semantics of static objects without semantics, only the semantics of static objects with semantics are used. This avoids assigning the semantics of static objects such as ground and trees to dynamic objects such as cars, ensuring the rationality and accuracy of the determined semantics, and thus improving the accuracy of occupancy labeling.

[0088] For example, after determining the semantics of the target point cloud data, each point cloud data in the sixth point cloud data has semantics. Occupancy tags corresponding to multiple frames of point cloud data can be determined based on the semantics of the third point cloud data, the semantics of the semantically meaningful point cloud data, and the semantics of the target point cloud data. That is, when determining the occupancy tags corresponding to multiple frames of point cloud data based on the third and sixth point cloud data, each dynamic object point cloud and each static object point cloud has semantics. Based on this, occupancy tags can be generated according to the semantics of each point cloud data.

[0089] In this embodiment, the nearest neighbor strategy is used to determine the semantics of target point cloud data that lacks semantics, thus completing the semantics of each point cloud data. When generating occupancy labels based on the third and sixth point cloud data, the accuracy of the generated occupancy labels can be improved.

[0090] It should be noted that when keyframe point cloud data is semantically annotated but non-keyframe point cloud data is not, all second point cloud data obtained based on the non-keyframe point cloud data are semantically empty. When concatenating the first and second point cloud data to obtain the third point cloud data, the semantics of the first point cloud data can be determined as the semantics of the corresponding third point cloud data. This can be understood as assigning the point cloud semantics of the first point cloud data representing the same dynamic object to the third point cloud data corresponding to the same dynamic object after point cloud concatenation.

[0091] For example, point cloud data is data obtained from radar equipment, which has a radar coordinate system, and point cloud data is used to represent information within that radar coordinate system. However, in real-world scenarios, static objects are stationary relative to the scene space. Therefore, stitching together point clouds of static objects in the real-world coordinate system can yield more accurate results.

[0092] In one embodiment, the fourth and fifth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object. Specifically, this can be achieved as follows: the fourth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the seventh point cloud data in the world coordinate system; the fifth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the eighth point cloud data in the world coordinate system; and the seventh and eighth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object.

[0093] Specifically, a transformation matrix can be used to transform the fourth point cloud data from the radar coordinate system to the world coordinate system, resulting in the seventh point cloud data in the world coordinate system. Similarly, a transformation matrix can be used to transform the fifth point cloud data from the radar coordinate system to the world coordinate system, resulting in the eighth point cloud data in the world coordinate system. The transformation matrix can be obtained based on the coordinate extrinsic parameters of the calibrated radar equipment, such as the installation height or angle of the lidar. The higher the calibration quality, i.e., the more accurate the calibrated extrinsic parameters, the higher the accuracy of the seventh and eighth point cloud data obtained based on the transformation matrix, resulting in more accurate and denser sixth point cloud data. The radar coordinate system can be understood as the LiDAR coordinate system.

[0094] Figure 3 This is a schematic diagram of stitching static object point cloud data provided in an embodiment of the present invention, such as... Figure 3 As shown, based on the static object point cloud data of each single frame, the point cloud data in the radar coordinate system is converted into point cloud data in the world coordinate system through a transformation matrix, and the obtained point cloud data are stitched together to obtain point cloud data representing static objects with higher density and higher accuracy.

[0095] In this embodiment, when the fourth and fifth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object, the coordinate system of the fourth and fifth point cloud data is transformed by a transformation matrix. This allows the point cloud data to be aligned and concatenated in the world coordinate system, so that the concatenated sixth point cloud data can more accurately represent the occupancy of the static object in the scene space, thereby improving the accuracy of the generated occupancy label.

[0096] For example, when determining the second point cloud data corresponding to the first point cloud data in the non-keyframe point cloud data, the second target bounding box in the non-keyframe point cloud data can be determined by the position of the first target bounding box. The point cloud data within the second target bounding box can include the second point cloud data. That is, without annotation, the second point cloud data can be selected by determining the bounding box. Based on this, the second point cloud data corresponding to the first point cloud data in the non-keyframe point cloud data can be determined quickly and accurately.

[0097] In one embodiment, based on the first point cloud data in the first target bounding box of the target keyframe point cloud data, the second point cloud data corresponding to the first point cloud data in the non-keyframe point cloud data is determined. Specifically, this can be done as follows: based on the position of the first target bounding box, a second target bounding box corresponding to the position of the first target bounding box is determined in the non-keyframe point cloud data, and the area of ​​the second target bounding box is greater than the area of ​​the first target bounding box; for each first point cloud data, the distance between the first point cloud data and each point cloud data in the second target bounding box is determined, and the point cloud data with a minimum distance less than a preset distance is determined as the second point cloud data.

[0098] Specifically, the second target bounding box can be a bounding box determined based on the first target bounding box, used to select the second point cloud data in the non-key frame point cloud data corresponding to the target key frame point cloud data.

[0099] For example, the coordinates of the first key point in the target keyframe point cloud data are determined, and the coordinates of the second key point corresponding to the key point of the second target bounding box are determined in the non-keyframe point cloud data based on the first key point coordinates. The key point of the first target bounding box can be the center point or a vertex of the first target bounding box. When determining the second key point coordinates of the second target bounding box in the non-keyframe point cloud data based on the first key point coordinates, for example, the position coordinates of the key point of the first target bounding box in the radar coordinate system can be used as the second key point coordinates corresponding to the key point of the second target bounding box.

[0100] After determining the coordinates of the second keypoint, a second target bounding box is defined based on these coordinates, having the same shape but a larger area as the first target bounding box. This larger bounding box can be understood as a bounding box obtained by enlarging the first target bounding box according to a preset ratio. Since the area of ​​the second target bounding box is larger than that of the first target bounding box, point cloud data representing the same dynamic object can be included within the area of ​​the second target bounding box when the object's position moves. Because the area of ​​the second target bounding box is large, some point cloud data within it may not be the second point cloud data; therefore, it is necessary to further determine the second point cloud data corresponding to the first point cloud data within the second target bounding box.

[0101] When determining the second point cloud data within the second target bounding box, the nearest neighbor filtering method can be used. Figure 4 This is a schematic diagram of determining the second point cloud data provided in an embodiment of the present invention, such as... Figure 4 As shown, for each non-keyframe point cloud data corresponding to the target keyframe point cloud data, the second target bounding box corresponding to each non-keyframe point cloud data is determined. This can also be understood as translating each non-keyframe point cloud data along the orientation direction of the first target bounding box. During translation, the second point cloud data can be determined by determining the translation matching cost. Determining the translation matching cost can be, for example, by calculating the distance between each point in the first point cloud data and the nearest point in the second target bounding box (e.g., Euclidean distance); iterating through each point in the first point cloud data to obtain the corresponding distance value; and summing all the distance values ​​to obtain the translation matching cost. When the calculated matching cost is less than a threshold, the points in the second target bounding box corresponding to this matching cost are the points in the second point cloud data. The matching cost at this point is the minimum distance, and the threshold is the preset distance. Based on this, the second point cloud data corresponding to the first point cloud data can be determined.

[0102] For example, such as Figure 4As shown, the second point cloud data in the non-keyframe point cloud data can be sequentially stitched together with the first point cloud data. During this sequential stitching, the second point cloud data in one frame of non-keyframe point cloud data is determined based on the first point cloud data in the target keyframe point cloud data. After determination, the first point cloud data is stitched together with the second point cloud data to obtain the third point cloud data. Based on this third point cloud data, the second point cloud data in another frame of non-keyframe point cloud data is determined as the first point cloud data in the target keyframe. These two are then stitched together to obtain the further stitched third point cloud data. This process is repeated for all the second point cloud data in the non-keyframe point cloud data to be stitched together with the first point cloud data in the target keyframe point cloud data, resulting in a third point cloud data with higher density. This method allows for sequential stitching to obtain highly accurate third point cloud data.

[0103] In this embodiment, based on the position of the first target bounding box, a second target bounding box corresponding to the position of the first target bounding box can be determined in the non-keyframe point cloud data. The point cloud of the same dynamic object in the non-keyframe point cloud data is then selected using the second target bounding box with a larger area. Furthermore, by determining the distance between the first point cloud data and each point cloud data in the second target bounding box, point cloud data with a minimum distance less than a preset distance can be identified as the second point cloud data. Based on this, a dynamic alignment and optimization strategy is adopted to adjust the alignment between the keyframe point cloud data and the point cloud data of dynamic objects in the non-keyframe point cloud data, improving the accuracy of the dynamic object point cloud data. This enables the rapid and accurate determination of the second point cloud data corresponding to the first point cloud data, thereby improving the accuracy of the determined second point cloud data and consequently improving the accuracy of the generated occupancy labels.

[0104] In practical applications, due to the characteristics of LiDAR sampling, some noise is inevitably introduced, affecting the accuracy of the generated occupancy labels. Noise affecting accuracy may exist in both the stitched third and sixth point cloud data. Therefore, denoising processing of the point cloud data is necessary to improve its accuracy, thereby improving the accuracy of the generated occupancy labels.

[0105] In one embodiment, the method further includes: acquiring an image corresponding to keyframe point cloud data; determining a third target bounding box corresponding to a target dynamic object in the image; determining a first semantic segmentation mask for the target dynamic object in the third target bounding box based on the semantics of a ninth point cloud data corresponding to the target dynamic object in the keyframe point cloud data; projecting the third point cloud data onto the first semantic segmentation mask; and deleting the mismatched point cloud data if there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask.

[0106] Specifically, when acquiring multi-frame point cloud data, the corresponding images can be acquired simultaneously. For example, by using multiple sensors such as cameras and LiDAR devices on a data acquisition vehicle, multiple frames of point cloud data and their corresponding images can be acquired in a hard-synchronized manner. Hard synchronization can be understood as the synchronization of the timestamps of the data acquired by the camera and LiDAR, with no deviation between the data. For keyframe point cloud data, bounding box annotation and sparse radar point cloud semantic annotation are performed.

[0107] For example, determining the third bounding box corresponding to a target dynamic object in an image can be understood as a bounding box used to label the target dynamic object in the image. For instance, image processing algorithms or models can be used to segment the image, dividing and labeling the image regions corresponding to each dynamic object, thus determining the third bounding box corresponding to the target dynamic object in the image. For example, a Segment Anything Model (SAM) can be used to segment the image corresponding to the keyframe point cloud data, labeling the third bounding box of each target dynamic object.

[0108] The first semantic segmentation mask of the target dynamic object within the third target bounding box is determined based on the semantics of the ninth point cloud data corresponding to the target dynamic object in the keyframe point cloud data. Here, the ninth point cloud data can be understood as the point cloud data corresponding to the dynamic object in the keyframe point cloud data.

[0109] In each keyframe point cloud data, the point cloud semantics of each dynamic object can be determined based on sparse LiDAR point cloud semantic annotation. The LiDAR point cloud semantics of the target dynamic object in the keyframe point cloud data are obtained, and this semantics is determined as the semantics of the target dynamic object in the corresponding image of the keyframe point cloud data. Once the target dynamic object in the third target bounding box has semantics, the first semantic segmentation mask of the target dynamic object is obtained. After determining the first semantic segmentation mask of each target dynamic object, the image can be understood as a semantic map or a driving semantic map, etc. Based on the semantic map, joint image point cloud filtering can be performed on the stitched third point cloud data to eliminate noise in the third point cloud data. The semantic map can be determined using formula (1).

[0110] Mask seg =SAM(imgs)*Semantics (1)

[0111] Among them, Mask seg SAM(imgs) represents the semantic graph; SAM(imgs) represents the image obtained after image segmentation with a third target bounding box; Semantics represents the point cloud semantics annotated in the target keyframe point cloud data.

[0112] For example, the third point cloud data is projected onto the first semantic segmentation mask, and if there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask, the mismatched point cloud data is deleted.

[0113] This can be understood as follows: the third point cloud data representing the dynamic target object is transformed from radar coordinates to image coordinates using a transformation matrix, and then projected onto the corresponding first semantic segmentation mask in the semantic map. The coordinates of the third point cloud data in the semantic map can be determined using formula (2).

[0114] (u, v) = T Egoto_img (P obj (2)

[0115] Among them, T Ego_to_img P represents the transformation matrix during projection; obj (u, v) represents the coordinates of the third point cloud data in the radar coordinate system; (u, v) represents the coordinates of the third point cloud data in the semantic graph after being projected onto the first semantic segmentation mask.

[0116] The semantics of the point cloud data in the third point cloud data are determined based on the semantics of the first semantic segmentation mask. As shown in the following formula (3), if the semantics of the point cloud data in the third point cloud data are the same as the semantics of the first semantic segmentation mask, it is a match, and the point is determined to be a non-noise point, so the point cloud data is retained; if the semantics of the point cloud data in the third point cloud data are not the same as the semantics of the first semantic segmentation mask, it is a mismatch, and the point cloud data is determined to be a noise point, so the point cloud data is deleted.

[0117] Mask point =Mask seg (u, v) == Object (3)

[0118] Among them, Mask point Indicates the first semantic segmentation mask; Mask seg (u, v) represents the semantics of the coordinates of (u, v) in the semantic graph; Object represents the semantics of the target dynamic object corresponding to the third point cloud data, that is, the category to which the target dynamic object belongs.

[0119] In this embodiment, the image corresponding to the keyframe point cloud data is acquired; the third target bounding box corresponding to the target dynamic object in the image is determined; based on the semantics of the ninth point cloud data corresponding to the target dynamic object in the keyframe point cloud data, a first semantic segmentation mask for the target dynamic object in the third target bounding box is determined; the third point cloud data is projected onto the first semantic segmentation mask; if there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask, the mismatched point cloud data is deleted. Based on the image acquired synchronously with the target keyframe point cloud data, joint image-point cloud filtering can be performed on the stitched third point cloud data to eliminate noise in the third point cloud data, thereby improving the accuracy of the third point cloud data. The joint point cloud filtering using mutual verification of point cloud images can enhance the consistency between the generated occupancy label and the camera image.

[0120] Alternatively, joint image and point cloud filtering can be performed on the stitched point cloud data corresponding to the static object to eliminate noise in the sixth point cloud data obtained after stitching, thereby improving the accuracy of the point cloud data corresponding to the static object.

[0121] In one embodiment, the method further includes: acquiring an image corresponding to the keyframe point cloud data; projecting the fourth point cloud data onto the image, and determining a second semantic segmentation mask corresponding to a static object in the image based on the semantics of the fourth point cloud data in the keyframe point cloud data; projecting the sixth point cloud data onto the second semantic segmentation mask, and deleting the mismatched point cloud data if there are point cloud data in the sixth point cloud data that do not semantically match the second semantic segmentation mask.

[0122] Specifically, when acquiring multi-frame point cloud data, the corresponding images can be acquired simultaneously. Using a transformation matrix similar to that used in projection, the fourth point cloud data is projected onto the image. Based on the semantics corresponding to the fourth point cloud data in the keyframe point cloud data, this semantics is assigned to the points at the corresponding coordinates in the image, thus determining the second semantic segmentation mask corresponding to the static object in the image. For each static object, any number of points can be randomly selected as cues, such as three points, to generate the corresponding second semantic segmentation mask.

[0123] Similarly, the sixth point cloud data can be projected onto the second semantic segmentation mask using a transformation matrix. Based on the determination method described in the above embodiments, noise is determined for each point cloud data in the sixth point cloud data. If there are point cloud data in the sixth point cloud data that do not semantically match the second semantic segmentation mask, the mismatched point cloud data can be deleted, thereby removing noise from the sixth point cloud data and obtaining sixth point cloud data with higher accuracy.

[0124] In this embodiment, the image corresponding to the keyframe point cloud data is acquired; the fourth point cloud data is projected onto the image, and based on the semantics of the fourth point cloud data in the keyframe point cloud data, a second semantic segmentation mask corresponding to the static object in the image is determined; the sixth point cloud data is projected onto the second semantic segmentation mask, and if there are point cloud data in the sixth point cloud data that do not semantically match the second semantic segmentation mask, the mismatched point cloud data is deleted. Based on this, the stitched sixth point cloud data can be filtered and denoised to eliminate noise and improve the accuracy of the sixth point cloud data. Assigning the nearest neighbor semantics to the static scene and dynamic objects respectively ensures the correctness of the semantic edges of the dynamic objects and can improve the accuracy of generating occupancy labels.

[0125] For example, when point cloud data contains missing or sparse regions, Poisson surface reconstruction can fill these missing regions through interpolation and smoothing to generate a complete surface model. Discrete point cloud data often contains noise or outliers, which can be removed and their impact reduced by smoothing the surface. Figure 2 As shown, after performing joint image point cloud filtering on the third and sixth point cloud data, Poisson surface reconstruction can be applied. This reconstructs continuous surface models from the discrete, unordered point cloud data, filling in the missing point cloud data. The point cloud data is reconstructed into a triangular mesh using Poisson surface reconstruction. The input is point cloud data with normal vectors, and the output is a triangular mesh. The obtained mesh uses uniformly distributed vertices to fill the gaps in the point cloud data, and the triangular mesh can be further converted into dense voxels. In this way, the generated occupancy labels utilize point cloud data from all frames in the sequence and fill in the gaps between point cloud data, resulting in more accurate occupancy labels.

[0126] Figure 5 This is a schematic diagram of the occupying tag provided in an embodiment of the present invention, such as... Figure 5 As shown, based on the dense third and sixth point cloud data, occupancy labels can be stitched together and generated. The generated occupancy labels are based on point cloud data of dynamic objects after alignment and stitching, employing an optimized strategy. Joint filtering of the point cloud images is used to separate and denoise the static and dynamic object point clouds. Compared to existing solutions, this significantly improves the accuracy of the generated occupancy labels, ensuring consistency between the image and the ground truth.

[0127] The occupation tag generation apparatus provided in the embodiments of the present invention is described below. The occupation tag generation apparatus described below can be referred to in correspondence with the occupation tag generation method described above. Figure 6This is a schematic diagram of the occupancy tag generation device provided in an embodiment of the present invention, with reference to... Figure 6 As shown, the label generation device 600 includes:

[0128] The first determining module 610 is used to determine, from all the key frame point cloud data in the multi-frame point cloud data, the target key frame point cloud data that is closest to the acquisition time of the non-key frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object.

[0129] The second determining module 620 is used to determine the second point cloud data in the non-key frame point cloud data that corresponds to the first point cloud data based on the first point cloud data in the first target bounding box in the target key frame point cloud data.

[0130] The splicing module 630 is used to splice the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object;

[0131] The third determination module 640 is used to determine the occupancy labels corresponding to multiple frames of point cloud data based on the third point cloud data.

[0132] In one example embodiment, the third determining module 640 is specifically used to: for each key frame point cloud data in the multi-frame point cloud data, determine the fourth point cloud data corresponding to the static object in the key frame point cloud data based on the point cloud data in each bounding box in the key frame point cloud data; determine the fifth point cloud data corresponding to the static object in the non-key frame point cloud data based on the second point cloud data in the non-key frame point cloud data; concatenate the fourth point cloud data and the fifth point cloud data to obtain the sixth point cloud data corresponding to the static object; and determine the occupancy label corresponding to the multi-frame point cloud data based on the third point cloud data and the sixth point cloud data.

[0133] In one example embodiment, the third determining module 640 is specifically used to: for each target point cloud data in the sixth point cloud data that has no semantics, determine the point cloud data with semantics that is closest to the target point cloud data in the sixth point cloud data; determine the semantics of the point cloud data with semantics as the semantics of the target point cloud data; and determine the occupancy tags corresponding to the multi-frame point cloud data based on the semantics of the third point cloud data, the semantics of the point cloud data with semantics, and the semantics of the target point cloud data.

[0134] In one example embodiment, the third determining module 640 is specifically used to: convert the fourth point cloud data from the radar coordinate system to the world coordinate system to obtain the seventh point cloud data in the world coordinate system; convert the fifth point cloud data from the radar coordinate system to the world coordinate system to obtain the eighth point cloud data in the world coordinate system; and concatenate the seventh point cloud data and the eighth point cloud data to obtain the sixth point cloud data corresponding to the static object.

[0135] In one example embodiment, the second determining module 620 is specifically used to: determine a second target bounding box corresponding to the position of the first target bounding box in the non-keyframe point cloud data based on the position of the first target bounding box, wherein the area of ​​the second target bounding box is greater than the area of ​​the first target bounding box; and for each first point cloud data, determine the distance between the first point cloud data and each point cloud data in the second target bounding box, and determine the point cloud data whose minimum distance is less than a preset distance as the second point cloud data.

[0136] In one example embodiment, the occupancy label generation device 600 further includes a first processing module; the first processing module is used to acquire an image corresponding to the keyframe point cloud data; the first processing module is also used to determine a third target bounding box corresponding to a target dynamic object in the image, and to determine a first semantic segmentation mask of the target dynamic object in the third target bounding box based on the semantics of the ninth point cloud data corresponding to the target dynamic object in the keyframe point cloud data; the first processing module is also used to project the third point cloud data onto the first semantic segmentation mask, and to delete the mismatched point cloud data if there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask.

[0137] In one example embodiment, the occupancy label generation device 600 further includes a second processing module; the second processing module is used to acquire an image corresponding to the keyframe point cloud data; the second processing module is also used to project the fourth point cloud data onto the image, and determine a second semantic segmentation mask corresponding to a static object in the image based on the semantics of the fourth point cloud data in the keyframe point cloud data; the second processing module is also used to project the sixth point cloud data onto the second semantic segmentation mask, and delete the mismatched point cloud data if there are point cloud data in the sixth point cloud data that do not match the semantics of the second semantic segmentation mask.

[0138] The apparatus of this embodiment can be used to execute the method of any embodiment in the side embodiment of the occupation tag generation method. Its specific implementation process and technical effects are similar to those in the side embodiment of the occupation tag generation method. For details, please refer to the detailed description in the side embodiment of the occupation tag generation method, which will not be repeated here.

[0139] Figure 7 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 7As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute an occupancy tag generation method, which includes: for each non-key frame point cloud data in multi-frame point cloud data, determining the target key frame point cloud data closest to the acquisition time of the non-key frame point cloud data from all key frame point cloud data in the multi-frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object; based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determining the second point cloud data in the non-key frame point cloud data corresponding to the first point cloud data; concatenating the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object; and determining the occupancy tag corresponding to the multi-frame point cloud data based on the third point cloud data.

[0140] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] On the other hand, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the occupancy tag generation method provided by the above methods. The method includes: for each non-key frame point cloud data in multi-frame point cloud data, determining, from all key frame point cloud data in the multi-frame point cloud data, a target key frame point cloud data whose acquisition time is closest to that of the non-key frame point cloud data; the key frame point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object; based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determining a second point cloud data in the non-key frame point cloud data corresponding to the first point cloud data; concatenating the first point cloud data and the second point cloud data to obtain a third point cloud data corresponding to the dynamic object; and determining the occupancy tag corresponding to the multi-frame point cloud data based on the third point cloud data.

[0142] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, the computer program being stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer is able to execute the occupancy tag generation method provided by the above methods, the method including: for each non-key frame point cloud data in multi-frame point cloud data, determining, from all key frame point cloud data in the multi-frame point cloud data, a target key frame point cloud data whose acquisition time is closest to that of the non-key frame point cloud data; the key frame point cloud data includes at least one bounding box, the bounding box including point cloud data corresponding to the same dynamic object; based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determining a second point cloud data in the non-key frame point cloud data corresponding to the first point cloud data; concatenating the first point cloud data and the second point cloud data to obtain a third point cloud data corresponding to the dynamic object; and based on the third point cloud data, determining the occupancy tag corresponding to the multi-frame point cloud data.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0144] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An occupancy tag generation method, characterized by, include: For each non-critical frame point cloud data in the multi-frame point cloud data, determine the target critical frame point cloud data that is closest to the acquisition time of the non-critical frame point cloud data from all the critical frame point cloud data in the multi-frame point cloud data. The keyframe point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object. Based on the first point cloud data in the first target bounding box of the target key frame point cloud data, determine the second point cloud data in the non-key frame point cloud data that corresponds to the first point cloud data. The first point cloud data and the second point cloud data are concatenated to obtain the third point cloud data corresponding to the dynamic object; Based on the third point cloud data, determine the occupancy labels corresponding to the multi-frame point cloud data; The steps for determining the second point of cloud data include: Based on the position of the first target bounding box, a second target bounding box corresponding to the position of the first target bounding box is determined in the non-keyframe point cloud data, and the area of ​​the second target bounding box is greater than the area of ​​the first target bounding box. For each of the first point cloud data, the distance between the first point cloud data and each point cloud data in the second target bounding box is determined, and the point cloud data with the minimum distance less than a preset distance is determined as the second point cloud data.

2. The occupancy tag generation method of claim 1, wherein, The step of determining the occupancy labels corresponding to the multi-frame point cloud data based on the third point cloud data includes: For each key frame point cloud data in multi-frame point cloud data, based on the point cloud data in each bounding box in the key frame point cloud data, determine the fourth point cloud data corresponding to the static object in the key frame point cloud data. Based on the second point cloud data in the non-key frame point cloud data, determine the fifth point cloud data corresponding to the static object in the non-key frame point cloud data. The fourth point cloud data and the fifth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object; Based on the third point cloud data and the sixth point cloud data, the occupancy labels corresponding to the multi-frame point cloud data are determined.

3. The occupancy tag generation method of claim 2, wherein, The step of determining the occupancy tags corresponding to the multi-frame point cloud data based on the third point cloud data and the sixth point cloud data includes: For each target point cloud data without semantic meaning in the sixth point cloud data, determine the point cloud data with semantic meaning that is closest to the target point cloud data in the sixth point cloud data; The semantics of the semantic point cloud data are determined as the semantics of the target point cloud data; Based on the semantics of the third point cloud data, the semantics of the semantic point cloud data, and the semantics of the target point cloud data, the occupancy tags corresponding to the multi-frame point cloud data are determined.

4. The occupancy tag generation method of claim 2, wherein, The step of concatenating the fourth point cloud data and the fifth point cloud data to obtain the sixth point cloud data corresponding to the static object includes: The fourth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the seventh point cloud data in the world coordinate system; The fifth point cloud data is transformed from the radar coordinate system to the world coordinate system to obtain the eighth point cloud data in the world coordinate system; The seventh point cloud data and the eighth point cloud data are concatenated to obtain the sixth point cloud data corresponding to the static object.

5. The occupancy tag generation method according to any one of claims 1-4, wherein, The method further includes: Obtain the image corresponding to the keyframe point cloud data; Determine the third target bounding box corresponding to the target dynamic object in the image, and determine the first semantic segmentation mask of the target dynamic object in the third target bounding box based on the semantics of the ninth point cloud data corresponding to the target dynamic object in the key frame point cloud data. The third point cloud data is projected onto the first semantic segmentation mask. If there are point cloud data in the third point cloud data that do not semantically match the first semantic segmentation mask, the mismatched point cloud data is deleted.

6. The occupancy tag generation method of claim 2 or 3, wherein, The method further includes: Obtain the image corresponding to the keyframe point cloud data; The fourth point cloud data is projected onto the image, and based on the semantics of the fourth point cloud data in the keyframe point cloud data, the second semantic segmentation mask corresponding to the static object in the image is determined. The sixth point cloud data is projected onto the second semantic segmentation mask. If there are point cloud data in the sixth point cloud data that do not semantically match the second semantic segmentation mask, the mismatched point cloud data is deleted.

7. An occupancy tag generation apparatus characterized by comprising: include: The first determining module is used to determine, from all the key frame point cloud data in the multi-frame point cloud data, the target key frame point cloud data that is closest to the acquisition time of the non-key frame point cloud data in the multi-frame point cloud data. The keyframe point cloud data includes at least one bounding box, and the bounding box includes point cloud data corresponding to the same dynamic object. The second determining module is used to determine, based on the first point cloud data in the first target bounding box of the target key frame point cloud data, the second point cloud data in the non-key frame point cloud data that corresponds to the first point cloud data; The splicing module is used to splice the first point cloud data and the second point cloud data to obtain the third point cloud data corresponding to the dynamic object; The third determining module is used to determine the occupancy label corresponding to the multi-frame point cloud data based on the third point cloud data. The second determining module is specifically used for: Based on the position of the first target bounding box, a second target bounding box corresponding to the position of the first target bounding box is determined in the non-keyframe point cloud data, and the area of ​​the second target bounding box is greater than the area of ​​the first target bounding box. For each of the first point cloud data, the distance between the first point cloud data and each point cloud data in the second target bounding box is determined, and the point cloud data with the minimum distance less than a preset distance is determined as the second point cloud data.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the occupation tag generation method as described in any one of claims 1 to 6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the occupancy tag generation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Point cloud data processing method and system

    CN116052155A