Occupied grid labeling method and device
By combining the inter-frame pose transformation matrix and the detection box tracking mark, static and motion annotation results are generated and fused, which solves the problems of coarse labels and poor generalization ability in the existing technology and realizes efficient and accurate occupancy grid annotation.
Patent Information
- Application Number
- CN202511028081.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-09-30
AI Technical Summary
Existing occupancy grid labeling schemes have rough labels, poor generalization ability, are prone to false detection, and require a lot of manual fine-tuning, making it difficult to meet the requirements of efficiency, accuracy and generalization.
By extracting the inter-frame pose transformation matrix and detection box tracking identifiers, static and motion annotation results are generated, and the occupancy grid annotations of static objects and motion trajectories are fused and propagated in combination with the inter-frame pose transformation matrix to generate efficient and accurate occupancy grid annotation results.
It significantly improves the labeling efficiency, reduces manual workload, improves the labeling accuracy and generalization ability, and meets the efficiency, accuracy and generalization requirements of occupied grid labeling.
Smart Images

Figure CN120723926A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of grid annotation, and in particular to an occupied grid annotation method and device. Background Art
[0002] In the development of autonomous driving technology, accurate environmental perception is crucial for safe driving. Traditional scene representation based on obstacle boxes struggles with complex scenarios such as crowded, suspended, and irregularly shaped obstacles. Occupancy grid technology, by dividing the three-dimensional space into a fine grid, can accurately describe obstacles and drivable areas, becoming a mainstream solution. Its fine granularity and adaptability make it crucial for redundant design to ensure vehicle safety.
[0003] Occupancy grid technology relies on deep learning and requires massive amounts of manually annotated training data. However, in real-world scenarios, tens to hundreds of thousands of grids must be annotated for each frame of data. Manual annotation is inefficient and costly, making the development of efficient annotation solutions a key bottleneck in promoting the large-scale application of this technology.
[0004] To this end, existing technologies typically rely on manual annotation based on semantic segmentation of point clouds and automatic annotation solutions such as OpenOCC. However, manual annotation is cumbersome and inefficient in complex scenes and time-series frame tasks, while automatic annotation solutions produce coarse labels, have poor generalization capabilities, are prone to false detections, and require extensive manual fine-tuning after annotation, making it difficult to meet the efficiency, accuracy, and generalization requirements of occupancy grid annotation. Summary of the Invention
[0005] The present invention provides an occupancy grid annotation method and device, which solves the technical problems of existing annotation schemes, such as rough labels, poor generalization ability, easy false detection, and still requiring a lot of manual fine-tuning after annotation, which makes it difficult to meet the requirements of occupancy grid annotation for efficiency, accuracy and generalization.
[0006] The present invention provides an occupancy grid annotation method, comprising:
[0007] When receiving multiple frames of time series data, extracting the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data;
[0008] Performing grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result;
[0009] Tracking the position of the detection frame in the time series data of each frame in real time according to the detection frame tracking identifier, and constructing the motion trajectory of the target to which each detection frame belongs;
[0010] Performing occupation grid annotation on the area where each motion trajectory is located in the time series data of each frame to generate a motion annotation result;
[0011] The static annotation result and the motion annotation result are fused to generate an occupancy grid annotation result corresponding to the time series data.
[0012] Optionally, when receiving multiple frames of time series data, extracting an inter-frame pose transformation matrix and a detection frame tracking identifier corresponding to the time series data includes:
[0013] When receiving multiple frames of time series data, obtaining the posture transformation relationship information corresponding to each frame of the time series data;
[0014] Calculating an inter-frame pose transformation matrix between two adjacent frames of time series data according to each of the pose transformation relationship information;
[0015] The target detection model is called to detect the detection frame at the foreground position from the time series data, and each of the detection frames is marked with a detection frame tracking identifier.
[0016] Optionally, the performing occupancy grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result includes:
[0017] Performing grid marking on the stationary object in the first frame of the time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary object;
[0018] Calculate, according to the inter-frame pose transformation matrix, second occupied grid information corresponding to the first occupied grid information in the time series data of the remaining frames;
[0019] The grid attribute information is mapped to each of the time series data according to each of the second occupied grid information to generate a stationary labeling result corresponding to the stationary object.
[0020] Optionally, it also includes:
[0021] If the same updated stationary object appears in multiple consecutive frames in the time series data of the remaining frames, an occupation grid is marked for the updated stationary object to obtain new first occupied grid information and new grid attribute information;
[0022] Jump to the step of calculating the second occupied grid information corresponding to the time series data of the remaining frames of the first occupied grid information according to the inter-frame pose transformation matrix.
[0023] Optionally, it also includes:
[0024] If the number of missing static objects in the time series data of any frame reaches a preset object threshold, the time series data is determined as the new first frame time series data;
[0025] Jump to the step of marking the occupied grid of the stationary object in the time series data of the first frame to obtain the first occupied grid information and grid attribute information corresponding to the stationary object.
[0026] Optionally, the method is applied to a labeling platform; the step of labeling the stationary objects in the time series data with an occupied grid according to the inter-frame pose transformation matrix to generate a stationary labeling result includes:
[0027] Get the current platform memory resources;
[0028] If the current platform memory resources do not exceed the preset annotation consumption threshold, all the time series data are converted to the global coordinate system of the time series data of the first frame according to the inter-frame pose transformation matrix and spliced to obtain global time series data;
[0029] Identifying the locations of stationary objects in the global time series data and performing occupation grid annotation to obtain initial annotation information;
[0030] The initial annotation information is inversely transformed into a single-frame coordinate system corresponding to each frame of the time series data according to the inter-frame pose transformation matrix to obtain a static annotation result corresponding to the static object.
[0031] Optionally, performing occupation grid annotation on the area where each motion trajectory is located in each frame of the time series data to generate a motion annotation result includes:
[0032] Expanding each of the motion trajectories in the time series data of each frame according to a preset expansion size to obtain a trajectory coverage area;
[0033] If the time series data of the current frame does not contain an obstacle detection frame that overlaps with the track coverage area, marking all occupied grids in the track coverage area as being in a drivable state;
[0034] If the time series data of the current frame includes an obstacle detection frame that overlaps with the track coverage area, locating the overlapping area between the three-dimensional position of the obstacle detection frame and the track coverage area;
[0035] Marking the occupied grids to which the overlapping area belongs as occupied and marking the obstacle category;
[0036] When the time series data of all frames are labeled, a motion labeling result is generated.
[0037] Optionally, the fusing the static annotation result and the motion annotation result to generate an occupancy grid annotation result corresponding to the time series data includes:
[0038] fusing the static annotation results and the motion annotation results frame by frame to determine whether there is an occupied grid annotation conflict;
[0039] When there is an occupied grid annotation conflict between the static annotation result and the moving annotation result, obtaining all occupied grid information corresponding to the conflicting grids;
[0040] Target grid information is selected from each of the occupied grid information according to a preset annotation priority and updated to generate an occupied grid annotation result corresponding to the time series data.
[0041] Optionally, the method further includes:
[0042] After the occupancy grid annotation result is generated, conflict optimization is performed on the occupancy grid annotation result according to a preset semantic rationality constraint to generate a new occupancy grid annotation result.
[0043] The present invention also provides an occupancy grid marking device, comprising:
[0044] A data preprocessing module is used to extract the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data when receiving multiple frames of time series data;
[0045] A stationary occupancy grid annotation module is used to perform occupancy grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result;
[0046] A motion trajectory construction module is used to track the position of the detection frame in the time series data of each frame in real time according to the detection frame tracking identifier, and construct the motion trajectory of the target to which each detection frame belongs;
[0047] A motion occupancy grid annotation module is used to perform occupancy grid annotation on the area where each motion trajectory is located in each frame of the time series data to generate a motion annotation result;
[0048] The annotation fusion module is used to fuse the static annotation result and the motion annotation result to generate an occupied grid annotation result corresponding to the time series data.
[0049] It can be seen from the above technical solutions that the present invention has the following advantages:
[0050] When receiving multiple frames of time-series data, the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time-series data are extracted. Occupancy grid annotation is performed on stationary objects within the time-series data based on the inter-frame pose transformation matrix to generate a static annotation result. The position of the corresponding detection frame in each frame of time-series data is tracked in real time according to the detection frame tracking identifier to construct the motion trajectory of the target belonging to each detection frame. Occupancy grid annotation is performed on the area where each motion trajectory lies in each frame of time-series data to generate a motion annotation result. The static and motion annotation results are fused to generate the occupancy grid annotation result corresponding to the time-series data. This method propagates the occupancy grid annotation of stationary objects by combining the inter-frame pose transformation matrix. By fusing the occupancy grid annotations of stationary objects and motion trajectories, the efficiency, accuracy, and generalization requirements of occupancy grid annotation are effectively met. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 A flowchart of a method for marking an occupied grid provided by an embodiment of the present invention;
[0053] Figure 2 This is a flow chart of the steps of marking the occupancy grid of a stationary object in an embodiment of the present invention;
[0054] Figure 3 This is a structural block diagram of an occupancy grid annotation device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0055] Embodiments of the present invention provide an occupancy grid annotation method and device for solving the technical problems of existing annotation schemes, such as rough labels, poor generalization ability, prone to false detection, and still requiring a lot of manual fine-tuning after annotation, which makes it difficult to meet the requirements of occupancy grid annotation for efficiency, accuracy and generalization.
[0056] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0057] See also Figure 1 , Figure 1 A flowchart of the steps of an occupancy grid annotation method provided by an embodiment of the present invention.
[0058] The present invention provides an occupancy grid annotation method, comprising:
[0059] Step 101: When receiving multiple frames of time series data, extract the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data;
[0060] Time series data refers to structured data collected continuously in chronological order by lidar or millimeter-wave radar, typically including timestamp information and point cloud sequences. In this embodiment, the time series data can be the point cloud sequence collected frame by frame by the lidar of an autonomous vehicle over a specific time period (e.g., 10 seconds, 1 minute, 1 hour).
[0061] The inter-frame pose transformation matrix refers to the matrix that describes the coordinate system transformation of two adjacent frames (such as the t-th frame and the t+1-th frame) in three-dimensional space. It contains translation (position change) and rotation (posture change) information and is used to convert the coordinates of the previous frame to the coordinate system of the next frame.
[0062] The detection frame tracking identifier refers to the unique identifier assigned to each moving target using the multi-target tracking algorithm, which is used to establish the correspondence between targets in time series frames.
[0063] In this embodiment of the present application, multiple frames of point cloud data can be collected in chronological order by a LiDAR and uploaded as time series data to a labeling platform or other device with computing capabilities. The device pre-processes the time series data to extract the inter-frame pose transformation matrix and detection box tracking markers, which serve as the basis for subsequent occupancy grid annotation of stationary objects and motion trajectories.
[0064] In an example of the present application, step 101 may include the following sub-steps:
[0065] When receiving multiple frames of time series data, obtain the posture transformation relationship information corresponding to each frame of time series data;
[0066] According to the pose transformation relationship information, the inter-frame pose transformation matrix between two adjacent frames of time series data is calculated;
[0067] The target detection model is called to detect the detection box in the foreground position from the time series data, and each detection box is marked with a detection box tracking identifier.
[0068] The pose transformation relationship information refers to the global coordinate system pose corresponding to the current frame point cloud, including the rotation matrix and translation vector, which can be provided by the on-board positioning module in the autonomous driving system (such as GNSS+IMU fusion system, RTK, etc.).
[0069] In an embodiment of the present application, when receiving multiple frames of time series data input, after obtaining the posture transformation relationship information corresponding to each frame of time series data, the inter-frame posture transformation matrix between the two adjacent frames of time series data is calculated according to the posture transformation relationship information of the two adjacent frames of time series data. The inter-frame posture transformation matrix includes a relative rotation matrix and a relative translation vector. The specific calculation process can be as follows:
[0070] In the global coordinate system, point P is in the time series data of the i-th frame The coordinates in , the j-th frame timing data The coordinates in ,satisfy:
[0071] Jointly: Transformed into: ;(Note: yes The inverse matrix of , because the rotation matrix is an orthogonal matrix, , that is, transpose equals inverse)
[0072] Further organized into: , compared with the "pose transformation formula ”, we can get:
[0073] Relative rotation matrix: (3×3 matrix, describing Relative to rotation).
[0074] Relative translation vector: (3×1 vector, describing Relative to translation).
[0075] At the same time, the target detection model can be called to detect the detection box in the foreground position from the time series data, and each detection box can be marked with a detection box tracking identifier.
[0076] It should be noted that the target detection model can be PointRCNN, the voxelized model VoxelNet, or the Transformer-based model VoteNet. Taking PointRCNN as an example, PointNet++ is used to directly extract key points from the original point cloud and generate rough 3D candidate boxes (only distinguishing between foreground and background). Refined feature extraction and bounding box regression are performed on the points within the candidate boxes, and the final detection results are output. At the same time, each detection box is annotated with a detection box tracking identifier. The detection box can select vehicles, pedestrians, or other moving targets in the foreground. In each time series data, the same obstacle, regardless of when it appears, has the same corresponding detection box tracking identifier (Track ID) and is different from other obstacles.
[0077] In addition, when the number of detection frames is small, the detection frames in the foreground position can be selected and detection frame tracking marks can be added through manual labeling.
[0078] Step 102, performing occupation grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result;
[0079] Occupancy grid annotation refers to a scene representation method that divides the environment into grid cells, with each cell labeled as "occupied" (indicating an object), "free" (indicating no obstacles), or "unknown." In this context, it is used to quantitatively describe the spatial occupancy of objects in a specific frame.
[0080] In this embodiment, after obtaining the inter-frame pose transformation matrix, the stationary objects in each time series data are marked with occupied grids according to the current platform memory resources and the inter-frame pose transformation matrix of the annotation platform to generate static annotation results corresponding to each stationary object.
[0081] See also Figure 2 In one example of the present application, step 102 may include the following sub-steps S11-S13:
[0082] S11, marking the occupied grid of the stationary object in the first frame of time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary object;
[0083] S12, calculating the second occupied grid information corresponding to the first occupied grid information in the remaining frame time series data according to the inter-frame pose transformation matrix;
[0084] S13 . Map the grid attribute information to each time series data according to each second occupied grid information, and generate a stationary annotation result corresponding to the stationary object.
[0085] In order to reduce resource consumption while ensuring the accuracy of static labeling results, the occupation grid of each static object in the first frame of time series data, such as buildings, green plants, street lights, traffic signs, etc., can be annotated through a labeling model or manual labeling to obtain the first occupied grid information (i.e., the specific three-dimensional coordinate position) and grid attribute information of the static object, such as occupancy status and category information.
[0086] According to the inter-frame pose transformation matrix, the second occupied grid information of the first occupied grid information in the remaining frame time series data except the first frame is calculated. Specifically, for the coordinates in the first frame ( ) of the occupied grid, and calculate its corresponding coordinates in the tth frame through the coordinate transformation formula ( ):
[0087]
[0088] Finally, the grid attribute information is mapped to each time series data according to each second occupied grid information to assign a specific category and occupancy status to each stationary object, and generate a stationary labeling result corresponding to each stationary object in each frame of time series data.
[0089] Furthermore, step 102 may also include the following sub-steps:
[0090] If the same updated stationary object appears in multiple consecutive frames in the remaining frame time series data, the updated stationary object is marked with an occupied grid to obtain new first occupied grid information and new grid attribute information;
[0091] Jump to the step of calculating the second occupied grid information corresponding to the first occupied grid information in the remaining frame timing data according to the inter-frame pose transformation matrix.
[0092] In this embodiment, there may be occluded stationary objects in the first frame of time series data. As the autonomous vehicle moves, these objects appear and remain in the time series data, which can easily lead to missing occupancy grid annotations for the stationary objects. Therefore, if the same updated stationary object appears in multiple consecutive frames of the remaining frame time series data, such as new green plants appearing in 10 consecutive frames, the updated stationary object can be annotated with an occupied grid to obtain new first occupied grid information and new grid attribute information. The process then skips to steps S12-S13 to annotate the occupied grids for the remaining frames of the subsequent time series data, effectively avoiding occlusion.
[0093] Furthermore, step 102 may also include the following sub-steps:
[0094] If the number of missing static objects in any frame of time series data reaches a preset object threshold, the time series data is determined as the new first frame of time series data;
[0095] Jump to the step of marking the occupied grid of the stationary object in the first frame of time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary object.
[0096] In this embodiment of the present application, when annotating the grid occupied by stationary objects in time series data, if the number of missing stationary objects detected in a certain frame of time series data reaches a preset object threshold, it indicates that the vehicle in the time series data may have undergone a significant scene change, such as turning or entering a tunnel. In this case, the new first frame of time series data can be determined based on this frame of time series data, and the execution jumps to steps S11-S13 to avoid stationary object annotation errors caused by significant scene changes.
[0097] In a specific implementation, after generating the static annotation results, each frame of the static annotation results can be fine-tuned separately to supplement the annotations of objects that are occluded or not labeled due to being far away, or mislabeled due to posture errors.
[0098] In another example of the present application, the method is applied to a labeling platform; step 102 may include the following sub-steps:
[0099] Get the current platform memory resources;
[0100] If the current platform memory resources do not exceed the preset annotation consumption threshold, all time series data are converted to the global coordinate system of the first frame time series data according to the inter-frame pose transformation matrix and spliced to obtain the global time series data;
[0101] Identify the locations of stationary objects in the global time series data and perform grid annotation to obtain initial annotation information;
[0102] According to the inter-frame pose transformation matrix, the initial annotation information is inversely transformed into the single-frame coordinate system corresponding to each frame of time series data to obtain the static annotation result corresponding to the stationary object.
[0103] In this embodiment, the consumption of the current platform's memory resources can be monitored in real time through the annotation platform. If the current platform's memory resources do not exceed the preset annotation consumption threshold, it indicates that high-resource-consuming annotation behavior can be performed at this time, in order to improve the annotation efficiency of stationary objects. The point cloud in each frame of time series data can be converted to the global coordinate system to which the first frame of time series data belongs according to the inter-frame position transformation matrix, and after frame-by-frame splicing with the first frame of time series data, the global time series data corresponding to the entire acquisition cycle is obtained. Then, by calling the object recognition model or manually identifying the position of the stationary object in the global time series number, and annotating the occupied grid for the position, the initial annotation information is obtained. Among them, the initial annotation information includes its position information, status and type information, and its content is the sum of the first occupied grid information and grid attribute information. Finally, the initial annotation information is inversely transformed into the single-frame coordinate system corresponding to each frame of time series data according to the inter-frame posture transformation matrix, thereby obtaining the static annotation results of the stationary objects in each frame of time series data.
[0104] The inter-frame pose transformation matrix automatically propagates the annotation information of stationary objects from the first frame to subsequent frames, avoiding repeated annotation of the same stationary objects in each frame. This mechanism, based on the physical prior that stationary objects have a fixed position in the world coordinate system (global coordinate system), ensures accurate propagation of annotation information through precise coordinate transformation.
[0105] Step 103: Track the position of the detection frame in each frame of time series data in real time according to the detection frame tracking identifier, and construct a motion trajectory of the target to which each detection frame belongs;
[0106] A detection box is the bounding box of an object detected in a single frame, typically using rectangular coordinates to represent the object's position, size, and extent (such as in an image or point cloud). Detection boxes are used to identify and locate targets and are the basis for constructing motion trajectories.
[0107] While stationary objects are marked with occupied grids, the positions of each detection frame in each frame of time series data are tracked in real time using the detection frame tracking identifier. The position changes of each detection frame reflect the position changes of the moving target in each frame of time series data. Therefore, the detection frame tracking identifier can be used to index the positions of all detection frames in the entire time series data and connect them to construct the motion trajectory of the target belonging to each detection frame.
[0108] The motion trajectory can be a sequence of positions of a moving target that changes over time, describing its motion path (such as position, speed, or direction).
[0109] Step 104 , performing occupation grid annotation on the area where each motion trajectory is located in each frame of time series data to generate a motion annotation result;
[0110] In this embodiment, after obtaining the motion trajectories corresponding to the tracking identifiers of each detection frame, occupancy grid annotation is performed on the area where each motion trajectory is located in each frame of time series data to dynamically annotate the area where the motion trajectory is located. For example, the grid unit changes over time to represent the occupancy of the moving object, thereby generating a motion annotation result that is different from the static annotation result.
[0111] In one example of the present application, step 104 may include the following sub-steps:
[0112] Expand each motion trajectory according to a preset expansion size in each frame of time series data to obtain a trajectory coverage area;
[0113] If the current frame time series data does not contain an obstacle detection frame that overlaps with the track coverage area, all occupied grids in the track coverage area are marked as drivable;
[0114] If the current frame time series data contains an obstacle detection frame that overlaps with the track coverage area, locate the overlapping area between the 3D position of the obstacle detection frame and the track coverage area;
[0115] Mark the occupied grids in the overlapping area as occupied and mark the obstacle category;
[0116] When all frame timing data are labeled, the motion labeling result is generated.
[0117] In an embodiment of the present application, taking into account that autonomous driving vehicles have different actual sizes and safe driving distances, the superposition value of the actual size and the safe driving distance can be used as a preset expansion size, and each motion trajectory in each time series data can be three-dimensionally expanded according to the preset expansion size. For example, the motion trajectory can be used as the center line to expand to a width of 0.5 meters and 1 meter on both sides, etc. The specific expansion size can be adaptively adjusted according to different vehicles to obtain the trajectory coverage area.
[0118] Each frame of time series data is used as a unit to check whether there are obstacle detection frames that overlap with any track coverage area. Since the detection frame in the foreground area is not limited to the current road, there may be areas where multiple objects overlap, such as intersections, turns, or zebra crossings. Therefore, the obstacle detection model can be used to perform obstacle detection on the foreground area to filter out the obstacle detection frames belonging to the foreground obstacles, and then determine whether they overlap with the track coverage area one by one.
[0119] If no obstacle detection box overlaps with the track coverage area, all occupied grid cells within the track coverage area are marked as drivable. This marking method accurately describes the drivable area (i.e., the area with flat road surface), which is beneficial for the safety redundancy design of the vehicle.
[0120] If there is an obstacle detection frame in the current frame time series data that overlaps with the trajectory coverage area, the overlapping area of the three-dimensional position of the obstacle detection frame and the trajectory coverage area is located, and the grid occupied by the overlapping area is marked as occupied and the corresponding obstacle category such as pedestrians, vehicles, etc. is marked, so as to ensure that the actual occupancy status of the time series data at all times is correctly marked.
[0121] This method uses target tracking technology to extract the trajectory of moving targets. Based on the physical prior that "areas traveled by vehicles are drivable," it automatically labels the areas covered by the trajectory as drivable. This method integrates the annotation information of foreground obstacles (the annotation cost is low at the granularity of the foreground frame), achieving efficient automatic labeling of drivable areas.
[0122] Step 105 : Fusing the static annotation result and the motion annotation result to generate an occupancy grid annotation result corresponding to the time series data.
[0123] Since the static annotation results and the motion annotation results belong to different targets and may overlap, in order to obtain accurate occupancy grid annotation results in the future, after obtaining the static annotation results and motion annotation results corresponding to each frame of time series data, the static annotation results and motion annotation results of each frame of time series data are fused frame by frame, and the annotation conflicts between the two are eliminated frame by frame to generate the occupancy grid annotation results corresponding to the time series data.
[0124] In one example of the present application, step 105 may include the following sub-steps:
[0125] Fuse the static and motion annotation results frame by frame to determine whether there is a conflict in the occupied grid annotations.
[0126] When there is an occupied grid annotation conflict between the static annotation result and the moving annotation result, obtain all occupied grid information corresponding to the conflicting grid;
[0127] According to the preset annotation priority, the target grid information is selected from each occupied grid information and updated to generate the occupied grid annotation results corresponding to the time series data.
[0128] Occupancy grid information refers to the annotation status, annotation category and other related attribute information of the target associated with the occupation grid.
[0129] In an embodiment of the present application, the static annotation results and the motion annotation results are fused frame by frame, and it is determined in real time whether there is an occupied grid annotation conflict. If so, it indicates that the conflicting grid may have both static annotations and motion annotations. At this time, the conflicting grid can be located and all the corresponding occupied grid information can be obtained. According to the preset annotation priority, the target grid information is selected from each occupied grid information and updated until all the static annotation results and motion annotation results are fused to generate the occupied grid annotation results corresponding to the time series data. Specifically, in the priority processing process, the actual occupancy state of the current frame has the highest priority, that is, the obstacle detection frame; the static object annotation is second, and the drivable area annotation has the lowest priority. Finally, a complete occupied grid label is generated, which contains information such as occupancy state, target category, and drivable attributes as the occupied grid annotation results corresponding to the time series data.
[0130] To address potential conflicts between static object annotation and motion trajectory annotation, a priority-based fusion strategy was designed to ensure consistency and accuracy of the final annotation results. This fusion mechanism takes into account the reliability differences between different annotation sources, ensuring annotation quality.
[0131] Furthermore, the method may further comprise the following steps:
[0132] After the occupancy grid annotation result is generated, the conflict optimization of the occupancy grid annotation result is performed according to the preset semantic rationality constraints to generate a new occupancy grid annotation result.
[0133] In an embodiment of the present application, after the occupancy grid annotation results are generated, a preset semantic rationality constraint can be introduced to perform conflict optimization on the occupancy grid annotation results to generate a new occupancy grid annotation result. This semantic rationality constraint can be encapsulated through a logical expression or conditional judgment rule. For example, if grid G is "wall occupied" in the static annotation and "vehicle occupied" in the moving annotation, a conflict is triggered and correction is required; if grid G is occupied by both "pedestrian" and "building" annotations, it is determined to be unreasonable and requires adjustment.
[0134] The semantic rationality constraints may include, but are not limited to, "vehicles (moving targets) cannot be embedded in walls (stationary objects)" and "pedestrians cannot penetrate street lamp bases."
[0135] Compared to existing point cloud semantic segmentation and annotation methods, this invention significantly improves annotation efficiency by introducing physical priors. Traditional methods require independent bounding box extraction and category labeling for each frame of data. However, this invention leverages the temporal invariance of stationary objects, automatically completing the annotation of stationary objects throughout the entire time series by simply labeling the first frame, reducing the annotation workload by over 90%. Furthermore, the automatic labeling of drivable areas based on motion trajectories further reduces manual intervention, improving overall annotation efficiency by approximately 5-8 times.
[0136] For a task labeling 15 seconds of time-series data from a certain urban road scene (acquisition frequency 2Hz, approximately 30 frames per segment, totaling 1200 frames), traditional methods would require a human annotator to complete the entire task in approximately 240 hours. Using our method, we first label the stationary objects in the first frame, which takes 2 hours. Then, we automatically label the drivable area using Track ID information. Manual verification and fine-tuning take 8 hours, for a total of 10 hours, representing a 24-fold efficiency improvement. The labeling accuracy, verified by human operators, reaches 96.8%, meeting practical application requirements.
[0137] As shown in Table 1 below:
[0138] Table 1
[0139]
[0140] Compared to automatic labeling solutions such as OpenOcc, this invention significantly improves labeling quality and reliability while maintaining high efficiency. Model-based approaches like OpenOcc suffer from poor generalization and coarse labeling, while this invention labels based on deterministic physical laws, avoiding the uncertainty of model predictions. Especially in complex scenarios and novel environments, this invention's labeling accuracy is 15-25% higher than pure model-based approaches, significantly reducing the workload of subsequent manual corrections.
[0141] In an embodiment of the present application, when multiple frames of time series data are received, the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data are extracted; the stationary objects in the time series data are marked with occupancy grids according to the inter-frame pose transformation matrix to generate a stationary marking result; the position of the detection frame in each frame of time series data is tracked in real time according to the detection frame tracking identifier to construct the motion trajectory of the target to which each detection frame belongs; the area where each motion trajectory is located is marked with occupancy grids in each frame of time series data to generate a motion marking result; the stationary marking result and the motion marking result are fused to generate the occupancy grid marking result corresponding to the time series data. Thus, the occupancy grid annotation of the stationary object is propagated in combination with the inter-frame pose transformation matrix, and at the same time, by fusing the occupancy grid annotations of the stationary object and the motion trajectory, the requirements of the occupancy grid annotation for efficiency, accuracy and generalization are effectively met.
[0142] See also Figure 3 , Figure 3 This is a structural block diagram of an occupancy grid annotation device in an embodiment of the present invention.
[0143] An embodiment of the present invention provides an occupancy grid annotation device, comprising:
[0144] The data preprocessing module 301 is used to extract the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data when receiving multiple frames of time series data;
[0145] The stationary occupancy grid annotation module 302 is used to perform occupancy grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result;
[0146] The motion trajectory construction module 303 is used to track the position of the detection frame in each frame of time series data in real time according to the detection frame tracking identifier, and construct the motion trajectory of the target to which each detection frame belongs;
[0147] The motion occupancy grid annotation module 304 is used to perform occupancy grid annotation on the area where each motion trajectory is located in each frame of time series data to generate a motion annotation result;
[0148] The annotation fusion module 305 is used to fuse the static annotation results and the motion annotation results to generate an occupancy grid annotation result corresponding to the time series data.
[0149] Optionally, the data preprocessing module 301 is specifically configured to:
[0150] When receiving multiple frames of time series data, obtain the posture transformation relationship information corresponding to each frame of time series data;
[0151] According to the pose transformation relationship information, the inter-frame pose transformation matrix between two adjacent frames of time series data is calculated;
[0152] The target detection model is called to detect the detection box in the foreground position from the time series data, and each detection box is marked with a detection box tracking identifier.
[0153] Optionally, the static occupancy grid labeling module 302 is specifically configured to:
[0154] Marking the occupied grids of the stationary objects in the first frame of time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary objects;
[0155] Calculate the second occupied grid information corresponding to the first occupied grid information in the remaining frame time series data according to the inter-frame pose transformation matrix;
[0156] The grid attribute information is mapped to each time series data according to each second occupied grid information to generate a stationary annotation result corresponding to the stationary object.
[0157] Optionally, the static occupied grid labeling module 302 is further configured to:
[0158] If the same updated stationary object appears in multiple consecutive frames in the remaining frame time series data, the updated stationary object is marked with an occupied grid to obtain new first occupied grid information and new grid attribute information;
[0159] Jump to the step of calculating the second occupied grid information corresponding to the first occupied grid information in the remaining frame timing data according to the inter-frame pose transformation matrix.
[0160] Optionally, the static occupied grid labeling module 302 is further configured to:
[0161] If the number of missing static objects in any frame of time series data reaches a preset object threshold, the time series data is determined as the new first frame of time series data;
[0162] Jump to the step of marking the occupied grid of the stationary object in the first frame of time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary object.
[0163] Optionally, the method is applied to a labeling platform; the static occupancy grid labeling module 302 is specifically used to:
[0164] Get the current platform memory resources;
[0165] If the current platform memory resources do not exceed the preset annotation consumption threshold, all time series data are converted to the global coordinate system of the first frame time series data according to the inter-frame pose transformation matrix and spliced to obtain the global time series data;
[0166] Identify the locations of stationary objects in the global time series data and perform grid annotation to obtain initial annotation information;
[0167] According to the inter-frame pose transformation matrix, the initial annotation information is inversely transformed into the single-frame coordinate system corresponding to each frame of time series data to obtain the static annotation result corresponding to the stationary object.
[0168] Optionally, the motion trajectory construction module 303 is specifically configured to:
[0169] Expand each motion trajectory according to a preset expansion size in each frame of time series data to obtain a trajectory coverage area;
[0170] If the current frame time series data does not contain an obstacle detection frame that overlaps with the track coverage area, all occupied grids in the track coverage area are marked as drivable;
[0171] If the current frame time series data contains an obstacle detection frame that overlaps with the track coverage area, locate the overlapping area between the 3D position of the obstacle detection frame and the track coverage area;
[0172] Mark the occupied grids in the overlapping area as occupied and mark the obstacle category;
[0173] When all frame timing data are labeled, the motion labeling result is generated.
[0174] Optionally, the annotation fusion module 305 is specifically configured to:
[0175] Fuse the static and motion annotation results frame by frame to determine whether there is a conflict in the occupied grid annotations.
[0176] When there is an occupied grid annotation conflict between the static annotation result and the moving annotation result, obtain all occupied grid information corresponding to the conflicting grid;
[0177] According to the preset annotation priority, the target grid information is selected from each occupied grid information and updated to generate the occupied grid annotation results corresponding to the time series data.
[0178] Optionally, the device further includes a semantic optimization module, specifically configured to:
[0179] After the occupancy grid annotation result is generated, the conflict optimization of the occupancy grid annotation result is performed according to the preset semantic rationality constraints to generate a new occupancy grid annotation result.
[0180] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0181] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0182] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of the present embodiment according to actual needs.
[0183] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for marking an occupied grid, characterized in that: include: When receiving multiple frames of time series data, extracting the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data; Performing grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result; Tracking the position of the detection frame in the time series data of each frame in real time according to the detection frame tracking identifier, and constructing the motion trajectory of the target to which each detection frame belongs; Performing occupation grid annotation on the area where each motion trajectory is located in the time series data of each frame to generate a motion annotation result; The static annotation result and the motion annotation result are fused to generate an occupancy grid annotation result corresponding to the time series data.
2. The method according to claim 1, characterized in that When receiving multiple frames of time series data, extracting the inter-frame pose transformation matrix and the detection frame tracking identifier corresponding to the time series data includes: When receiving multiple frames of time series data, obtaining the posture transformation relationship information corresponding to each frame of the time series data; Calculating an inter-frame pose transformation matrix between two adjacent frames of time series data according to each of the pose transformation relationship information; The target detection model is called to detect the detection frame at the foreground position from the time series data, and each of the detection frames is marked with a detection frame tracking identifier.
3. The method according to claim 1, characterized in that The performing occupation grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result includes: Performing grid marking on the stationary object in the first frame of the time series data to obtain first occupied grid information and grid attribute information corresponding to the stationary object; Calculate, according to the inter-frame pose transformation matrix, second occupied grid information corresponding to the first occupied grid information in the time series data of the remaining frames; The grid attribute information is mapped to each of the time series data according to each of the second occupied grid information to generate a stationary labeling result corresponding to the stationary object.
4. The method according to claim 3, characterized in that Also includes: If the same updated stationary object appears in multiple consecutive frames in the time series data of the remaining frames, an occupation grid is marked for the updated stationary object to obtain new first occupied grid information and new grid attribute information; Jump to the step of calculating the second occupied grid information corresponding to the time series data of the remaining frames of the first occupied grid information according to the inter-frame pose transformation matrix.
5. The method according to claim 3, characterized in that Also includes: If the number of missing static objects in the time series data of any frame reaches a preset object threshold, the time series data is determined as the new first frame time series data; Jump to the step of marking the occupied grid of the stationary object in the time series data of the first frame to obtain the first occupied grid information and grid attribute information corresponding to the stationary object.
6. The method according to claim 1, characterized in that The method is applied to a labeling platform; the method performs occupation grid labeling on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary labeling result, including: Get the current platform memory resources; If the current platform memory resources do not exceed the preset annotation consumption threshold, all the time series data are converted to the global coordinate system of the time series data of the first frame according to the inter-frame pose transformation matrix and spliced to obtain global time series data; Identifying the locations of stationary objects in the global time series data and performing occupation grid annotation to obtain initial annotation information; The initial annotation information is inversely transformed into a single-frame coordinate system corresponding to each frame of the time series data according to the inter-frame pose transformation matrix to obtain a static annotation result corresponding to the static object.
7. The method according to claim 1, characterized in that The step of performing occupation grid annotation on the area where each motion trajectory is located in each frame of the time series data to generate a motion annotation result includes: Expanding each of the motion trajectories in the time series data of each frame according to a preset expansion size to obtain a trajectory coverage area; If the time series data of the current frame does not contain an obstacle detection frame that overlaps with the track coverage area, marking all occupied grids in the track coverage area as being in a drivable state; If the time series data of the current frame includes an obstacle detection frame that overlaps with the track coverage area, locating the overlapping area between the three-dimensional position of the obstacle detection frame and the track coverage area; Marking the occupied grids to which the overlapping area belongs as occupied and marking the obstacle category; When the time series data of all frames are labeled, a motion labeling result is generated.
8. The method according to claim 1, characterized in that The fusing the static annotation result and the motion annotation result to generate an occupancy grid annotation result corresponding to the time series data includes: fusing the static annotation results and the motion annotation results frame by frame to determine whether there is an occupied grid annotation conflict; When there is an occupied grid annotation conflict between the static annotation result and the moving annotation result, obtaining all occupied grid information corresponding to the conflicting grids; Target grid information is selected from each of the occupied grid information according to a preset annotation priority and updated to generate an occupied grid annotation result corresponding to the time series data.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: After the occupancy grid annotation result is generated, conflict optimization is performed on the occupancy grid annotation result according to a preset semantic rationality constraint to generate a new occupancy grid annotation result.
10. An occupancy grid annotation device, characterized in that: include: A data preprocessing module is used to extract the inter-frame pose transformation matrix and detection frame tracking identifier corresponding to the time series data when receiving multiple frames of time series data; A stationary occupancy grid annotation module is used to perform occupancy grid annotation on the stationary objects in the time series data according to the inter-frame pose transformation matrix to generate a stationary annotation result; A motion trajectory construction module is used to track the position of the detection frame in the time series data of each frame in real time according to the detection frame tracking identifier, and construct the motion trajectory of the target to which each detection frame belongs; A motion occupancy grid annotation module is used to perform occupancy grid annotation on the area where each motion trajectory is located in each frame of the time series data to generate a motion annotation result; The annotation fusion module is used to fuse the static annotation result and the motion annotation result to generate an occupied grid annotation result corresponding to the time series data.
Citation Information
Cited By
Perception labeling method of parking scene, related device and storage medium
CN121459353A