Label processing method, device and equipment, readable storage medium and program product

By displaying a 3D view on a web page and responding to selection actions, the system automatically filters and transforms point cloud data, solving the complexity problem of desktop applications and achieving efficient point cloud annotation processing.

CN121582931APending Publication Date: 2026-02-27CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511760261.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing desktop application point cloud annotation technology is complex and requires professional knowledge, resulting in low annotation processing efficiency.

Method used

By displaying a 3D view on a webpage, the system automatically filters object point cloud data in response to selection operations, converts it into point cloud data in the vehicle coordinate system, and marks the 3D bounding boxes that determine the object type.

Benefits of technology

It simplifies the annotation process, improves annotation processing efficiency, and enables automated point cloud data filtering and annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582931A_ABST
    Figure CN121582931A_ABST
Patent Text Reader

Abstract

The invention relates to an annotation processing method and device, computer equipment, a computer readable storage medium and a computer program product. The method comprises the steps of displaying an annotated webpage containing a three-dimensional view, wherein the three-dimensional view is obtained by rendering point cloud data of a scene where a target object is located; in response to a frame selection operation on a target object in the three-dimensional view, displaying a rectangular frame including the target object, and based on frame coordinate data of the rectangular frame in a radar coordinate system, screening object point cloud data about the target object from the point cloud data, the radar coordinate system being a coordinate system constructed according to a radar on the target vehicle; converting the object point cloud data into converted point cloud data under a vehicle coordinate system, wherein the vehicle coordinate system is a coordinate system constructed according to the target vehicle; and determining the object stereo frame matched with the object type of the target object based on the converted point cloud data, and labeling the scene image of the scene based on the object stereo frame, thereby improving the labeling processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a labeling processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology

[0002] With the development of 3D data processing technology, point cloud annotation technology has emerged. This technology can annotate point cloud data in 3D space, providing key data support for subsequent 3D modeling, target recognition and other tasks.

[0003] In related technologies, desktop applications are mostly used for point cloud annotation. However, some desktop applications are quite complex to use, requiring users to have relevant professional knowledge and a series of operations, which makes the annotation process cumbersome and complicated, making it difficult to improve the efficiency of annotation processing. Summary of the Invention

[0004] Therefore, it is necessary to provide a labeling processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the efficiency of labeling processing in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides an annotation processing method, including:

[0006] Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located;

[0007] In response to a selection operation of a target object in the 3D view, a rectangular frame including the target object is displayed. Based on the frame coordinate data of the rectangular frame in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0008] The object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0009] Based on the transformed point cloud data, a stereo bounding box matching the object type of the target object is determined, and the scene image of the scene is annotated based on the stereo bounding box.

[0010] Secondly, this application also provides an annotation processing apparatus, comprising:

[0011] The view display module is used to display annotated web pages containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located;

[0012] The data filtering module is used to respond to the selection operation of the target object in the three-dimensional view, display a rectangular box including the target object, and filter out the object point cloud data about the target object from the point cloud data based on the box coordinate data of the rectangular box in the radar coordinate system. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0013] The data conversion module is used to convert the object point cloud data into transformed point cloud data in the vehicle coordinate system, wherein the vehicle coordinate system is a coordinate system constructed based on the target vehicle.

[0014] The image annotation module is used to determine an object 3D bounding box that matches the object type of the target object based on the converted point cloud data, and to annotate the scene image of the scene based on the object 3D bounding box.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0016] Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located;

[0017] In response to a selection operation of a target object in the 3D view, a rectangular frame including the target object is displayed. Based on the frame coordinate data of the rectangular frame in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0018] The object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0019] Based on the transformed point cloud data, a stereo bounding box matching the object type of the target object is determined, and the scene image of the scene is annotated based on the stereo bounding box.

[0020] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0021] Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located;

[0022] In response to a selection operation of a target object in the 3D view, a rectangular frame including the target object is displayed. Based on the frame coordinate data of the rectangular frame in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0023] The object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0024] Based on the transformed point cloud data, a stereo bounding box matching the object type of the target object is determined, and the scene image of the scene is annotated based on the stereo bounding box.

[0025] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:

[0026] Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located;

[0027] In response to a selection operation of a target object in the 3D view, a rectangular frame including the target object is displayed. Based on the frame coordinate data of the rectangular frame in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0028] The object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0029] Based on the transformed point cloud data, a stereo bounding box matching the object type of the target object is determined, and the scene image of the scene is annotated based on the stereo bounding box.

[0030] The aforementioned annotation processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product display an annotation webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located. In response to a bounding box selection operation on the target object in the 3D view, a rectangular bounding box including the target object is directly displayed. Then, based on the bounding box coordinates in the radar coordinate system, object point cloud data related to the target object is filtered from the point cloud data. The radar coordinate system is constructed based on the radar on the target vehicle, eliminating the need for manual filtering of point cloud data. Next, the object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is constructed based on the target vehicle. Based on the transformed point cloud data, the 3D bounding box matching the object type of the target object can be accurately determined, and the scene image is annotated based on the object 3D bounding box. Throughout the process, by performing a bounding box selection operation once on the annotation webpage, point cloud data matching the selected target object can be automatically filtered, and the 3D bounding box related to the target object can be determined through coordinate system transformation, thus achieving standard image processing, greatly simplifying the annotation process and effectively improving annotation processing efficiency. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is an application environment diagram of the annotation processing method in one embodiment;

[0033] Figure 2 This is a flowchart illustrating the annotation processing method in one embodiment;

[0034] Figure 3 This is a schematic diagram of point cloud annotation in one embodiment;

[0035] Figure 4 This is a schematic diagram of fused annotations in one embodiment;

[0036] Figure 5 This is a schematic diagram of point cloud segmentation in one embodiment;

[0037] Figure 6 This is a structural block diagram of the annotation processing device in one embodiment;

[0038] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0040] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0041] Before introducing the embodiments of this application, the following terms will be explained:

[0042] 3D point cloud (point cloud data): A collection of points with coordinate information in three-dimensional space, which are usually acquired by 3D scanning devices (such as laser scanners, depth cameras, etc.). Each point contains at least coordinates on the X, Y, and Z axes, and may sometimes contain other information such as color, intensity, or normals.

[0043] Web-based: Devices accessed and logged in using a web browser on a computer.

[0044] Point cloud segmentation: Dividing 3D point cloud data into different objects or regions so that each object or region can be processed separately.

[0045] Fusion annotation: Point cloud data and image data are registered and fused to map the targets constructed on the 3D point cloud onto the image, while annotating the 3D and 2D targets.

[0046] The annotation processing method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with radar 104 and camera device 106 via a network.

[0047] In some embodiments, terminal 102 acquires point cloud data of the scene where the target object is located, collected by radar 104. Terminal 102 renders a 3D view based on the point cloud data and displays a labeled webpage containing the 3D view. In response to a selection operation of the target object in the 3D view, terminal 102 displays a rectangular frame including the target object. Based on the frame coordinates of the rectangular frame in the radar coordinate system, it filters out object point cloud data related to the target object from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle. Terminal 102 converts the object point cloud data into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle. Based on the transformed point cloud data, terminal 102 determines an object stereo frame matching the object type of the target object. Terminal 102 acquires a scene image sent by camera device 106 and labels the scene image based on the object stereo frame.

[0048] The terminal 102 can be an in-vehicle terminal mounted on the target vehicle, or it can be another terminal that communicates with the in-vehicle terminal of the target vehicle. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc.

[0049] In one exemplary embodiment, such as Figure 2 As shown, a labeling processing method is provided, which can be applied to... Figure 1 Taking terminal 102 as an example, the explanation includes the following steps 202 to 208. Wherein:

[0050] Step 202: Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located.

[0051] The annotation webpage is the webpage on which the annotation is performed. The scene where the target object is located can be the scene around the target vehicle, such as the scene in front of the target vehicle. The target object is an object that belongs to the scene. For example, the target object is another vehicle located in front of the target vehicle, or the target object is a roadblock located in front of the target vehicle.

[0052] Optionally, the terminal uses React (a library for building dynamic, interactive single-page applications) to build a basic web-based single-page application to create an annotation webpage for annotation. After rendering the 3D view, the terminal displays the 3D view in the 2D screen drawing area of ​​the annotation webpage. For example, the step of determining the 2D screen drawing area includes: the terminal using Canvas (a native API provided by HTML5 for drawing graphics, animations, and rendering visual content on a webpage) to construct a 2D screen drawing area, which is used to record the drawing shape.

[0053] Optionally, the terminal uses threejs (a web-based 3D graphics library) to load point cloud files (in PCD format, a compressed binary PCD format, specifically designed for storing and processing 3D point cloud data), and renders and constructs a 3D view based on the point cloud data in the point cloud file.

[0054] Optionally, when the terminal receives a labeling task, it parses the task to obtain a timestamp, retrieves the point cloud data collected at that timestamp from the radar, and renders the point cloud data to obtain a 3D view. For example, when the terminal is an in-vehicle terminal of the target vehicle, the labeling task is automatically generated after the radar on the target vehicle collects the point cloud data. When the terminal is another terminal not in a vehicle, the labeling task can be generated by the server after acquiring data uploaded from multiple vehicles, and then broken down into multiple labeling tasks, each labeling task based on data from a specific vehicle. This data includes point cloud data sent by the radar on the corresponding vehicle and scene images sent by the camera equipment, both acquired at the same timestamp. The server acquires multiple other terminals in an idle state and assigns each labeling task to each of them. After acquiring the corresponding labeling task, each other terminal can begin executing step 202. This standardizes the annotation process, breaks down annotation tasks, and enables multiple team members to collaborate on annotating the same batch of data, so that subsequent annotation results can be output in a unified data format.

[0055] It should be noted that the labeled webpage is compatible with all common browsers currently available.

[0056] Step 204: In response to the selection operation of the target object in the 3D view, display the rectangular box including the target object. Based on the box coordinate data of the rectangular box in the radar coordinate system, filter out the object point cloud data about the target object from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0057] The bounding box selection operation involves selecting target objects within a 3D view. The bounding box is a 2D rectangle on the annotation grid. The radar coordinate system is a coordinate system constructed with a point on the radar of the target vehicle as its origin; for example, the radar's center can be used as the origin. Object point cloud data can be understood as the point cloud data of points belonging to the target object within the radar coordinate system.

[0058] Optionally, in response to a click operation on a target object in the 3D view, the terminal displays an initial rectangle of a preset size, with the top-left corner of the initial rectangle being the clicked point. In response to zooming in or out of the initial rectangle, a new rectangle including the target object is obtained, with the top-left corner of this new rectangle still being the clicked point. The terminal records the two-dimensional coordinate data of the rectangle, which is coordinate data in a two-dimensional coordinate system constructed with the center of the labeled webpage as the origin.

[0059] Optionally, the terminal converts the two-dimensional coordinate data of the rectangle in a two-dimensional coordinate system into coordinate data in a radar coordinate system to obtain the frame coordinate data. Based on the frame coordinate data, the terminal filters out the point cloud data projected within the rectangle from the point cloud data, and uses the filtered point cloud data as the object point cloud data.

[0060] In some embodiments, the step of determining the frame coordinate data of the rectangle in the radar coordinate system includes: determining the two-dimensional coordinate data of the rectangle, wherein the two-dimensional coordinate data is the coordinate data of each vertex of the rectangle in the two-dimensional coordinate system of the labeled webpage; obtaining a first coordinate system transformation matrix between the two-dimensional coordinate system and the radar coordinate system, and converting the two-dimensional coordinate data into frame coordinate data in the radar coordinate system.

[0061] For example, the terminal determines the two-dimensional coordinate data of the rectangle based on the selection operation. The terminal obtains a first coordinate system transformation matrix between the two-dimensional coordinate system and the radar coordinate system from the local machine, and multiplies the two-dimensional coordinate data with the first coordinate system transformation matrix to obtain the rectangle coordinate data in the radar coordinate system. It can be understood that the rectangle coordinate data includes the rectangle coordinates of each vertex, and the coordinate values ​​of each rectangle are the same under a certain radar coordinate axis.

[0062] In this embodiment, the transformation between the two-dimensional coordinate system and the radar coordinate system is realized through the first coordinate system transformation matrix, so as to transform the rectangle in the two-dimensional coordinate system to the radar coordinate system, unify the coordinate system, and ensure the effectiveness and accuracy of point cloud data filtering.

[0063] In some embodiments, based on the bounding box coordinates data of the rectangle in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data, including: based on the bounding box coordinates data of the rectangle in the radar coordinate system, determining the range of each radar coordinate under different radar coordinate axes in the radar coordinate system; based on each radar coordinate range, filtering point cloud data that satisfies each radar coordinate range from the point cloud data, and the filtered point cloud data is object point cloud data about the target object.

[0064] For example, for each radar coordinate axis in the radar coordinate system, the terminal determines the maximum and minimum values ​​on the radar coordinate axis in the frame coordinate data of the rectangle in the radar coordinate system, so as to obtain the radar coordinate range of the radar coordinate axis.

[0065] For example, for each radar coordinate axis, the terminal obtains the coordinate data (coordinate data in the radar coordinate system) of each point from the point cloud data. Based on the coordinate data of each point, it determines the coordinate value of each point on the radar coordinate axis. Based on the coordinate values ​​of each point on the radar coordinate axis, it filters out points whose coordinate values ​​are within the corresponding radar coordinate range, obtaining a filter set for that radar coordinate axis. This filter set includes the filtered points whose coordinate values ​​are within the corresponding radar coordinate range. The terminal obtains the filter set for each radar coordinate axis, determines the intersection of the filter sets, takes the points in the intersection as object points, and obtains the point cloud data of the object points from the point cloud data to serve as object point cloud data about the target object.

[0066] In the above embodiments, by using the frame coordinate data of the rectangle in the radar coordinate system, the point cloud data projected onto the rectangle can be filtered from the point cloud data, thereby automatically and accurately filtering the object point cloud data about the target object, improving the processing efficiency of the filtering stage.

[0067] Step 206: Convert the object point cloud data into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0068] Among them, the vehicle coordinate system is a coordinate system constructed with a certain point of the target vehicle as the origin, such as the vehicle coordinate system constructed with the center of the target vehicle as the origin.

[0069] For example, the terminal acquires a three-dimensional coordinate system rotation matrix, which is a second coordinate system transformation matrix for converting the radar coordinate system to the vehicle coordinate system. The terminal calculates the product of the object point cloud data and the three-dimensional coordinate system rotation matrix to obtain the transformed point cloud data in the vehicle coordinate system. This transformed point cloud data can be understood as the point cloud data of points belonging to the target object in the vehicle coordinate system.

[0070] Step 208: Based on the transformed point cloud data, determine the object 3D bounding box that matches the object type of the target object, and annotate the scene image of the scene based on the object 3D bounding box.

[0071] The object type reflects the type of the target object; for example, an object could be a vehicle or a pedestrian. The object bounding box is the three-dimensional frame that encloses the target object.

[0072] For example, the terminal identifies the object type of the target object based on the transformed point cloud data, and then matches the object 3D bounding box that matches the object type from multiple candidate 3D bounding boxes. The terminal then overlays and displays the object 3D bounding box in a 3D view. Figure 3 The image shown is a schematic diagram of point cloud annotation in one embodiment. Figure 3 The object's bounding box is a cuboid frame, and the object's bounding box represents the target object.

[0073] For example, the terminal uses object bounding boxes to annotate the scene image, obtaining an annotated scene image. This annotated scene image is obtained by overlaying object bounding boxes onto the target objects in the scene image. The object bounding box is a two-dimensional box projected onto the camera coordinate system by the object bounding box. It should be noted that the scene image is understood to be a two-dimensional image in the camera coordinate system. Therefore, annotating the scene image using object bounding boxes essentially projects or fuses the point cloud data of the target object, which is located in radar coordinates, onto or into the two-dimensional image in the camera coordinate system, thus achieving fused annotation. Figure 4 The image shown is a schematic diagram of fused annotations in one embodiment. Figure 4 The white circle on the left side outlines the target object selected in the 3D view. Figure 4 The white circle on the right outlines the labeled scene image. This is the effect of projecting the object's 3D bounding box onto the scene image's 2D bounding box and superimposing it onto the target object. As you can see, the 2D bounding box outlines the target object.

[0074] In some embodiments, the object type determination step for the target object includes: for each vehicle coordinate axis in the vehicle coordinate system, filtering the coordinate data of the target point from the transformed point cloud data, wherein the coordinate values ​​of the target point belonging to the vehicle coordinate axis are the maximum values; performing edge snapping processing based on the coordinate data of each of the filtered target points to determine the size of the three-dimensional model of the target object; and identifying the object type of the target object based on the size of the three-dimensional model.

[0075] For example, for each vehicle coordinate axis in the vehicle coordinate system, the coordinate values ​​of each point under that vehicle coordinate axis are obtained from the coordinate data of each point in the transformed point cloud data. The minimum or maximum value is selected from multiple coordinate values, and the point corresponding to the minimum value is taken as the target point, and the point corresponding to the maximum value is also taken as the target point. For example, 8 target points are selected from the coordinate data of each point in the transformed point cloud data. The target points can be understood as vertices.

[0076] For example, the terminal determines the initial length, width, and height of the 3D model based on each focal point, performs edge snapping based on the length, width, and height to fine-tune the shape, and then constructs the 3D model based on the final length, width, and height. Based on the dimensions of the 3D model (i.e., the length, width, and height of the 3D model), the terminal selects candidate object radars that match the dimensions from a preset pool of candidate object types; these are the object types of the target object.

[0077] In this embodiment, the size of the 3D model is determined by filtering the coordinate data of the target points from the transformed point cloud data. This allows for the automatic identification of the object type of the target object without the need for manual identification, thus improving the efficiency of identification while ensuring accuracy.

[0078] In some embodiments, annotating a scene image based on an object stereo frame includes: acquiring a scene image of the scene, wherein the scene image is obtained by a camera device deployed on the target vehicle capturing the scene; determining object camera data of the object stereo frame in the camera coordinate system where the camera device is located; and, based on the object camera data, performing target object annotation processing on the scene image to obtain an annotated scene image.

[0079] For example, the terminal determines the timestamp of the radar's point cloud data acquisition and acquires the scene image acquired at that timestamp. In other words, the radar and camera equipment deployed on the target vehicle acquire data at the same time.

[0080] For example, the terminal acquires the coordinate data of the object's 3D bounding box in the vehicle coordinate system and converts this coordinate data into object camera data in the camera coordinate system. Based on the object camera data, the terminal performs target object annotation processing on the scene image to obtain an annotated scene image, which is an image obtained by projecting the 3D bounding box onto the two-dimensional scene image.

[0081] In this embodiment, by determining the object's stereo frame in the camera coordinate system, the object's stereo frame can be projected onto the scene image, enabling point cloud data and image fusion annotation, and automatically generating an annotated scene image.

[0082] In some embodiments, determining the object camera data of the object stereo frame in the camera coordinate system where the camera device is located includes: obtaining the intrinsic and extrinsic parameter matrix of the camera device; obtaining the stereo frame coordinate data of the object stereo frame in the vehicle coordinate system; and calculating the product of the stereo frame coordinate data and the intrinsic and extrinsic parameter matrix to obtain the object camera data of the object stereo frame in the camera coordinate system where the camera device is located.

[0083] The intrinsic and extrinsic parameter matrices include the intrinsic parameter matrix and extrinsic parameter matrix of the camera device. The intrinsic parameter matrix includes the internal parameters of the camera device, and the extrinsic parameter matrix includes the position and attitude parameters of the camera device.

[0084] For example, the terminal obtains the 3D frame coordinate data of the object in the vehicle coordinate system, calculates the product of the 3D frame coordinate data and the extrinsic parameter matrix to obtain an intermediate product, normalizes the intermediate product, calculates the product of the normalized result and the intrinsic parameter matrix to obtain the object camera data of the object 3D frame in the camera coordinate system where the camera device is located.

[0085] In this embodiment, by using the intrinsic and extrinsic parameter matrix of the camera device, the coordinate data of the object's 3D bounding box can be converted from the vehicle coordinate system to the camera coordinate system, thus achieving automatic coordinate system conversion.

[0086] In the above annotation method, an annotation webpage containing a 3D view is displayed. This 3D view is obtained by rendering point cloud data of the scene where the target object is located. In response to a bounding box selection operation on the target object in the 3D view, a rectangular bounding box containing the target object is directly displayed. Then, based on the bounding box coordinates in the radar coordinate system, object point cloud data related to the target object is filtered from the point cloud data. The radar coordinate system is constructed based on the radar on the target vehicle, eliminating the need for manual point cloud data filtering. Next, the object point cloud data is converted to transformed point cloud data in the vehicle coordinate system, constructed based on the target vehicle. Based on the transformed point cloud data, the 3D bounding box matching the object type of the target object can be accurately determined. The scene image is then annotated based on the object 3D bounding box. Throughout this process, by performing a bounding box selection operation once on the annotation webpage, point cloud data matching the selected target object can be automatically filtered, and a 3D bounding box related to the target object can be determined through coordinate system transformation, achieving standard image processing. This greatly simplifies the annotation process and effectively improves annotation efficiency.

[0087] In some embodiments, the method further includes: the terminal drawing a polygon on a two-dimensional screen drawing area of ​​the labeled webpage, filtering out all points projected into the polygon from the point cloud data, adding color to the points, and identifying them as individual objects, thereby realizing the point cloud segmentation operation. Figure 5 The diagram shown is a schematic representation of point cloud segmentation in one embodiment. Figure 5 The white ellipse in the middle outlines the segmented object.

[0088] In one specific embodiment, the interaction between the server and various processing terminals is involved. The server acquires collected data within the current processing cycle. This collected data includes sub-data collected by at least one vehicle-mounted terminal at different acquisition times. The sub-data includes point cloud data and scene images simultaneously acquired by radar and camera equipment in the same vehicle. The server divides the collected data into sub-data segments and establishes corresponding annotation tasks based on each sub-data segment. Each sub-data segment includes point cloud data and scene images acquired on the same vehicle, with the point cloud data and scene images acquired at the same time. Based on the number of annotation tasks, the server acquires idle processing terminals. The number of processing terminals is the same as the number of annotation tasks, and each processing terminal handles one annotation task. For each processing terminal, the following steps are performed:

[0089] The processing terminal displays a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located.

[0090] In response to a selection operation of a target object in a 3D view, a rectangular box including the target object is displayed. The processing terminal determines the 2D coordinate data of the rectangular box, which is the coordinate data of each vertex of the rectangular box in the 2D coordinate system of the annotation webpage. The first coordinate system transformation matrix between the 2D coordinate system and the radar coordinate system is obtained, and the 2D coordinate data is converted into the box coordinate data in the radar coordinate system.

[0091] The processing terminal determines the radar coordinate range for each radar axis in the radar coordinate system based on the bounding box coordinates data of the rectangle. Based on each radar coordinate range, it filters point cloud data that satisfies each radar coordinate range; the filtered point cloud data is the object point cloud data related to the target object. The radar coordinate system is constructed based on the radar on the target vehicle.

[0092] The processing terminal converts the object point cloud data into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0093] For each vehicle coordinate axis in the vehicle coordinate system, the processing terminal filters the coordinate data of the target point from the transformed point cloud data, and the coordinate values ​​of the target point belonging to the vehicle coordinate axis are the maximum values; based on the coordinate data of each target point, edge snapping processing is performed to determine the size of the three-dimensional model of the target object; based on the size of the three-dimensional model, the object type of the target object is identified.

[0094] The processing terminal acquires the intrinsic and extrinsic parameter matrices of the camera device; acquires the coordinate data of the object's 3D bounding box in the vehicle coordinate system; calculates the product of the 3D bounding box coordinate data and the intrinsic and extrinsic parameter matrices to obtain the object's 3D bounding box in the camera coordinate system where the camera device is located.

[0095] The processing terminal acquires scene images of the scene, which are captured by cameras deployed on the target vehicle at the same time as the point cloud data acquisition. Based on the object camera data, the scene images are annotated with the target objects to obtain an annotated scene image.

[0096] The processing terminal stores the labeled scene images, point cloud data, and scene images according to the annotation requirements, and performs multi-level access control on these data to prevent data leakage; the data is backed up in the cloud to reduce the risk of data loss.

[0097] It should be noted that the above process can be applied to autonomous driving scenarios, and it is based on web-based endpoint cloud annotation technology. This technology typically has good cross-platform compatibility, running on different operating systems and browsers without requiring additional software installation; it can also automatically preprocess data of different formats. The annotation webpage is designed to be simple and intuitive, allowing users to easily start annotating and processing point cloud data.

[0098] It should be noted that the above-mentioned use of different processing terminals to collaboratively process the collected data means that it supports real-time collaboration among multiple people. Team members can jointly annotate and process point cloud data, improve work efficiency, achieve remote collaboration and real-time communication, standardize annotation operations, and ensure that the final annotation results are consistent.

[0099] Of course, after the server obtains the labeled scene images uploaded by each processing terminal, it can evaluate each labeled scene image separately to assess whether the corresponding target object is accurately labeled in the labeled scene image, and obtain the evaluation score of each labeled scene image. If there is an evaluation score that is less than the score threshold, the processing terminal with the highest evaluation score of the labeled scene image is determined, and this processing terminal is designated as the target processing terminal. The server sends the sub-data corresponding to the labeled scene images with scores less than the score threshold to the target processing terminal, so that the target processing terminal can re-execute the annotation processing method of this application and re-execute the annotation task for the sub-data.

[0100] In this embodiment, a webpage containing a 3D view is displayed. This 3D view is obtained by rendering point cloud data of the scene where the target object is located. In response to a selection operation of the target object in the 3D view, a rectangular bounding box containing the target object is directly displayed. Then, based on the bounding box coordinates in the radar coordinate system, object point cloud data related to the target object is filtered from the point cloud data. The radar coordinate system is constructed based on the radar on the target vehicle, eliminating the need for manual filtering of point cloud data. Next, the object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, constructed based on the target vehicle. Based on the transformed point cloud data, the 3D bounding box matching the object type of the target object can be accurately determined. The scene image is then annotated based on the object 3D bounding box. Throughout this process, by performing a single selection operation on the annotation webpage, point cloud data matching the selected target object can be automatically filtered, and a 3D bounding box related to the target object can be determined through coordinate system transformation, achieving standard image processing. This greatly simplifies the annotation process and effectively improves annotation efficiency.

[0101] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0102] Based on the same inventive concept, this application also provides an annotation processing apparatus for implementing the annotation processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more annotation processing apparatus embodiments provided below can be found in the limitations of the annotation processing method described above, and will not be repeated here.

[0103] In one exemplary embodiment, such as Figure 6 As shown, a labeling processing device 600 is provided, including: a view display module 602, a data filtering module 604, a data conversion module 606, and an image labeling module 608, wherein:

[0104] The view display module 602 is used to display an annotation webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located.

[0105] The data filtering module 604 is used to respond to the selection operation of the target object in the 3D view, display the rectangle including the target object, and filter the object point cloud data about the target object from the point cloud data based on the rectangle's frame coordinate data in the radar coordinate system. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle.

[0106] The data conversion module 606 is used to convert object point cloud data into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle.

[0107] The image annotation module 608 is used to determine the object 3D bounding box that matches the object type of the target object based on the transformed point cloud data, and to annotate the scene image of the scene based on the object 3D bounding box.

[0108] In some embodiments, the apparatus further includes a coordinate transformation module for determining two-dimensional coordinate data of a rectangular frame, wherein the two-dimensional coordinate data is the coordinate data of each vertex of the rectangular frame in the two-dimensional coordinate system of the labeled webpage; obtaining a first coordinate system transformation matrix between the two-dimensional coordinate system and the radar coordinate system, and converting the two-dimensional coordinate data into frame coordinate data in the radar coordinate system.

[0109] In some embodiments, the data filtering module 604 is used to determine the range of each radar coordinate under different radar coordinate axes in the radar coordinate system based on the frame coordinate data of the rectangle in the radar coordinate system; based on each radar coordinate range, it filters out point cloud data that meets each radar coordinate range from the point cloud data, and the filtered point cloud data is object point cloud data about the target object.

[0110] In some embodiments, the apparatus further includes a type determination module, configured to, for each vehicle coordinate axis in the vehicle coordinate system, filter out the coordinate data of the target point from the transformed point cloud data, wherein the coordinate values ​​of the target point belonging to the vehicle coordinate axis are the maximum values; perform edge snapping processing based on the coordinate data of each selected target point to determine the size of the three-dimensional model of the target object; and identify the object type of the target object based on the size of the three-dimensional model.

[0111] In some embodiments, the image annotation module 608 is used to acquire a scene image of the scene, which is obtained by a camera device deployed on the target vehicle to capture the scene; determine the object camera data of the object's 3D bounding box in the camera coordinate system where the camera device is located; and based on the object camera data, perform target object annotation processing on the scene image to obtain an annotated scene image.

[0112] In some embodiments, the image annotation module 608 is used to obtain the intrinsic and extrinsic parameter matrix of the camera device; obtain the coordinate data of the object's 3D bounding box in the vehicle coordinate system; and calculate the product of the 3D bounding box coordinate data and the intrinsic and extrinsic parameter matrix to obtain the object's 3D bounding box in the camera coordinate system where the camera device is located.

[0113] Each module in the aforementioned annotation processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0114] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 7 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a labeling processing method.

[0115] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0116] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0117] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0118] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0120] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0121] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0122] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A labeling processing method, characterized in that, The method includes: Display a labeled webpage containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located; In response to a selection operation of a target object in the 3D view, a rectangular frame including the target object is displayed. Based on the frame coordinate data of the rectangular frame in the radar coordinate system, object point cloud data about the target object is filtered from the point cloud data. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle. The object point cloud data is converted into transformed point cloud data in the vehicle coordinate system, which is a coordinate system constructed based on the target vehicle. Based on the transformed point cloud data, a stereo bounding box matching the object type of the target object is determined, and the scene image of the scene is annotated based on the stereo bounding box.

2. The method according to claim 1, characterized in that, The steps for determining the frame coordinates of the rectangular frame in the radar coordinate system include: Determine the two-dimensional coordinate data of the rectangle, wherein the two-dimensional coordinate data is the coordinate data of each vertex of the rectangle in the two-dimensional coordinate system of the labeled webpage; Obtain the first coordinate system transformation matrix between the two-dimensional coordinate system and the radar coordinate system, and convert the two-dimensional coordinate data into frame coordinate data in the radar coordinate system.

3. The method according to claim 1, characterized in that, The step of filtering object point cloud data about the target object from the point cloud data based on the bounding box coordinate data in the radar coordinate system includes: Based on the frame coordinate data of the rectangle in the radar coordinate system, the range of each radar coordinate under different radar coordinate axes in the radar coordinate system is determined respectively. Based on the coordinate range of each radar, point cloud data that meets the coordinate range of each radar is selected from the point cloud data. The selected point cloud data is object point cloud data about the target object.

4. The method according to claim 1, characterized in that, The step of determining the object type of the target object includes: For each vehicle coordinate axis in the vehicle coordinate system, the coordinate data of the target point is filtered out from the transformed point cloud data, and the coordinate values ​​of the target point that belong to the vehicle coordinate axis are the maximum values. Based on the coordinate data of each selected target point, edge snapping processing is performed to determine the size of the three-dimensional model of the target object; Based on the dimensions of the 3D model, the object type of the target object is identified.

5. The method according to claim 1, characterized in that, The annotation of the scene image based on the object's 3D bounding box includes: Acquire a scene image of the scene, which is obtained by a camera device deployed on the target vehicle capturing the scene; The object's 3D bounding box is determined in the camera coordinate system where the camera device is located. Based on the object's 3D bounding box, the scene image is annotated with the target object to obtain an annotated scene image.

6. The method according to claim 5, characterized in that, The process of determining the object's stereo frame in the camera coordinate system of the camera device includes: Obtain the intrinsic and extrinsic parameter matrix of the camera device; Obtain the coordinate data of the object's 3D bounding box in the vehicle coordinate system; The product of the 3D frame coordinate data and the intrinsic and extrinsic parameter matrices is calculated to obtain the object camera data of the object 3D frame in the camera coordinate system where the camera device is located.

7. A labeling processing device, characterized in that, The device includes: The view display module is used to display annotated web pages containing a 3D view, which is obtained by rendering point cloud data of the scene where the target object is located; The data filtering module is used to respond to the selection operation of the target object in the three-dimensional view, display a rectangular box including the target object, and filter out the object point cloud data about the target object from the point cloud data based on the box coordinate data of the rectangular box in the radar coordinate system. The radar coordinate system is a coordinate system constructed based on the radar on the target vehicle. The data conversion module is used to convert the object point cloud data into transformed point cloud data in the vehicle coordinate system, wherein the vehicle coordinate system is a coordinate system constructed based on the target vehicle. The image annotation module is used to determine an object 3D bounding box that matches the object type of the target object based on the converted point cloud data, and to annotate the scene image of the scene based on the object 3D bounding box.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.