A data annotation method, electronic device, and storage medium

By unifying and integrating the initial annotation box information of multiple sensor devices in the bird's-eye view coordinate system, the problems of low efficiency and low accuracy of 3D point cloud data and 2D image data annotation in road test intelligent perception technology are solved, and efficient and high-precision image data annotation is achieved.

CN122312992APending Publication Date: 2026-06-30ZHEJIANG DAHUA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2026-04-09
Publication Date
2026-06-30

Smart Images

  • Figure CN122312992A_ABST
    Figure CN122312992A_ABST
Patent Text Reader

Abstract

This application discloses a data annotation method, electronic device, and storage medium. The method includes: annotating raw sensor data collected by multiple sensor devices to obtain several initial bounding box information corresponding to the same target; fusing the several initial bounding box information in a bird's-eye view coordinate system to obtain fused bounding box information of the target; and updating the initial bounding box information of the target in each raw sensor data based on adjustment information of the fused bounding box information. This approach can improve the efficiency and accuracy of image data annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image annotation technology, and in particular to a data annotation method, electronic device, and storage medium. Background Technology

[0002] In roadside intelligent sensing technology scenarios, multi-sensor devices and edge computing units are combined to achieve all-weather, high-precision perception of roadside traffic participants and road conditions. The variety of multi-sensor devices and their raw sensor data is diverse, including 3D point cloud data and 2D image data. Traditional annotation methods involve separately annotating 3D point cloud data and 2D image data, resulting in low annotation efficiency. Furthermore, when converting 2D image data to 3D-assisted annotation, it is difficult to obtain the true 3D positioning information of the target, leading to low annotation accuracy.

[0003] Therefore, there is an urgent need in the market for an image data annotation method to solve the above-mentioned technical problems. Summary of the Invention

[0004] This application provides at least one data annotation method, electronic device, and storage medium that can improve the efficiency and accuracy of image data annotation.

[0005] The first aspect of this application provides a data annotation method, which includes: annotating raw sensor data collected by multiple sensor devices to obtain several initial annotation box information corresponding to the same target; fusing the several initial annotation box information in a bird's-eye view coordinate system to obtain fused annotation box information of the target; and updating the initial annotation box information of the target in each raw sensor data based on the adjustment information of the fused annotation box information.

[0006] The process involves labeling raw sensor data collected by multiple sensor devices to obtain several initial bounding boxes for the same target. This includes: identifying key points in the raw sensor data collected by each sensor device to obtain several target key points; performing bird's-eye view coordinate transformation on the target key points to obtain target bird's-eye view projection points; generating the target region in the bird's-eye view coordinate system based on the target bird's-eye view projection points; and mapping the target region to the sensor coordinate system corresponding to the raw sensor data to obtain the initial bounding box information of the target in the raw sensor data.

[0007] The target region is a three-dimensional region, and the initial bounding box information includes the projection area of ​​the target in the original sensor data and the height information of the target.

[0008] The initial bounding box information is the 3D rectangular bounding box of the target in the original sensor data. The bottom surface of the 3D rectangular bounding box is the projection area, and the height of the 3D rectangular bounding box represents the height information of the target.

[0009] Before performing bird's-eye view coordinate transformation on several target key points to obtain the target bird's-eye view projection points, the process includes: determining the first spatial coordinate transformation relationship between the coordinate system corresponding to each sensor device and the bird's-eye view coordinates; performing bird's-eye view coordinate transformation on several target key points to obtain the target bird's-eye view projection points includes: based on the first spatial coordinate transformation relationship, transforming several target key points into the bird's-eye view coordinate system to obtain the target bird's-eye view projection points.

[0010] Before fusing several initial bounding box information in the bird's-eye view coordinate system to obtain the target's fused bounding box information, the process includes: identifying the target sensor device and the sensor device to be converted among multiple sensor devices, wherein the sensor device to be converted is the device other than the target sensor device among the multiple sensor devices; determining a third spatial coordinate transformation relationship based on the coordinate system corresponding to the target sensor device and the bird's-eye view coordinate system; and determining a second spatial coordinate transformation relationship based on the coordinate system corresponding to the sensor device to be converted and the coordinate system corresponding to the target sensor device, respectively.

[0011] The process involves fusing several initial bounding box information in the bird's-eye view coordinate system to obtain the target's fused bounding box information. This includes: based on a second spatial coordinate transformation relationship, transforming the initial bounding box information corresponding to the original sensor data collected by the sensor device to be transformed into the coordinate system corresponding to the target sensor device to obtain several first transformed bounding box information; based on a third spatial coordinate transformation relationship, transforming the several first transformed bounding box information and the initial bounding box information corresponding to the target sensor device into the bird's-eye view coordinate system to obtain several second transformed bounding box information; and fusing the several second transformed bounding box information to obtain the first fused bounding box information.

[0012] The adjustment information includes data on the center position, size, and orientation of the target within the fused annotation box.

[0013] The second aspect of this application provides an electronic device including a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the data annotation method in the first aspect described above.

[0014] A third aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a processor, implement the data annotation method described in the first aspect above.

[0015] The above scheme involves labeling raw sensor data collected by multiple sensor devices to obtain several initial bounding boxes corresponding to the same target. These initial bounding boxes are then uniformly transformed into a bird's-eye view coordinate system and fused within that system to obtain a fused 3D bounding box for the target. Finally, adjustments to the fused 3D bounding box information are needed to update the initial bounding boxes for the target in each of the original sensor data sets. Compared to existing technologies, this approach eliminates the need for separate labeling of 3D and 2D image data; instead, it unifies the data into a bird's-eye view coordinate system for labeling and adjustment, thereby improving labeling efficiency and accuracy.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0018] Figure 1 This is a flowchart illustrating an embodiment of the data annotation method of this application; Figure 2 This is a schematic diagram of an embodiment of the initial annotation box information in this application; Figure 3 This is a schematic diagram of the framework of an embodiment of the display interface of the image annotation device of this application; Figure 4 This is a flowchart illustrating another embodiment of the data annotation method of this application; Figure 5 This is a schematic diagram of the framework of an embodiment of the data annotation device of this application; Figure 6 This is a schematic diagram of the framework of an embodiment of the electronic device of this application; Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0019] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0020] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0021] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0022] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the data annotation method of this application. Specifically, it may include the following steps: Step S110: Mark the target on the raw sensor data collected by multiple sensor devices to obtain several initial annotation boxes corresponding to the same target.

[0023] This method is primarily applied to the field of roadside intelligent sensing technology, particularly involving 3D annotation and BEV (Bird's Eye View) perception technology. Roadside intelligent sensing technology combines multi-sensor devices and edge computing units at the roadside to achieve all-weather, high-precision perception of road users and road conditions. For example, sensor devices are installed on multi-functional integrated poles or L-shaped poles along both sides of the road to collect data on motor vehicles, non-motor vehicles, and pedestrians. These sensor devices can be cameras, millimeter-wave radar, and lidar, and the raw sensor data collected can be 2D image data, point cloud data, etc.

[0024] This application transforms the initial bounding box information corresponding to each target in the raw sensor data collected by different sensor devices into the bird's-eye view coordinate system and then fuses them. Therefore, only one adjustment is needed in the bird's-eye view coordinate system, instead of adjusting separately in each raw sensor data, thereby reducing the amount of manual annotation. Furthermore, all raw sensor data in the bird's-eye view are converted into three-dimensional data, and the annotation data of three-dimensional data is more accurate than that of two-dimensional data.

[0025] The initial bounding box information consists of a 3D rectangular box representing the target in the original sensor data. The bottom surface of the 3D rectangular box is the projection area, and the height of the 3D rectangular box represents the target's height information. (See also...) Figure 2 The vehicle in the image is the target, and the 3D rectangle on the vehicle is the initial annotation information of the vehicle.

[0026] The data annotation method in this application can be executed by an image annotation device. For example, the image annotation method can be executed by a terminal device, a server, or other processing device. The terminal device can be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, the data annotation method can be implemented by a processor calling computer-readable instructions stored in memory.

[0027] In some implementations, the raw sensor data collected may include two-dimensional image data. However, three-dimensional target annotation of two-dimensional image data is inaccurate. Therefore, it is necessary to uniformly convert all data to a three-dimensional coordinate system for target annotation. Specifically, refer to steps S111 to S113.

[0028] Step S111: For the raw sensor data collected by each sensor device, perform key point identification on the raw sensor data to obtain several target key points corresponding to the target.

[0029] Key points of a target are representative feature locations of a target in an image, such as the center point or corner points of the target, which are obtained through image processing or deep learning algorithms.

[0030] Step S112: Perform bird's-eye view coordinate transformation on several target key points to obtain the target bird's-eye view projection points.

[0031] Bird's-eye view coordinate transformation uses the sensor device calibration relationship to map the sensor coordinates to a bird's-eye view coordinate system perpendicular to the ground, thus obtaining the coordinate representation of the target in the bird's-eye view.

[0032] Furthermore, before performing bird's-eye view coordinate transformation on several key target points to obtain the target bird's-eye view projection points, it is necessary to determine the first spatial coordinate transformation relationship between the coordinate system corresponding to each sensor device (i.e., the sensor coordinate system) and the bird's-eye view coordinates. This relationship is used to achieve the mapping of data from different devices under a unified coordinate system. For example, for a pure vision camera, the bird's-eye view coordinate system can be defined as a northeast-sky coordinate system with the camera facing due north. Combining intrinsic parameters and camera height, the first spatial coordinate transformation relationship between the camera and the BEV is calibrated using a parallel line self-calibration method.

[0033] After determining the first spatial coordinate transformation relationship, several target key points can be transformed into the bird's-eye view coordinate system based on the first spatial coordinate transformation relationship to obtain the target bird's-eye view projection points. This design ensures that the data from multiple devices maintains geometric consistency during the coordinate transformation process, avoids positioning deviations caused by inconsistencies in coordinate systems, and makes the transformation process precisely controllable.

[0034] By establishing a first spatial coordinate transformation relationship, the coordinate systems of each sensor can be accurately mapped to the bird's-eye view coordinate system, thereby improving the accuracy of multi-sensor data conversion. Data conversion based on this first spatial coordinate transformation relationship improves the geometric consistency of the representation of several target key points in the bird's-eye view coordinate system, thus enhancing the quality of basic target detection data.

[0035] Step S113: Based on the target's bird's-eye view projection points, generate the target area corresponding to the target in the bird's-eye view coordinate system.

[0036] Based on the target's bird's-eye view projection points, a target region in the bird's-eye view coordinate system is generated through geometric constraints. This region conforms to a rotating rectangular structure with height, ensuring the rigidity of the target's positioning.

[0037] The target region can be a three-dimensional region. The initial bounding box information includes the projection area of ​​the target onto the original sensor data and the target's height information. The target region refers to the three-dimensional spatial region representing the target in the bird's-eye view coordinate system. This region can be defined by its center point, orientation, length, width, and height, ensuring a rigid structure for target positioning. The projection area is the two-dimensional bounding box corresponding to the target in the original sensor data, generated based on the target's projection onto the original sensor. The height information is the target's height value in the vertical direction, used to describe the target's height in three-dimensional space. In this embodiment, the target region is a three-dimensional region, and its three-dimensional coordinates are mapped to the original sensor coordinate system through the sensor device calibration relationship to generate the initial bounding box information. Height information can be obtained directly from LiDAR point cloud data, the target height can be predicted using a deep learning model, or the height information can be combined with the projection area to generate the initial bounding box information; no specific limitations are made here. During application, the target area is defined as a rotated rectangle with height in the bird's-eye view coordinate system. When mapped back to the original sensor coordinate system, the initial annotation box information includes the projected area (two-dimensional boundary) and height information (vertical dimension), so that the annotation can fully reflect the three-dimensional characteristics of the target and avoid the positioning ambiguity caused by using only a two-dimensional box.

[0038] By defining the target area as a three-dimensional region, the target's positioning in three-dimensional space becomes more precise, thereby improving annotation accuracy. Including the projected area in the initial annotation box information makes the target's boundary positioning in the original sensor more realistic, thus improving annotation boundary matching. Including height information in the initial annotation box information makes the target's vertical dimension representation more accurate, thereby improving the consistency of three-dimensional annotation.

[0039] Step S114: Map the target area to the sensor coordinate system corresponding to the original sensor data to obtain the initial bounding box information of the target in the original sensor data.

[0040] By identifying key points, the projection area and orientation of the target on the ground can be determined more accurately, thereby improving the positioning accuracy of the initial bounding box. Bird's-eye view coordinate transformation makes the target representation more consistent in a unified coordinate system, thus improving the geometric consistency of the annotation. Mapping the target area back to the original sensor coordinate system results in a higher degree of matching between the initial bounding box and the actual target boundary, thereby improving the accuracy of the annotation.

[0041] Step S120: Merge several initial annotation box information in the bird's-eye view coordinate system to obtain the merged annotation box information of the target.

[0042] In some implementations, the Hungarian matching algorithm can be used to match and fuse several initial bounding box information in the bird's-eye view coordinate system to obtain the target's fused bounding box information. The specific fusion method is not specifically limited here.

[0043] In other implementations, to ensure that the initial bounding box information corresponding to each original image can be properly fused in the bird's-eye view coordinate system, a target sensor device can be selected from among many sensor devices. The initial bounding box information corresponding to the original sensor data of the other sensor devices can be converted to the target sensor coordinate system, and then the target sensor coordinate system can be converted to the bird's-eye view coordinate system. Before this, the conversion relationship between the target sensor device and the other sensor devices, as well as the conversion relationship between the target sensor device and the bird's-eye view, needs to be determined. Specifically, a target sensor device and a sensor device to be converted are determined among the multiple sensor devices, wherein the sensor device to be converted is the device other than the target sensor device among the multiple sensor devices. A second spatial coordinate conversion relationship is determined based on the coordinate system corresponding to the sensor device to be converted and the coordinate system corresponding to the target sensor device, respectively. A third spatial coordinate conversion relationship is determined based on the coordinate system corresponding to the target sensor device and the bird's-eye view coordinate system.

[0044] Subsequently, based on the second spatial coordinate transformation relationship, the initial annotation box information corresponding to the original sensor data collected by the sensor device to be transformed is transformed into the coordinate system corresponding to the target sensor device, resulting in several first transformed annotation box information; based on the third spatial coordinate transformation relationship, the several first transformed annotation box information and the initial annotation box information corresponding to the target sensor device are transformed into the bird's-eye view coordinate system, resulting in several second transformed annotation box information; the several second transformed annotation box information are fused to obtain first fused annotation box information.

[0045] The target sensor device is typically a high-precision device such as a lidar unit, serving as the reference for coordinate system transformation. The sensor device to be transformed is another sensor device, such as a camera. Data mapping between the target sensor device and the sensor device to be transformed is achieved through a second spatial coordinate transformation relationship, which is determined using multi-sensor calibration techniques, such as calculating the relative positions and angles between the sensor devices using point cloud data. Data mapping between the coordinate system corresponding to the target sensor device and the bird's-eye view coordinate system is achieved through a third spatial coordinate transformation relationship, which is determined through calibration methods, such as using a parallel line self-calibration method combined with camera intrinsic parameters and height parameters.

[0046] In the application process, the initial bounding box information of the sensor device to be converted is first mapped to the coordinate system corresponding to the target sensor device through a second spatial coordinate transformation relationship, ensuring that the data is aligned in the target sensor device's reference system. Next, the mapped first converted bounding box information and the initial bounding box information corresponding to the target sensor device are mapped to the bird's-eye view coordinate system through a third spatial coordinate transformation relationship, unifying all data to the bird's-eye view perspective. Finally, the mapped second converted bounding box information is fused. This step-by-step coordinate transformation avoids the coordinate inconsistency problems caused by directly processing in the original sensor coordinate system, making the fusion process more reliable.

[0047] Step S130: Based on the adjustment information of the fused bounding box information, update the initial bounding box information of the target in each original sensor data.

[0048] In some embodiments, the image annotation device also includes a display interface, such as... Figure 3As shown, the left side of the display interface displays raw sensor data (such as image data and 3D point cloud data) collected by multi-sensor devices, while the right side displays the fused bounding box information of each target in the BEV coordinate system. The image annotation device converts the fused bounding box information of each target back into the raw sensor data based on the first spatial coordinate transformation relationship. The operator uses the initial bounding box information in the raw sensor data on the left as a reference to adjust the fused bounding box information of each target in the BEV view on the right, thereby generating adjustment information. This adjustment information will synchronously update the initial bounding box information of the target in each raw sensor data in real time. For example, if the initial bounding box information of the same target in the two image data on the left does not completely cover the target, the operator will zoom in on the fused bounding box information of the target in the BEV view on the right, so that the initial bounding box information of the same target in the two image data on the left can completely cover the target.

[0049] The adjustment information includes data on the center position, size, and orientation of the target within the fused annotation box.

[0050] Furthermore, to ensure the accuracy of the adjustment information, it can be verified. Specifically, based on several initial bounding box information corresponding to the target, the initial adjustment information fused with the bounding box information is verified to obtain a verification result. The initial adjustment information includes at least one of the following: adding, deleting, or adjusting the bounding box information corresponding to the target. If the verification result is successful, the initial adjustment information is used as the adjustment information.

[0051] Initial bounding box information refers to a set of initial bounding boxes for the same target obtained from raw sensor data from different sensor devices. This information reflects the target's preliminary location in the raw sensor data. Fusion bounding box information is a unified representation obtained by fusing multiple initial bounding box information in the bird's-eye view coordinate system, used to integrate multi-sensor device data. Initial adjustment information refers to the modification operations to be performed on the fusion bounding boxes, including adding targets, deleting targets, or adjusting target bounding boxes. The verification process verifies the rationality of the adjustment information by comparing the consistency between the actual target in the raw sensor data and the fusion bounding box information, ensuring that the adjustment operation conforms to the actual situation of the target. A successful verification indicates that the adjustment information matches the actual target and no further correction is needed. For example, when the initial bounding box information shows that the target is occluded in the image, the verification process verifies whether the adjusted fusion bounding box is reasonable, avoiding labeling errors caused by occlusion. Another example is when the initial bounding box information comes from raw sensor data from multiple sensor devices, the verification process checks the consistency of data from different sensor devices to ensure that the adjustment information is applicable to all sensor devices. When the adjustment information includes adding a new target, the verification process checks whether the new target overlaps with the existing target position to prevent labeling conflicts.

[0052] By validating the initial adjustment information of the merged annotation boxes based on several initial annotation box information corresponding to the target, the adjusted information is made consistent with the actual target, thereby improving the accuracy of annotation correction. The validation process verifies the rationality of the initial adjustment information, preventing erroneous operations during the annotation update process and thus improving the reliability of annotation. When the initial adjustment information includes adding, deleting, or adjusting the annotation box information corresponding to the target, the validation ensures that these operations are compatible with the actual target, thereby reducing the annotation error rate.

[0053] This scheme first annotates the raw sensor data to obtain initial bounding boxes, then fuses these initial bounding boxes in the bird's-eye view coordinate system, and finally adjusts the initial bounding boxes in the raw sensor data based on the information of the fused bounding boxes. This design reduces the amount of manual annotation because it only requires adjustment once in the bird's-eye view coordinate system, instead of adjusting each piece of raw sensor data separately. LiDAR data can be used as a reference; each piece of raw sensor data is first converted to the LiDAR coordinate system, and then uniformly converted to the bird's-eye view coordinate system. Alternatively, a pure vision camera can be used, and the conversion relationship between the camera and the bird's-eye view can be calibrated using a parallel line self-calibration method based on intrinsic parameters and camera height. Generating initial bounding boxes by combining the target's ground-level projection points with the rigid structural constraints of the target in the bird's-eye view coordinate system makes the annotations on the image more accurate.

[0054] Please see Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the data annotation method of this application. Specifically, it may include the following steps: Step S410: Preprocess the raw sensor data collected by multiple sensor devices to obtain preprocessed sensor data.

[0055] In some implementations, distortion correction and noise reduction processing can be performed on the raw sensor data acquired by the camera device. Feature enhancement processing can be performed on the 3D point cloud data acquired by the LiDAR. Here, no specific limitations are made regarding the method of preprocessing the raw sensor data.

[0056] Step S420: Label the preprocessed sensor data to obtain several initial label boxes corresponding to the same target.

[0057] This step is the same as step S110 above, and will not be repeated here.

[0058] Step S430: Merge several initial annotation box information in the bird's-eye view coordinate system to obtain the merged annotation box information of the target.

[0059] This step is the same as step S120 above, and will not be repeated here.

[0060] Step S440: Based on the adjustment information of the fused bounding box information, update the initial bounding box information of the target in each preprocessed sensor data.

[0061] This step is the same as step S130 above, and will not be repeated here.

[0062] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0063] Please see Figure 5 , Figure 5 This is a schematic diagram of a framework of an embodiment of the data annotation device 500 of this application. The data annotation device 500 includes a target annotation module 510, a fusion module 520, and an adjustment module 530. The target annotation module 510 performs target annotation on raw sensor data collected by multiple sensor devices respectively, obtaining several initial annotation box information corresponding to the same target. The fusion module 520 performs fusion of the several initial annotation box information in a bird's-eye view coordinate system to obtain fused annotation box information of the target. The adjustment module 530 performs updates to the initial annotation box information of the target in each raw sensor data based on the adjustment information of the fused annotation box information.

[0064] In some implementations, the target annotation module 510 performs target annotation on the raw sensor data collected by multiple sensor devices to obtain several initial annotation box information corresponding to the same target. This includes: for the raw sensor data collected by each sensor device, performing key point identification on the raw sensor data to obtain several target key points corresponding to the target; performing bird's-eye view coordinate transformation on the several target key points to obtain target bird's-eye view projection points; generating a target area corresponding to the target in the bird's-eye view coordinate system based on the target bird's-eye view projection points; and mapping the target area to the sensor coordinate system corresponding to the raw sensor data to obtain the initial annotation box information of the target in the raw sensor data.

[0065] In some implementations, the target annotation module 510 executes a three-dimensional region as the target region, and the initial annotation box information includes the projection area of ​​the target in the original sensor data and the height information of the target.

[0066] In some implementations, the target annotation module 510 executes initial annotation box information as a three-dimensional rectangular box of the target in the original sensor data, the bottom surface of the three-dimensional rectangular box is the projection area, and the height of the three-dimensional rectangular box represents the height information of the target.

[0067] In some implementations, before performing bird's-eye view coordinate transformation on several target key points to obtain target bird's-eye view projection points, the target annotation module 510 includes: determining a first spatial coordinate transformation relationship between the coordinate system corresponding to each sensor device and the bird's-eye view coordinates; performing bird's-eye view coordinate transformation on several target key points to obtain target bird's-eye view projection points includes: transforming several target key points into the bird's-eye view coordinate system based on the first spatial coordinate transformation relationship to obtain target bird's-eye view projection points.

[0068] In some embodiments, before fusing several initial bounding box information in the bird's-eye view coordinate system to obtain the target's fused bounding box information, the fusion module 520 performs the following steps: determining a target sensor device and a sensor device to be converted among a plurality of sensor devices, wherein the sensor device to be converted is a device other than the target sensor device among the plurality of sensor devices; determining a third spatial coordinate transformation relationship based on the coordinate system corresponding to the target sensor device and the bird's-eye view coordinate system; and determining a second spatial coordinate transformation relationship based on the coordinate system corresponding to the sensor device to be converted and the coordinate system corresponding to the target sensor device, respectively.

[0069] In some implementations, the fusion module 520 performs the fusion of several initial bounding box information in the bird's-eye view coordinate system to obtain the first fused bounding box information of the target, including: based on a second spatial coordinate transformation relationship, transforming the initial bounding box information corresponding to the original sensor data collected by the sensor device to be transformed to the coordinate system corresponding to the target sensor device to obtain several first transformed bounding box information; based on a third spatial coordinate transformation relationship, transforming the several first transformed bounding box information and the initial bounding box information corresponding to the target sensor device to the bird's-eye view coordinate system to obtain several second transformed bounding box information; and fusing the several second transformed bounding box information to obtain the first fused bounding box information.

[0070] In some implementations, the adjustment module 530 performs adjustment information including adjustment data for the target's center position, size, and orientation within the fused annotation box information.

[0071] Please see Figure 6 , Figure 6 This is a schematic diagram of an embodiment of the electronic device 60 of this application. The electronic device 60 includes a memory 61 and a processor 62 coupled to each other. The processor 62 is used to execute program instructions stored in the memory 61 to implement the steps in any of the above-described data annotation method embodiments. In a specific implementation scenario, the electronic device 60 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 60 may also include mobile devices such as laptops and tablets, which are not limited here.

[0072] Specifically, processor 62 controls itself and memory 61 to implement the steps in any of the above-described data annotation method embodiments. Processor 62 can also be referred to as a CPU (Central Processing Unit). Processor 62 may be an integrated circuit chip with signal processing capabilities. Processor 62 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 62 can be implemented using integrated circuit chips.

[0073] Please see Figure 7 , Figure 7 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium 70 of this application. The computer-readable storage medium 70 stores program instructions 701 that can be executed by a processor. The program instructions 701 are used to implement the steps in any of the above-described data annotation method embodiments.

[0074] In some embodiments, the functions or modules of the apparatus provided in this disclosure can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.

[0075] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.

[0076] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0077] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0078] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A data annotation method, characterized in that, include: The raw sensor data collected by multiple sensor devices are labeled to obtain several initial label box information corresponding to the same target; The initial bounding box information is fused in the bird's-eye view coordinate system to obtain the fused bounding box information of the target; Based on the adjustment information of the fused bounding box information, the initial bounding box information of the target in each of the original sensor data is updated.

2. The method according to claim 1, characterized in that, The process of labeling raw sensor data collected by multiple sensor devices to obtain several initial bounding box information corresponding to the same target includes: For the raw sensor data collected by each of the sensor devices, key point identification is performed on the raw sensor data to obtain several target key points corresponding to the target; The bird's-eye view coordinates of the aforementioned key target points are transformed to obtain the target bird's-eye view projection points; Based on the target's bird's-eye view projection points, generate the target area corresponding to the target in the bird's-eye view coordinate system; The target region is mapped to the sensor coordinate system corresponding to the original sensor data to obtain the initial bounding box information of the target in the original sensor data.

3. The method according to claim 2, characterized in that, The target region is a three-dimensional region, and the initial annotation box information includes the projection area of ​​the target in the original sensor data and the height information of the target.

4. The method according to claim 3, characterized in that, The initial bounding box information is the 3D rectangular bounding box of the target in the original sensor data, the bottom surface of the 3D rectangular bounding box is the projection area, and the height of the 3D rectangular bounding box represents the height information of the target.

5. The method according to claim 2, characterized in that, Before performing bird's-eye view coordinate transformation on the aforementioned key target points to obtain the target bird's-eye view projection points, the process includes: Determine the first spatial coordinate transformation relationship between the coordinate system corresponding to each of the sensor devices and the coordinate system of the bird's-eye view; The step of performing a bird's-eye view coordinate transformation on the aforementioned key target points to obtain the target bird's-eye view projection points includes: Based on the first spatial coordinate transformation relationship, the several target key points are transformed into the bird's-eye view coordinate system to obtain the target bird's-eye view projection points.

6. The method according to claim 1, characterized in that, Before fusing the initial bounding box information in the bird's-eye view coordinate system to obtain the fused bounding box information of the target, the process includes: Identify a target sensor device and a sensor device to be converted among the plurality of sensor devices, wherein the sensor device to be converted is a device other than the target sensor device among the plurality of sensor devices; Based on the coordinate system corresponding to the target sensor device and the bird's-eye view coordinate system, determine the third spatial coordinate transformation relationship; The second spatial coordinate transformation relationship is determined based on the coordinate system corresponding to the sensor device to be transformed and the coordinate system corresponding to the target sensor device, respectively.

7. The method according to claim 6, characterized in that, The step of fusing the initial bounding box information in the bird's-eye view coordinate system to obtain the fused bounding box information of the target includes: Based on the second spatial coordinate transformation relationship, the initial annotation box information corresponding to the original sensor data collected by the sensor device to be transformed is transformed to the coordinate system corresponding to the target sensor device to obtain several first transformation annotation box information; Based on the third spatial coordinate transformation relationship, the plurality of first transformation annotation box information and the initial annotation box information corresponding to the target sensor device are transformed to the bird's-eye view coordinate system to obtain a plurality of second transformation annotation box information; The information of the several second transformation annotation boxes is merged to obtain the information of the first merged annotation box.

8. The method according to claim 1, characterized in that, The adjustment information includes adjustment data for the target's center position, size, and orientation within the fused annotation frame information.

9. An electronic device, characterized in that, It includes a memory and a processor coupled to each other, the processor being used to execute program instructions stored in the memory to implement the data annotation method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the data annotation method according to any one of claims 1 to 8.