Three-dimensional image labeling method and system, storage medium and electronic equipment

By acquiring and processing the camera calibration parameters and the conversion relationship between the two-dimensional image and the three-dimensional space in the scene image set, combining the two-dimensional frame coordinate information and object categories, adjusting the three-dimensional information of the target object to fit with the two-dimensional frame, solving the existing three-dimensional object detection and annotation methods that are cumbersome, time-consuming and cost-effective, and achieving efficient, low-cost and high-applicability three-dimensional image annotation.

CN120107973APending Publication Date: 2025-06-06GRG BANKING EQUIPMENT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510173173.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing three-dimensional object detection and annotation methods are cumbersome, time-consuming and cost-effective, and the model using simulation data is poorly applicable.

Method used

By obtaining the scene image set to be marked, determining the camera calibration parameters and the conversion relationship between the two-dimensional image and the three-dimensional space, obtaining the two-dimensional box coordinate information and object categories of the target object, adjusting the three-dimensional information of the target object to make it fit with the two-dimensional frame, and outputting the annotated scene image set.

Benefits of technology

It realizes three-dimensional image annotation that is efficient, simple, low-cost and high-profile, avoids the cumbersome and high-cost problems of using lidar point cloud data or multi-scene perspective data, and improves the applicability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107973A_ABST
    Figure CN120107973A_ABST
Patent Text Reader

Abstract

The invention discloses a three-dimensional image labeling method and device, a storage medium and electronic equipment. The method comprises the steps that a scene image set to be labeled is acquired; determining a camera calibration parameter corresponding to each scene according to the to-be-labeled scene image set, and obtaining a conversion relationship between a two-dimensional image and a three-dimensional space in each scene according to the camera calibration parameter; obtaining two-dimensional frame coordinate information and an object category of at least one target object in each scene image; adjusting the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relation, and obtaining a target three-dimensional frame attached to the two-dimensional frame of the target object based on the three-dimensional information; and obtaining and outputting a labeled scene image set based on each target three-dimensional frame. The labeling method is efficient, simple, low in cost and high in model applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image processing technology, and in particular relates to a three-dimensional image annotation method, device, storage medium and electronic device. Background Art

[0002] Three-dimensional object detection can obtain three-dimensional information such as the spatial position, three-dimensional size and direction angle of the target. Compared with two-dimensional object detection, it is conducive to improving the accuracy of fine-grained classification of similar two-dimensional feature targets. As an important part of three-dimensional object detection, three-dimensional annotation can provide the basis of training data for three-dimensional object detection algorithms.

[0003] However, most of the current annotations are done using LiDAR point cloud data or multi-scene view data, which is a cumbersome, time-consuming, and costly process. Although the use of simulated data can improve the annotation efficiency to a certain extent, there are certain differences between the data scene and the real scene, resulting in poor applicability of the model trained using 3D annotated data in simulated scenes. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems existing in the related art. To this end, the present application proposes a three-dimensional image annotation method, device, storage medium and electronic device, which can achieve the effects of high efficiency, simplicity, low cost and high model applicability.

[0005] In a first aspect, the present application provides a three-dimensional image annotation method, the method comprising:

[0006] Acquire a scene image set to be annotated, wherein each scene image in the scene image set to be annotated corresponds to a scene;

[0007] Determining camera calibration parameters corresponding to each of the scenes according to the scene image set to be annotated, and obtaining a conversion relationship between a two-dimensional image and a three-dimensional space in each of the scenes according to the camera calibration parameters;

[0008] Acquire two-dimensional frame coordinate information and object category of at least one target object in each of the scene images;

[0009] Adjust the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtain a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information;

[0010] A set of labeled scene images is obtained and output based on each of the target three-dimensional frames.

[0011] In a second aspect, the present application provides a three-dimensional image annotation device, the device comprising:

[0012] A data acquisition module, used to acquire a set of scene images to be annotated, wherein each scene image in the set of scene images to be annotated corresponds to a scene;

[0013] A camera calibration module, used to determine camera calibration parameters corresponding to each of the scenes according to the scene image set to be annotated, and obtain a conversion relationship between a two-dimensional image and a three-dimensional space in each of the scenes according to the camera calibration parameters;

[0014] A two-dimensional target detection module, used to obtain two-dimensional frame coordinate information and object category of at least one target object in each of the scene images;

[0015] A three-dimensional annotation module, configured to adjust the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtain a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information;

[0016] A result output module is used to obtain and output a marked scene image set based on each of the target three-dimensional boxes.

[0017] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the three-dimensional image annotation method as described in the first aspect above is implemented.

[0018] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the three-dimensional image annotation method as described in the first aspect above is implemented.

[0019] In a fifth aspect, the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the three-dimensional image annotation method as described in the first aspect.

[0020] In a sixth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the three-dimensional image annotation method as described in the first aspect above.

[0021] The above one or more technical solutions in the embodiments of the present application have at least the following technical effects:

[0022] The above-mentioned technical scheme of the present application obtains the camera calibration parameters corresponding to each scene and the conversion relationship between the two-dimensional image and the three-dimensional space by using the scene images of each scene in the scene image set to be annotated, obtains the two-dimensional frame coordinate information and object category of the marked target object in each scene image as an auxiliary for three-dimensional annotation, and then adjusts the three-dimensional information of the target object in each scene image based on the two-dimensional frame coordinate information, object category, camera calibration parameters and conversion relationship of each target object, so that the three-dimensional frame of the target object is as close to the two-dimensional frame as possible, obtains the target three-dimensional annotation information and outputs the annotated scene image set. This method does not require additional lidar point cloud data and simulation data as an auxiliary for annotation, and only uses the scene image under the perspective of a monocular camera to complete the annotation of the three-dimensional image. It not only avoids the problems of high cost, long time consumption and cumbersome annotation when using lidar point cloud data or multi-scene perspective data for annotation, but also avoids the problem of poor model adaptability caused by using simulation data. It can be applied to the three-dimensional annotation of target objects in scene images under the perspective of a variety of different roadside monocular cameras, and has good scene versatility.

[0023] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0025] Figure 1 This is one of the flowcharts of the three-dimensional image annotation method provided in the embodiment of the present application;

[0026] Figure 2 is a schematic diagram of a coordinate system and a camera model established according to an embodiment of the present application;

[0027] FIG3( a ) is a top view schematic diagram of a camera coordinate system provided in an embodiment of the present application;

[0028] FIG3( b ) is a schematic top view of a two-dimensional coordinate system provided in an embodiment of the present application;

[0029] Figure 4 is a schematic diagram of the structure of a two-dimensional target detection model provided in an embodiment of the present application;

[0030] Figure 5 This is the second flow chart of the three-dimensional image annotation method provided in the embodiment of the present application;

[0031] Figure 6 is a schematic diagram of the structure of a three-dimensional image annotation device provided in an embodiment of the present application;

[0032] Figure 7 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0034] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0035] The following, in conjunction with the accompanying drawings, describes in detail the three-dimensional image annotation method, three-dimensional image annotation device, electronic device, and readable storage medium provided by the embodiments of the present application through specific embodiments and their application scenarios.

[0036] The three-dimensional image annotation method may be applied to a terminal, and may be specifically executed by hardware or software in the terminal.

[0037] The three-dimensional image annotation method provided in the embodiment of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the three-dimensional image annotation method. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The three-dimensional image annotation method provided in the embodiment of the present application is described below using an electronic device as an example of the execution subject.

[0038] like Figure 1 As shown, the three-dimensional image annotation method includes the following steps:

[0039] S100, obtaining a scene image set to be labeled {I i , i=1,2,LN}, each scene image I in the scene image set to be labeled i Corresponding to a scene, where N represents the number of scene images and N is an integer greater than or equal to 1.

[0040] S200, determining camera calibration parameters corresponding to each of the scenes according to the set of scene images to be annotated, and obtaining a conversion relationship between a two-dimensional image and a three-dimensional space in each of the scenes according to the camera calibration parameters.

[0041] Specifically, step S200 includes: determining a scene image of a target scene from the scene image set to be annotated, extracting point coordinate information of a vanishing point in the scene image of the target scene, obtaining target camera calibration parameters of the target scene using the point coordinate information, road width information, and length information of a certain distance in the road, and obtaining a target conversion relationship between a two-dimensional image and a three-dimensional space in the target scene based on the target camera calibration parameters and a scale factor.

[0042] See also Figure 2 , establish the camera model, three-dimensional world coordinate system O w -X w Y w Z w , camera coordinate system O c -X c Y c Z c , two-dimensional coordinate system O i -uv, the camera model is simplified to a pinhole model, the coordinate systems are all right-handed, and the principal point r is located at the center of the image by default. The camera calibration parameters are set as follows: the focal length of the camera is f, the height of the camera origin from the ground is h, the camera pitch angle is φ, and the camera deflection angle (the angle between the projection of the camera optical axis on the road plane and the extension direction of the road) is θ. Since the camera spin angle ρ can be represented by a simple image rotation and has no effect on the calibration result, it is not considered.

[0043] Referring to Figure 3(a) and Figure 3(b), let the coordinate information of the vanishing point along the road be (u 0 ,v 0 ), l is the length information of a certain distance in the road, w is the width information of the road. It should be noted that the length information and width information of the road are both physical distances. There is a quartic equation for the unknown parameter f:

[0044]

[0045] Among them, the intermediate variable k is introduced for the convenience of calculation V =δτl / wv 0 , δ is the pixel distance corresponding to the road width information w on the image, τ=(v f -v 0 )(v b -v 0 ) / (v f -v b ), where v band v f They respectively represent the coordinate values ​​of the two endpoints of l corresponding to the v-axis in the two-dimensional coordinate system;

[0046]

[0047] The target camera calibration parameters of the target scene are obtained by the above formulas (1), (2) and (3). It can be understood that the target scene can be any of the above scenes.

[0048] Assume that the three-dimensional world coordinates of any point on the image are (x, y, z), introduce the scale factor α, and the projection relationship between the point in the three-dimensional world coordinate system and the point in the two-dimensional coordinate system is:

[0049]

[0050] According to the above formulas (4)-(7), the target conversion relationship between the two-dimensional image and the three-dimensional space in the target scene is finally obtained.

[0051] In addition, according to the above formulas (4)-(7), the transformation matrix H from three-dimensional coordinates to two-dimensional coordinates can be obtained:

[0052]

[0053] The specific derivation process of the above formulas (1)-(8) can be found in the paper "A Taxonomy and Analysis of Camera Calibration Methods for Traffic Monitoring Applications", which will not be described in detail here.

[0054] It should be noted that the reason why the above method is used to obtain the camera calibration parameters of the camera in each scene is that under normal circumstances, the roadside camera is of gimbal type, and there is a zoom problem for different scenes. Therefore, it is necessary to calculate the camera calibration parameters of each scene separately for different unknown scenes.

[0055] S300, obtaining two-dimensional frame coordinate information and object category of at least one target object in each of the scene images.

[0056] In some embodiments, the target object of the scene image set can be completed by manual annotation. Preferably, to improve efficiency, the target object can be completed by manual annotation. Figure 4 The 2D object detection model and manual annotations shown are obtained. Figure 4The two-dimensional target detection model shown can automatically output the two-dimensional box of the target object of the required object category, and supplement the two-dimensional annotation information of the target object that is not annotated by the two-dimensional target detection model through manual annotation, so as to ensure that the two-dimensional annotation information of each target object in the input scene image set is comprehensive, thereby improving the efficiency and accuracy of three-dimensional image annotation.

[0057] In some embodiments, step S300 specifically includes: using the annotation results output by the two-dimensional target detection model as the initial scene image set. If there are unlabeled target objects in the initial scene image set, adding corresponding two-dimensional frame coordinate information and object categories for the unlabeled target objects to obtain the two-dimensional frame coordinate information and object category of each target object. In this way, target objects with complete annotation information are efficiently obtained to better assist subsequent three-dimensional image annotation.

[0058] S400, adjusting the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtaining a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information.

[0059] In some embodiments, the two-dimensional frame coordinate information includes at least two-dimensional center point coordinates.

[0060] Specifically, in some embodiments, see Figure 5 , step S400 includes:

[0061] S410: Determine an initial three-dimensional size and an initial direction angle based on the object category to which the target object belongs.

[0062] S420, calculating the three-dimensional center point coordinates of the target object according to the two-dimensional center point coordinates of the target object, target camera calibration parameters of the target scene where the target object is located, and the corresponding target transformation relationship and initial three-dimensional size.

[0063] Specifically, in some embodiments, step S420 includes: combining the camera calibration parameters and the initial three-dimensional size, converting the two-dimensional center point coordinates of the target object into three-dimensional center point coordinates [x yz] according to formulas (6)-(7) in step S200. T The camera calibration parameters involved in formulas (6)-(7) are the target camera calibration parameters in the target scene, and u and v are the values ​​of the two-dimensional center point coordinates of the target object.

[0064] S430: Calculate an initial three-dimensional frame of the target object based on the three-dimensional center point coordinates, the initial three-dimensional size and the direction angle.

[0065] S440, adjusting the initial three-dimensional frame according to a preset rule to obtain a target three-dimensional frame that best fits the two-dimensional frame of the target object, wherein the three-dimensional information of the target three-dimensional frame includes at least target three-dimensional center point coordinates, target three-dimensional vertex coordinates and target direction angles.

[0066] In this embodiment, after determining the target object that needs to be three-dimensionally labeled, the approximate initial three-dimensional size and initial direction angle are first determined based on the object category to which the target object belongs. For example: if the object category is a car, the three-dimensional size of the car can be known. In some more specific embodiments, the target objects can also be classified according to the type of vehicle, such as cars, trucks, trucks, non-motor vehicles, etc., so that more accurate three-dimensional dimensions can be obtained. The three-dimensional dimensions of target objects of each object type can be stored in a database and can be retrieved and used according to the object type identifier when needed. The initial direction angle is usually the same as the direction of movement of the target object by default.

[0067] It should be noted that after obtaining the initial three-dimensional size and initial direction angle of the target object according to the object category and movement direction of the target object, in order to obtain more accurate three-dimensional size and direction angle, the three-dimensional frame of the target object will be continuously adjusted according to the preset rules so that it fits the two-dimensional frame as closely as possible. Specifically, the coordinates of the center point of the three-dimensional frame of the target object, the three-dimensional size and direction angle will be adjusted so that the adjusted three-dimensional frame fits the two-dimensional frame best, thereby obtaining the target three-dimensional information annotation data. It is understandable that in the process of adjusting to obtain the target three-dimensional information annotation data, the three-dimensional vertex coordinates of the three-dimensional frame of the target object will be continuously calculated and projected to the image space to obtain the minimum circumscribed rectangle, and then the degree of fit between it and the two-dimensional frame area is calculated. When the degree of fit is greater than the threshold, the obtained target three-dimensional center point coordinates, three-dimensional size, target direction angle, three-dimensional vertices and projection vertices combined with the target object category are the final three-dimensional information annotation data.

[0068] Specifically, in some embodiments, step S430 includes:

[0069] According to the target three-dimensional center point coordinates obtained in step S420 and the formula Calculate and obtain the three-dimensional vertex coordinates of the target object, where: R rot is the rotation matrix calculated by the direction angle rot, diag(l o ,w o ,h o ) represents the three-dimensional size (l o ,w o ,h o ) is a diagonal matrix, [xyz] Tis the coordinate of the three-dimensional center point, H is the transformation matrix, which is determined by formula (8). The parameter f in formula (8) is the focal length of the camera, h is the height of the camera origin from the ground, φ is the camera pitch angle, and P i 3D are the three-dimensional vertex coordinates of the three-dimensional box of the target object, The three-dimensional box vertex projection coordinates of the target object.

[0070] As can be seen from the above, the various parameters involved in the above formula (9) have been obtained from the above steps, thereby calculating the three-dimensional vertex coordinates of the target object to be three-dimensionally labeled and the projection coordinates of the three-dimensional vertex coordinates relative to the two-dimensional box.

[0071] Step S440 includes: during the adjustment process, firstly, by adjusting the three-dimensional center point coordinates, the initial three-dimensional size and the initial direction angle, the adjusted target three-dimensional frame can be obtained according to formula (9), and then based on formula Calculate the intersection and union ratio of the adjusted 3D box and 2D box, where B 3D Represents the minimum circumscribed rectangular area of ​​the 3D vertex coordinate projection of the 3D box, B 2D Represents a two-dimensional box area.

[0072] When IoU is greater than the threshold, the target 3D center point coordinates, 3D size, direction angle, 3D vertex and projection vertex are the final output 3D annotation information. The above information and object category are output together as an XML file to obtain a set of annotated scene images.

[0073] S500: Obtain and output a set of labeled scene images based on each of the target three-dimensional frames.

[0074] Furthermore, step S500 includes: outputting the target 3D center point coordinates, 3D size, target direction angle and 3D vertices and projection vertices, and object category of each target 3D box into an XML file to obtain a labeled scene image set.

[0075] The 3D coordinates of the center point, 3D size, target direction angle, 3D vertex and projection vertex, and object category of each 3D target object are output in XML format to form structured 3D image annotation data, which is convenient for subsequent efficient retrieval and application. It can be understood that the 3D size and 3D vertex and projection vertex can all be represented by the target 3D vertex coordinates.

[0076] The three-dimensional image annotation method provided in the embodiment of the present application first uses the scene images of each scene in the scene image set to be annotated to obtain the camera calibration parameters corresponding to each scene and the conversion relationship between the two-dimensional image and the three-dimensional space, and obtains the two-dimensional frame coordinate information and object category of the annotated target object in each scene image as an auxiliary for three-dimensional annotation, and then adjusts the three-dimensional information of the target object in each scene image based on the two-dimensional frame coordinate information, object category, camera calibration parameters and conversion relationship of each target object, so that the three-dimensional frame of the target object is as close to the two-dimensional frame as possible, obtains the target three-dimensional annotation information and outputs the annotated scene image set. The method does not require additional lidar point cloud data and simulation data as an annotation auxiliary, and only uses the scene image under the perspective of a monocular camera to complete the annotation of the three-dimensional image. It not only avoids the problems of high cost, long time consumption and cumbersome annotation when using lidar point cloud data or multi-scene perspective data for annotation, but also avoids the problem of poor model adaptability caused by using simulation data. The method can be applied to the three-dimensional annotation of target objects in scene images under the perspective of a variety of different roadside monocular cameras, and has good scene versatility.

[0077] The 3D image annotation method provided in the embodiment of the present application can be executed by a 3D image annotation device. In the embodiment of the present application, the 3D image annotation method performed by the 3D image annotation device is taken as an example to illustrate the 3D image annotation device provided in the embodiment of the present application.

[0078] The present application also provides a three-dimensional image annotation device. Figure 6 As shown, the three-dimensional image annotation device includes:

[0079] The data acquisition module 100 is used to acquire a set of scene images to be annotated, wherein each scene image in the set of scene images to be annotated corresponds to a scene. The camera calibration module 200 is used to determine the camera calibration parameters corresponding to each of the scenes according to the set of scene images to be annotated, and obtain the conversion relationship between the two-dimensional image and the three-dimensional space in each of the scenes according to the camera calibration parameters. The two-dimensional target detection module 300 is used to obtain the two-dimensional frame coordinate information and object category of at least one target object in each of the scene images. The three-dimensional annotation module 400 is used to adjust the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtain a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information. The result output module 500 is used to obtain and output the annotated scene image set based on each of the target three-dimensional frames.

[0080] According to the three-dimensional image annotation device provided by the embodiment of the present application, by using the scene images of each scene in the scene image set to be annotated, the camera calibration parameters corresponding to each scene and the conversion relationship between the two-dimensional image and the three-dimensional space are obtained, and the two-dimensional frame coordinate information and object category of the annotated target object in each scene image are obtained as an auxiliary for three-dimensional annotation, and then the three-dimensional information of the target object in each scene image is adjusted based on the two-dimensional frame coordinate information, object category, camera calibration parameters and conversion relationship of each target object, so that the three-dimensional frame of the target object is as close to the two-dimensional frame as possible, the target three-dimensional annotation information is obtained and the annotated scene image set is output. The method does not require additional lidar point cloud data and simulation data as an annotation auxiliary, and the annotation of the three-dimensional image can be completed only by using the scene image under the perspective of a monocular camera. It not only avoids the problems of high cost, long time consumption and cumbersome annotation when using lidar point cloud data or multi-scene perspective data for annotation, but also avoids the problem of poor model adaptability caused by using simulation data. It can be applied to the three-dimensional annotation of target objects in scene images under the perspective of a variety of different roadside monocular cameras, and has good scene versatility.

[0081] In some embodiments, the two-dimensional frame coordinate information includes at least two-dimensional center point coordinates. The three-dimensional annotation module 400 is specifically used to determine an initial three-dimensional size and an initial direction angle based on the object category to which the target object belongs;

[0082] The three-dimensional center point coordinates of the target object are calculated according to the two-dimensional center point coordinates of the target object, the target camera calibration parameters of the target scene where the target object is located, the corresponding target transformation relationship and the initial three-dimensional size; the initial three-dimensional frame of the target object is calculated based on the three-dimensional center point coordinates, the initial three-dimensional size and the initial direction angle; the initial three-dimensional frame is adjusted according to preset rules to obtain a target three-dimensional frame that best fits the two-dimensional frame of the target object, and the three-dimensional information of the target three-dimensional frame includes at least the target three-dimensional center point coordinates, the target three-dimensional vertex coordinates and the target direction angle.

[0083] In some embodiments, adjusting the initial three-dimensional frame according to a preset rule to obtain a target three-dimensional frame that best fits the two-dimensional frame of the target object includes:

[0084] Based on the formula The initial 3D frame is adjusted, and the 3D frame corresponding to the IoU closest to 1 is used as the target 3D frame, where IoU represents the intersection-over-union ratio of the 2D frame and the 3D frame, B 3D Represents the minimum circumscribed rectangular area of ​​the 3D vertex coordinate projection of the 3D box, B 2D Represents a two-dimensional box area.

[0085] In some embodiments, the two-dimensional frame coordinate information also includes the two-dimensional frame vertex coordinates. The three-dimensional annotation module 400 is further specifically configured to: The three-dimensional vertex coordinates of the target object are calculated, where: R rot is the rotation matrix calculated by the direction angle rot, diag(l o ,w o ,h o ) represents the three-dimensional size (l o ,w o ,h o ) is a diagonal matrix, [xyz] T is the coordinate of the three-dimensional center point, H is the transformation matrix, f is the focal length of the camera, h is the height of the camera origin from the ground, φ is the camera pitch angle, P i 3D are the three-dimensional vertex coordinates of the three-dimensional box, The vertex projection coordinates of the three-dimensional box of the target object.

[0086] In some embodiments, the result output module 500 is further specifically used to output the target 3D center point coordinates, the target 3D vertex coordinates, the target direction angle and the object category of each target 3D box as an XML file to obtain a labeled scene image set.

[0087] In some embodiments, the camera calibration module 200 is further specifically used to determine the scene image of the target scene from the set of scene images to be annotated, and extract the point coordinate information of the vanishing point in the scene image of the target scene; obtain the target camera calibration parameters of the target scene using the point coordinate information, the width information of the road and the length information of a certain distance in the road; and obtain the target conversion relationship between the two-dimensional image and the three-dimensional space in the target scene based on the target camera calibration parameters and the scale factor.

[0088] In some embodiments, the two-dimensional target detection module 300 is specifically used to use the labeling results output by the two-dimensional target detection model as the initial scene image set; if there are unlabeled target objects in the initial scene image set, the corresponding two-dimensional frame coordinate information and object category are added to the unlabeled target objects to obtain the two-dimensional frame coordinate information and object category of each of the target objects.

[0089] The three-dimensional image annotation device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0090] The 3D image annotation device in the embodiment of the present application may be a device having an operating system. The operating system may be a Microsoft (Windows) operating system, an Android (Android) operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0091] The three-dimensional image annotation device provided in the embodiment of the present application can achieve Figures 1 to 5 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0092] In some embodiments, Figure 7 As shown, an embodiment of the present application also provides an electronic device 800, including a processor 801, a memory 802, and a computer program stored in the memory 802 and executable on the processor 801. When the program is executed by the processor 801, each process of the above-mentioned three-dimensional image annotation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0093] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0094] The embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned three-dimensional image annotation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0095] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0096] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned three-dimensional image annotation method when executed by a processor.

[0097] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0098] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned three-dimensional image annotation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0099] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0100] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0102] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0103] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0104] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A three-dimensional image annotation method, characterized in that: include: Acquire a scene image set to be annotated, wherein each scene image in the scene image set to be annotated corresponds to a scene; Determining camera calibration parameters corresponding to each of the scenes according to the scene image set to be annotated, and obtaining a conversion relationship between a two-dimensional image and a three-dimensional space in each of the scenes according to the camera calibration parameters; Acquire two-dimensional frame coordinate information and object category of at least one target object in each of the scene images; Adjust the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtain a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information; A set of labeled scene images is obtained and output based on each of the target three-dimensional frames.

2. The method according to claim 1, characterized in that The two-dimensional frame coordinate information at least includes the two-dimensional center point coordinates; The method further comprises: adjusting the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameter, and the conversion relationship, and obtaining a target three-dimensional frame that matches the two-dimensional frame of the target object based on the three-dimensional information, including: Determine an initial three-dimensional size and an initial direction angle based on the object category to which the target object belongs; The three-dimensional center point coordinates of the target object are calculated according to the two-dimensional center point coordinates of the target object, the target camera calibration parameters of the target scene where the target object is located, the corresponding target transformation relationship and the initial three-dimensional size; Calculate the initial three-dimensional frame of the target object based on the three-dimensional center point coordinates, the initial three-dimensional size and the initial direction angle; The initial three-dimensional frame is adjusted according to preset rules to obtain a target three-dimensional frame that best fits the two-dimensional frame of the target object, wherein the three-dimensional information of the target three-dimensional frame includes at least the coordinates of the target three-dimensional center point, the coordinates of the target three-dimensional vertices and the target direction angle.

3. The method according to claim 2, characterized in that The initial three-dimensional frame is adjusted according to a preset rule to obtain a target three-dimensional frame that best fits the two-dimensional frame of the target object, including: Based on the formula The initial 3D frame is adjusted, and the 3D frame corresponding to the IoU closest to 1 is used as the target 3D frame, where IoU represents the intersection-over-union ratio of the 2D frame and the 3D frame, B 3D Represents the minimum circumscribed rectangular area of ​​the 3D vertex coordinate projection of the 3D box, B 2D Represents a two-dimensional box area.

4. The method according to claim 2, characterized in that: The two-dimensional frame coordinate information also includes the two-dimensional frame vertex coordinates; The initial three-dimensional frame of the target object is calculated based on the three-dimensional center point coordinates, the initial three-dimensional size and the initial direction angle, including: According to the formula The three-dimensional vertex coordinates of the target object are calculated, where: R rot is the rotation matrix calculated by the direction angle rot, diag(l o ,w o ,h o ) represents the three-dimensional size (l o ,w o ,h o ) is a diagonal matrix, [xyz] T is the coordinate of the three-dimensional center point, H is the transformation matrix, f is the focal length of the camera, h is the height of the camera origin from the ground, φ is the camera pitch angle, are the three-dimensional vertex coordinates of the three-dimensional box, The vertex projection coordinates of the three-dimensional box of the target object.

5. The method according to claim 4, characterized in that The step of obtaining and outputting a labeled scene image set based on each of the target three-dimensional frames includes: The target three-dimensional center point coordinates, the target three-dimensional vertex coordinates, the target direction angle and the object category of each target three-dimensional frame are output as an XML file to obtain a labeled scene image set.

6. The method according to any one of claims 1 to 5, characterized in that: The obtaining of two-dimensional frame coordinate information and object category of at least one target object in each of the scene images includes: The annotation results output by the two-dimensional object detection model are used as the initial scene image set; If there are unlabeled target objects in the initial scene image set, corresponding two-dimensional frame coordinate information and object category are added to the unlabeled target objects to obtain the two-dimensional frame coordinate information and object category of each target object.

7. A three-dimensional image annotation device, characterized in that: include: A data acquisition module, used to acquire a set of scene images to be annotated, wherein each scene image in the set of scene images to be annotated corresponds to a scene; A camera calibration module, used to determine camera calibration parameters corresponding to each of the scenes according to the scene image set to be annotated, and obtain a conversion relationship between a two-dimensional image and a three-dimensional space in each of the scenes according to the camera calibration parameters; A two-dimensional target detection module, used to obtain two-dimensional frame coordinate information and object category of at least one target object in each of the scene images; A three-dimensional annotation module, configured to adjust the three-dimensional information of the target object according to the two-dimensional frame coordinate information, the object category, the camera calibration parameters and the conversion relationship, and obtain a target three-dimensional frame that fits the two-dimensional frame of the target object based on the three-dimensional information; A result output module is used to obtain and output a marked scene image set based on each of the target three-dimensional boxes.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the three-dimensional image annotation method according to any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the three-dimensional image annotation method according to any one of claims 1 to 6 is implemented.

10. A chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, characterized in that: The processor is used to run a program or instruction to implement the three-dimensional image annotation method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Automatic labeling method and device, electronic equipment and storage medium

    CN121330425A

  • Image processing method and device, visual model training method and electronic equipment

    CN121482502A