A method and related device for converting a target 3D frame into a 2D frame
By extracting and mapping point cloud data to generate 2D boxes, the inaccuracy problem when converting 3D boxes into 2D boxes in the prior art is solved, the detection accuracy and fitting degree are improved, and the accuracy of 2D boxes is ensured.
Patent Information
- Application Number
- CN202211731828.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-12-30
AI Technical Summary
When converting a 3D frame into a 2D frame, the generated 2D frame often exceeds the actual range of the target object, resulting in a decrease in detection accuracy. Especially in the case of occlusion and shadowing, problems of missing marks and over-range marking are prone to occur.
By obtaining image data, original point cloud data, calibration data, and the target 3D box under the radar coordinate system, the target point cloud data is extracted, and the target point cloud data is mapped to the image space based on the calibration data, and 4 most value points are extracted to generate a 2D box.
The fit and accuracy of 2D frame generation is improved, ensuring that the generated 2D frame is close to the physical limit, reducing the situation of missing marks and over-range annotations, and improving the stability and accuracy of subsequent processing.
Smart Images

Figure CN115984097B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data fusion, and in particular to a method and a related device for converting a target 3D frame into a 2D frame. Background Art
[0002] Autonomous driving / assisted driving is a mainstream application in the field of artificial intelligence. Motor vehicles perceive environmental information through sensors such as cameras, radars, and monitoring devices. An important part of perception is to detect targets of interest in the scene (such as vehicles, pedestrians, etc.), so it is necessary to mark the targets with reasonable bounding boxes.
[0003] At present, the main method used to complete the above tasks is BEV-based neural network learning. Compared with traditional target detection under 2D boxes, target detection under BEV introduces the characteristics of spatial perception. In the current BEV neural network training, 3D boxes are often used as the supervision target of learning, but in recent studies, 2D boxes corresponding to 3D boxes have been introduced to assist in learning supervision, thereby achieving better results. When converting 3D boxes to 2D boxes, the converted 2D boxes may be inaccurate. The 2D boxes converted using existing methods are often significantly larger than the target object itself, which affects the stability of the deep learning network. Therefore, the demand for "converting 3D boxes into 2D boxes" in various scenarios has arisen.
[0004] Regarding the conversion of 3D boxes into 2D boxes, the main approach in the prior art is to directly map the 3D box to the 2D image plane, and then take the maximum value of the 4 directions of the 3D bbox box in the pixel coordinate system to generate the 2D box. However, the generated 2D box obviously exceeds the range of the actual 2D box of the target, and there are omissions and out-of-range annotations for occluded objects (partial occlusion and shadow occlusion) (such as Figure 1 When there is a large error between the 2D box and the actual target, it will seriously affect the post-processing results (such as network learning and training). Summary of the invention
[0005] The technical problem to be solved by the present application is to provide a method and a related device for converting a target 3D frame into a 2D frame in view of the deficiencies of the prior art.
[0006] In order to solve the above technical problem, a first aspect of an embodiment of the present application provides a method for converting a target 3D frame into a 2D frame, the method comprising:
[0007] Obtain image data, original point cloud data, calibration data, and the 3D frame of the target in the radar coordinate system;
[0008] Extracting point cloud data belonging to the target from the original point cloud data to generate target point cloud data, wherein the target point cloud data does not contain discrete points;
[0009] Based on the calibration data, mapping the target point cloud data to an image space corresponding to the image data to obtain mapped point cloud data;
[0010] Extracting four extreme points of the mapped point cloud data, wherein the four extreme points are respectively the leftmost, rightmost, bottommost and topmost boundary points of the target point cloud data mapped to the image space;
[0011] A 2D box of the target is generated based on the four maximum points.
[0012] The method for converting the 3D frame into a 2D frame, wherein after obtaining the 3D frame, the method further comprises:
[0013] Selecting a center point of the 3D frame and parameters of the 3D frame, wherein the parameters include length information, width information, and height information of the 3D frame;
[0014] The coordinate information of the eight vertices of the 3D frame is obtained according to the center point and the parameters.
[0015] The method according to claim 2, characterized in that the step of extracting point cloud data belonging to the target from the original point cloud data to generate target point cloud data comprises:
[0016] Generate 4 mapping vertices based on the BEV perspective from the 8 vertices of the 3D box;
[0017] Calculate the angle sum between each of the original point cloud data and the four mapping vertices based on the horizontal axis coordinate information and the vertical axis coordinate information;
[0018] If the angle sum is equal to 360°, the original point cloud data is set as the target candidate point cloud data;
[0019] A vertical axis coordinate threshold is set according to the vertical axis coordinate information of the eight vertices, and the target candidate point cloud data is screened based on the vertical axis coordinate threshold to generate target point cloud data.
[0020] The method for converting the 3D frame into a 2D frame, wherein the calibration data includes joint calibration data and camera calibration data, and mapping the target point cloud data to an image space corresponding to the image data based on the calibration data to obtain mapped point cloud data, comprises:
[0021] Convert the target point cloud data in the radar coordinate system to the camera coordinate system according to the joint calibration data;
[0022] The target point cloud data in the camera coordinate system is converted to the image coordinate system according to the camera calibration data.
[0023] The method for converting the 3D frame into a 2D frame, wherein before mapping the target point cloud data to the image space corresponding to the image data based on the calibration data to obtain the mapped point cloud data, the method further comprises:
[0024] If the target point cloud data is lower than a preset condition, correcting the 3D frame;
[0025] The vertex information of the modified 3D box is obtained, and the vertices of the modified 3D box are used as the target point cloud data.
[0026] The method for converting a 3D frame into a 2D frame, wherein the preset threshold is set according to the target point cloud density and / or the target distance.
[0027] The method for converting the 3D frame into a 2D frame, wherein, if the target is a motor vehicle type target, correction is performed according to the motor vehicle parameters, and the motor vehicle parameters at least include rearview mirror parameters and vehicle type parameters of the motor vehicle.
[0028] A second aspect of an embodiment of the present application provides a device for converting a target 3D frame into a 2D frame, the device comprising:
[0029] The acquisition module is used to obtain image data, original point cloud data, calibration data and the 3D frame of the target in the radar coordinate system.
[0030] A generation module is used to generate target point cloud data by extracting original point cloud data belonging to the target, wherein if the target point cloud data is lower than a preset condition, the 3D box of the target is corrected, and the vertices of the corrected 3D box are used as the target point cloud data.
[0031] A conversion module is used to convert the 3D box of the target in the radar coordinate system into the 2D box of the target in the image coordinate system, wherein the target point cloud data obtained by the generation module is mapped to the image space corresponding to the image data, and the 2D box is generated based on 4 extreme points.
[0032] A third aspect of an embodiment of the present application provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in any of the methods for converting a target 3D frame into a 2D frame as described above.
[0033] A fourth aspect of the embodiments of the present application provides a terminal device, comprising: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;
[0034] The communication bus realizes the connection and communication between the processor and the memory;
[0035] When the processor executes the computer-readable program, the processor implements the steps in any of the above methods for converting a target 3D frame into a 2D frame.
[0036] Beneficial effects: Compared with the prior art, this application utilizes the discrete characteristics of point clouds and uses the point clouds themselves to connect 3D frames and 2D frames, thus getting rid of the inconsistency caused by directly using the mapping relationship to change the 3D frame into a 2D frame (the converted 2D frame is significantly larger than the target object), and improving the fit of the 2D frame generation, thereby improving the accuracy of converting the target 3D frame into a 2D frame based on point cloud fusion, close to the physical limit. The generated 2D frame can be applied to various subsequent tasks (such as target detection in assisted driving scenarios, etc.) to improve the accuracy of post-processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other structural principle diagrams can be obtained based on these drawings without inventive work.
[0038] Figure 1 Schematic diagram of the effect of converting a 3D box into a 2D box.
[0039] Figure 2 A flow chart of an embodiment provided for this application.
[0040] Figure 3 Another embodiment flow chart provided for this application.
[0041] Figure 4 This is a schematic diagram of the structure of the device for converting a target 3D frame into a 2D frame provided in this application.
[0042] Figure 5 This is a schematic diagram of the structure of the terminal device provided in this application. DETAILED DESCRIPTION
[0043] The present application provides a method and a related device for converting a target 3D frame into a 2D frame. In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0045] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0046] The following is an explanation of some of the terms used in this application:
[0047] 2D box: The "2D box" mentioned in this application refers to a two-dimensional annotation box used to mark targets on image data, where the "box" in "2D box" should be understood as the English "bonding box", which can be interpreted as an annotation box, bounding box or detection box.
[0048] 3D box: The "3D box" mentioned in this application refers to a three-dimensional annotation box used to annotate targets on point cloud data, where the "box" of "3D box" should be understood as the English "bonding box", which can be interpreted as an annotation box, bounding box or detection box.
[0049] Point cloud fusion: The "point cloud fusion" mentioned in this application refers to the fusion of point cloud and image, which can be broadly understood as associating point cloud data with image data through a mapping relationship.
[0050] Joint calibration: refers to the calibration of the external parameters from the LiDAR coordinate system to the camera coordinate system. After the camera and laser are jointly calibrated, the LiDAR measurement value can be accurately projected into the camera image, thereby realizing the association between the laser point and the three-channel color information. Conversely, the pixel in the camera image can obtain the depth value by querying the nearest laser.
[0051] BEV: It is the abbreviation of Bird's Eye View, also known as God's perspective. It is a perspective or coordinate system (3D) used to describe the perceived world. BEV is also used to refer to an end-to-end technology in the field of computer vision that uses a neural network to convert visual information from image space to BEV space.
[0052] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0053] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.
[0054] This embodiment provides a method for converting a target 3D frame into a 2D frame, such as Figure 2 As shown, the method includes:
[0055] S10, acquiring image data, original point cloud data, calibration data, and a 3D frame of the target in a radar coordinate system.
[0056] Specifically, the above 4 items are all the input items required by this method. Among them, since this method is based on point cloud fusion technology, the image data and original point cloud data are required; the calibration data includes joint calibration data and camera internal and external parameter data, and the joint calibration data is the coordinate transformation relationship between the radar coordinate system and the camera coordinate system, which can be represented by a matrix; the 3D frame of the target in the radar coordinate system (hereinafter referred to as "3D frame") only contains the point cloud data of the target, which is included in the original point cloud data;
[0057] Furthermore, in order to accurately obtain the point cloud data of the target in the original point cloud data, in one implementation, the vertex information of the 3D box needs to be accurately obtained. For a 3D box, there are 8 vertices, so the coordinate information of the 8 vertices needs to be obtained.
[0058] Based on this, in one implementation, the method for obtaining 3D frame vertex information specifically includes:
[0059] S11, selecting a center point of the 3D frame and parameters of the 3D frame, wherein the parameters include length information, width information, and height information of the 3D frame;
[0060] S12. Obtain coordinate information of eight vertices of the 3D frame according to the center point and the parameters.
[0061] Specifically, after obtaining the center point of the 3D frame and the length information, width information and height information of the 3D frame, the eight vertices of the 3D frame can be obtained by multiplying the two. At this time, the coordinate information of the eight vertices in the radar coordinate system can be obtained by (x l ,y l , z l ) is used to represent the horizontal and vertical coordinate information, the vertical axis coordinate information, and the vertical axis coordinate threshold of the vertex. The vertex coordinate information can be used to extract the point cloud data of the target.
[0062] It is worth noting that different coordinate systems can have different coordinate expressions. For example, in the camera coordinate system, (x c ,y c , z c ), in image coordinates it is (u, v), and so on, without limitation here.
[0063] S20, extracting point cloud data belonging to the target from the original point cloud data to generate target point cloud data, wherein the target point cloud data does not contain discrete points.
[0064] Specifically, based on the target as the classification basis, the several point cloud points existing in the original point cloud can be divided into target point cloud data and non-target point cloud data. The ultimate goal of the embodiment of the present application is to obtain a 2D box that is highly fitted to the target. The 2D box should be highly fitted to the physical limit, that is, highly fitted to the target. Therefore, it is necessary to accurately extract the point cloud data belonging to the target in the original point cloud. Since the point cloud data at this time is in the radar coordinate system, that is, (x l ,y l , z l ), and the 2D box is the image data in the image coordinate system, so it is impossible to accurately obtain the target point cloud data directly by limiting the coordinate threshold of the point cloud in the radar coordinate system.
[0065] In one embodiment, the method for generating target point cloud data specifically includes:
[0066] S21 generates four mapping vertices from the eight vertices of the 3D box based on the BEV perspective;
[0067] S22 calculates the angle sum between each of the original point clouds and the four mapping vertices based on the horizontal axis coordinate information and the vertical axis coordinate information;
[0068] S23: if the angle sum is equal to 360°, setting the point cloud as target candidate point cloud data;
[0069] S24 sets a vertical axis coordinate threshold according to the vertical axis coordinate information of the eight vertices, and screens the target candidate point cloud data based on the threshold to generate target point cloud data.
[0070] After obtaining the 8 vertices of the 3D box, the 8 vertices are mapped based on the BEV perspective. The BEV perspective is a bird's-eye view, which can be understood as looking at the 8 point vertices from a top-down angle. The original 8 vertices will be changed to 4 mapped vertices.
[0071] In one embodiment, firstly only consider (x, y) projected under the BEV perspective, that is, do not consider the height information (vertical axis coordinate information z) to calculate the sum of the angles of each point cloud and the four vertices. If the sum of the four angles (that is, the "angle sum") is 360°, then the point is an inner point (that is, the "target candidate point cloud data" of claim 3); if the sum of the four angles (that is, the "angle sum") is not 360°, then the point is an outer point (that is, not the "target candidate point cloud data" of claim 3). The calculation result is a set of points distributed in the rectangle surrounded by the above four points with no height restrictions under the bev perspective, and then by setting a height threshold, the points on the target object can be screened out. In one embodiment, the method of setting a threshold for height can find the minimum value z of the vertical axis z based on the 8 vertex positions mentioned above. min and the maximum value z max , select the internal points that satisfy "z min ≤z≤z max " is the target point cloud data. For example, assuming that the original point cloud data has 100,000 point clouds, there are 3,000 point clouds that meet the angle sum of 360° under the BEV perspective, and the minimum vertical axis z of the 3D box is min and the maximum value z max If they are 1 and 10 respectively, then the z of the 3000 point clouds is limited to [1,10]. Assuming that 2500 point clouds meet this requirement, the number of point cloud data of the target object is 2500.
[0072] In addition, in order to obtain the final target point cloud data, the target object point cloud data has been pre-processed to remove discrete points in the point cloud data on the target object. This is to improve the accuracy of the subsequent 2D frame. Continuing with the example in the previous paragraph, assuming that the target object point cloud data is 2500, and 200 point clouds are removed during the preprocessing (noise reduction processing), the final target point cloud data is 2300.
[0073] S30. Based on the calibration data, map the target point cloud data to an image space corresponding to the image data to obtain mapped point cloud data.
[0074] Specifically, the step of mapping the target point cloud data to the image space is to first convert the target point cloud data in the radar coordinate system to the camera coordinate system according to the joint calibration data, and then convert the target point cloud data in the camera coordinate system to the image coordinate system according to the camera calibration data. The camera calibration data includes camera intrinsic parameters and camera extrinsic parameters. The function of the camera intrinsic parameters is to convert from the camera coordinate system to the pixel coordinate system, and the function of the camera extrinsic parameters is to convert from the world coordinate system to the camera coordinate system.
[0075] In one embodiment, if the target point cloud data is lower than a preset threshold, the 3D frame is corrected, specifically including:
[0076] S31 obtains vertex information of the modified 3D frame;
[0077] S32 replaces the target point cloud data with the corrected 3D box vertices;
[0078] S33 maps the corrected 3D box vertices to the image space based on the calibration data.
[0079] In one embodiment, the preset threshold is set according to the target point cloud density and / or the target distance. In one embodiment, a preliminary limitation is preferentially made according to the (Euclidean) distance. Assuming that the preliminary limitation value is 25m, if the target distance>25m, the 3D box is directly corrected; if the target distance is <25m, first determine whether the target point cloud density (of the target point cloud data) is reliable. If the point cloud density is lower than the reliable value, the 3D box is corrected. If the point cloud density is greater than or equal to the reliable value, no correction is required. The reliability value can be set according to different tasks and different scenarios, and can also be set by empirical values. The setting method is not limited.
[0080] The preset threshold value can be set according to the target point cloud density and the target distance. The target point cloud density can be used to preliminarily limit the distance, or the preset threshold value can be set only by the density of the target point cloud data. There is no restriction on the setting steps and combinations of the preset threshold values, which will not be listed here.
[0081] Furthermore, in one embodiment, Figure 3 , set the preset conditions according to the target distance / target point cloud density. When the conditions are met, choose to directly map the target point cloud data to the image space. When the conditions are not met, choose to correct the 3D frame before mapping. It is worth noting that in another embodiment, when the conditions are not met, you can also choose to directly map the target point cloud data to the image space. However, the effect of the final 2D frame (such as accuracy) will be reduced. Therefore, whether the preset conditions are met is not a necessary item. In order to pursue a high-precision 2D frame, you can choose to perform classification processing under the conditions that meet the preset conditions and under the conditions that do not meet the preset conditions. Of course, you can also choose other processing methods. This is hereby explained.
[0082] In one embodiment, such as a vehicle driving scene, for the vehicle itself, the target may be a pedestrian, a motor vehicle, a traffic sign, etc. Since the motor vehicle is one of the main target types, it is hereby explained how to correct the 3D frame when the motor vehicle is the target.
[0083] In one embodiment, if the target is a motor vehicle type target, the correction is performed according to the motor vehicle parameters, and the motor vehicle parameters at least include the rearview mirror parameters and vehicle model parameters of the motor vehicle. Specifically, motor vehicle models generally have two left and right outer rearview mirrors and streamlines of the front and rear vehicles, which makes it easy for the 2D box of the motor vehicle to be significantly larger than the actual range when it is mapped to the image space.
[0084] In one embodiment, the specific method includes: cropping the width of the 3D frame according to the rearview mirror of the car; cropping the width of the 3D frame according to the car model; and reprojecting the corrected point cloud.
[0085] In another embodiment, the specific method includes: performing an initial correction based on the correction condition that the correction amplitude of the front of the vehicle is greater than the correction amplitude of the rear of the vehicle to obtain the left extreme point and the right extreme point of the corrected 3D frame; selecting the lowest point and the highest point based on the camera and the wheelbase from point to camera; and correcting the 3D frame according to the left extreme point, the right extreme point, the lowest point and the highest point.
[0086] Since the width of the front and rear of the vehicle is generally smaller than the width of the vehicle body, the error can be corrected using empirical values when projected in the camera coordinate system. Since the width of the vehicle's rearview mirror is mainly increased, and most of the front of the vehicle will have additional streamlined designs, the correction of the front of the vehicle is greater than the correction of the rear of the vehicle, and the leftmost and rightmost points are obtained; in the height direction, the correction of the lowest point in the axis closest to the camera is the lowest point, and the highest value of the axis away from the vehicle is the highest point. Finally, the camera intrinsic parameters are used to reproject back to the image coordinate system and the 2D box limit values in four directions.
[0087] S40, extracting four extreme points of the mapped point cloud data, wherein the four extreme points are respectively the leftmost, rightmost, bottommost and topmost boundary points of the target point cloud data mapped to the image space.
[0088] Specifically, the four extreme points are the leftmost, rightmost, bottommost and topmost boundary points of the target point cloud data mapped to the image space. Further, if the target point cloud data is lower than the preset threshold, the mapping result is that the 8 vertices in the radar coordinate system of the modified 3D frame are first mapped to the camera coordinate system, and then mapped to the 8 vertices in the image coordinate system in the image space. The coordinate relationship transformation of the vertices in this process can be expressed as (x l ,y l , z l)→(x c ,y c , z c )→(u, v), therefore, finally select the maximum points in four directions (up, down, left, and right) according to (u, v);
[0089] If the target point cloud data reaches the preset threshold, the mapping result is that the target point cloud data is finally mapped to several target point clouds in the image coordinate system in the image space. The mapping process is the same as above and will not be repeated here. Finally, the maximum points in four directions (up, down, left, and right) are selected according to (u, v).
[0090] S50: Generate a 2D box of the target based on the four maximum points.
[0091] Specifically, in the image space, when the maximum points in four directions are obtained and used as four boundary points, a 2D box can be formed. Figure 5 As shown in FIG. 1 , the effect of the method for converting the target 3D frame into a 2D frame provided by the application is that the target is completely fitted; for example, Figure 1 The effect diagram of the prior art is shown in the figure. By comparison, it can be clearly seen that the method can accurately convert the target 3D frame into a 2D frame without any occlusion or excessive standard range. Figure 1 As shown, frame A is the effect after transformation by this method, and frame B is the effect of general transformation by the prior art. It can be clearly seen that frame A is close to the physical limit, while frame B is obviously not in line with reality, and the effect of the transformation has been significantly improved.
[0092] In summary, the present application discloses a method and related device for converting a target 3D frame into a 2D frame, the method comprising acquiring image data, original point cloud data, calibration data, and a 3D frame of the target in a radar coordinate system, extracting the point cloud data of the target in the original point cloud data, projecting the target point cloud data into the image space according to the calibration data, and generating a 2D frame based on the four extreme points in the projection result. The present application utilizes the discrete characteristics of the point cloud and uses the point cloud itself to connect the 3D frame and the 2D frame, thereby getting rid of the inconsistency caused by directly using a mapping relationship to change the 3D frame into a 2D frame (the converted 2D frame is significantly larger than the target object), and improving the fitting degree of the 2D frame generation, thereby improving the accuracy of converting the target 3D frame into a 2D frame based on point cloud fusion, which is close to the physical limit.
[0093] Based on the above method for converting a target 3D frame into a 2D frame, this embodiment provides a device for converting a target 3D frame into a 2D frame, such as Figure 4 As shown, the system comprises:
[0094] The acquisition module is used to obtain image data, original point cloud data, calibration data and the 3D frame of the target in the radar coordinate system.
[0095] A generation module is used to generate target point cloud data by extracting original point cloud data belonging to the target, wherein if the target point cloud data is lower than a preset threshold, the 3D frame of the target is corrected to generate a corrected 3D frame.
[0096] The conversion module is used to convert the 3D frame of the target in the radar coordinate system into the 2D frame of the target in the image coordinate system, wherein the target point cloud data obtained by the generation module or the corrected 3D frame vertices are mapped to the image space, and the 2D frame is generated based on the four extreme points.
[0097] Based on the above-mentioned method of converting a target 3D frame into a 2D frame, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the method of converting a target 3D frame into a 2D frame as described in the above-mentioned embodiment.
[0098] Based on the above method of converting the target 3D frame into a 2D frame, the present application also provides a terminal device, such as Figure 5 As shown, it includes at least one processor 20; a display screen 21; and a memory 22, and may also include a communications interface 23 and a bus 24. The processor 20, the display screen 21, the memory 22, and the communications interface 23 may communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communications interface 23 may transmit information. The processor 20 may call the logic instructions in the memory 22 to execute the method in the above embodiment.
[0099] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0100] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.
[0101] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.
[0102] In addition, the specific process of loading and executing the multiple instruction processors in the above storage medium and the terminal device has been described in detail in the above method and will not be described here one by one.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for converting a target 3D bounding box into a 2D bounding box, characterized in that, the method includes: obtaining image data, original point cloud data, calibration data, and the 3D bounding box of the target in the radar coordinate system; extracting the point cloud data belonging to the target from the original point cloud data to generate target point cloud data, and the target point cloud data does not include discrete points; mapping the target point cloud data to the image space corresponding to the image data based on the calibration data to obtain mapped point cloud data; extracting 4 extreme value points of the mapped point cloud data, and the 4 extreme value points are respectively the leftmost, rightmost, bottommost, and uppermost 4 boundary points of the target point cloud data mapped to the image space; generating a 2D bounding box of the target based on the 4 extreme value points; the extracting the point cloud data belonging to the target from the original point cloud data to generate target point cloud data includes: generating 4 mapped vertices of the 8 vertices of the 3D bounding box based on the BEV perspective; calculating the sum of the angles between each piece of the original point cloud data and the 4 mapped vertices based on the horizontal axis coordinate information and the vertical axis coordinate information; if the sum of the angles is equal to 360°, setting this piece of the original point cloud data as target candidate point cloud data; setting a vertical axis coordinate threshold according to the vertical axis coordinate information of the 8 vertices, and screening the target candidate point cloud data based on the vertical axis coordinate threshold to generate target point cloud data.
2. The method according to claim 1, characterized in that, after obtaining the 3D bounding box, it further includes: selecting the center point of the 3D bounding box and the parameters of the 3D bounding box, and the parameters include the length information, width information, and height information of the 3D bounding box; obtaining the coordinate information of the 8 vertices of the 3D bounding box according to the center point and the parameters.
3. The method according to claim 1, characterized in that, the calibration data includes joint calibration and camera calibration data, and the mapping the target point cloud data to the image space corresponding to the image data based on the calibration data to obtain mapped point cloud data includes: converting the target point cloud data in the radar coordinate system to the camera coordinate system according to the joint calibration data; converting the target point cloud data in the camera coordinate system to the image coordinate system according to the camera calibration data.
4. The method according to claim 1, characterized in that, before the mapping the target point cloud data to the image space corresponding to the image data based on the calibration data to obtain mapped point cloud data, the method further includes: if the target point cloud data is lower than a preset condition, correcting the 3D bounding box; obtaining the vertex information of the corrected 3D bounding box, and using the vertices of the corrected 3D bounding box as the target point cloud data.
5. The method according to claim 4, characterized in that, the preset condition is set according to the target point cloud density and / or the target distance.
6. The method according to claim 4, characterized in that, if the target is a motor vehicle type target, correcting according to the motor vehicle parameters, and the motor vehicle parameters at least include the rearview mirror parameters and the vehicle type parameters of the motor vehicle.
7. An apparatus for converting a target 3D bounding box into a 2D bounding box, It is characterized in that The device comprises: The acquisition module is used to obtain image data, original point cloud data, calibration data and the 3D frame of the target in the radar coordinate system; A generating module, configured to generate target point cloud data by extracting original point cloud data belonging to the target, wherein if the target point cloud data is lower than a preset condition, the 3D box of the target is corrected, and the vertices of the corrected 3D box are used as the target point cloud data; A conversion module, used for converting a 3D frame of the target in the radar coordinate system into a 2D frame of the target in the image coordinate system, wherein the target point cloud data obtained by the generation module is mapped to the image space corresponding to the image data, and the 2D frame is generated based on four maximum points; The generating module comprises: Generate 4 mapping vertices based on the BEV perspective from the 8 vertices of the 3D box; Calculate the angle sum between each of the original point cloud data and the four mapping vertices based on the horizontal axis coordinate information and the vertical axis coordinate information; If the angle sum is equal to 360°, the original point cloud data is set as the target candidate point cloud data; A vertical axis coordinate threshold is set according to the vertical axis coordinate information of the eight vertices, and the target candidate point cloud data is screened based on the vertical axis coordinate threshold to generate target point cloud data.
8. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for converting a target 3D frame into a 2D frame as claimed in any one of claims 1 to 6.
9. A terminal device, It is characterized in that include: Processor, memory and communication bus; The memory stores a computer-readable program that can be executed by the processor; the communication bus realizes the connection and communication between the processor and the memory; when the processor executes the computer-readable program, it implements the steps in the method for converting a target 3D frame into a 2D frame as described in any one of claims 1-6.
Citation Information
Patent Citations
Target detection method and electronic equipment
CN113903028A
Target detection method and device, roadside base station and storage medium
CN114332579A