Monocular 3D detection method, device, electronic device and storage medium

Through projection and iterative adjustment technology, the existing monocular 3D detection methods have been solved in terms of stability and accuracy, and more efficient optimization of 3D detection results and improvement of accuracy are achieved.

CN113887290BActive Publication Date: 2025-05-13JILUO TECH (SHANGHAI) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111013236.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-05-13
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

The existing monocular 3D detection methods have shortcomings in terms of stability and accuracy, and often have detection results with errors or large errors.

Method used

By obtaining the 2D detection results and 3D detection results of the target in the picture to be detected, the 3D detection results are projected on the picture to be detected, the 2D detection results after projection are obtained, and the 3D detection results are iteratively adjusted based on the 2D detection results and the 2D detection results after projection, and the adjusted 3D detection results are obtained.

Benefits of technology

Effectively optimize the output 3D detection results, improve the rationality and stability of 3D detection results, and improve the overall accuracy of monocular 3D detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887290B_ABST
    Figure CN113887290B_ABST
Patent Text Reader

Abstract

The present invention provides a monocular 3D detection method, device, electronic device and storage medium. The method comprises: obtaining a 2D detection result and a 3D detection result of a target in a picture to be detected; projecting the 3D detection result onto the picture to be detected to obtain a projected 2D detection result; iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result; and performing 2D frame annotation and 3D frame annotation on the target on the picture to be detected according to the 2D detection result and the adjusted 3D detection result, so as to achieve effective optimization of the output 3D detection result, improve the rationality and stability of the 3D detection result, and improve the overall accuracy of monocular 3D detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a monocular 3D detection method, system, electronic equipment and storage medium. Background Art

[0002] The existing mainstream monocular 3D detection method is an end-to-end solution represented by CenterNet, which completes the 2D target detection task on the image while regressing the 3D properties of the target and finally outputs the 3D detection result.

[0003] However, the stability and accuracy of 3D detection results are required to be high. The current network model cannot guarantee the lower limit of the detection results, and some errors or large errors in detection results often occur. Summary of the invention

[0004] In view of the problems existing in the prior art, the present invention provides a monocular 3D detection method, system, electronic device and storage medium.

[0005] In a first aspect, the present invention provides a monocular 3D detection method, comprising:

[0006] Obtain 2D detection results and 3D detection results of the target in the image to be detected;

[0007] Projecting the 3D detection result onto the image to be detected to obtain a projected 2D detection result;

[0008] Iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result;

[0009] The target on the to-be-detected screen is labeled with a 2D frame and a 3D frame according to the 2D detection result and the adjusted 3D detection result.

[0010] In one embodiment, iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result includes:

[0011] Determine loss information according to the 2D detection result and the projected 2D detection result;

[0012] The target parameters in the 3D detection result are adjusted according to the loss information and the preset iteration step size and termination step size to obtain an adjusted 3D detection result.

[0013] In one embodiment, determining loss information according to the 2D detection result and the projected 2D detection result includes:

[0014] Determine a 2D bounding box and a projected 2D bounding box based on the 2D detection result and the projected 2D detection result;

[0015] Determine a loss value based on the coordinate points of the 2D bounding box and the coordinate points of the projected 2D bounding box.

[0016] In one embodiment, the target parameters in the 3D detection result include a depth value and an orientation angle. Correspondingly, adjust the target parameters in the 3D detection result according to the loss information, a preset iteration step size, and a termination step size to obtain an adjusted 3D detection result, including:

[0017] Set an initial iteration step size step_d corresponding to the depth value depth, an iteration termination step size step_d_end, an initial iteration step size step_r corresponding to the orientation angle rot, an iteration termination step size step_r_end, and set an iteration step size attenuation coefficient η; L is the loss value determined by the 2D bounding box and the initial projected 2D bounding box;

[0018] 1) If step_d > step_d_end, let depth_neg = depth - step_d, depth_pos = depth + step_d, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos between the projected 2D bounding box and the 2D bounding box;

[0019] 2) If L_neg <= L and L_pos < L, let step_d = step_d * η, jump to 4), otherwise go to 3);

[0020] 3) If L_pos > L and L_pos > L_neg, let depth = depth + step_d, L = L_pos; otherwise let depth = depth - step_d, L = L_neg;

[0021] 4) If step_r > step_r_end, let rot_neg = rot - step_r, rot_pos = rot + step_r, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos between the projected 2D bounding box and the 2D bounding box;

[0022] 5) If L_neg <= L and L_pos < L, let step_r = step_r * η, jump to 1), otherwise go to 6);

[0023] 6) If L_pos > L and L_pos > L_neg, let rot = rot + step_r, L = L_pos; otherwise let rot = rot - step_r, L = L_neg;

[0024] Among them, depth_neg and depth_pos are the adjustment values ​​of the depth value in the two adjustment directions, rot_neg and rot_pos are the adjustment values ​​of the orientation angle in the two adjustment directions; L_neg and L_pos are the loss values ​​between the 2D box corresponding to the adjustment values ​​of the depth value or the orientation value in the two adjustment directions and the projected 2D box.

[0025] In one embodiment, projecting the 3D detection result onto the image to be detected to obtain a projected 2D detection result includes:

[0026] The internal parameters of the image acquisition device are obtained, and the 3D detection result is projected onto the image to be detected according to the internal parameters to obtain the projected 2D detection result.

[0027] In one embodiment, projecting the 3D detection result onto a to-be-detected screen according to the internal reference to obtain a projected 2D detection result includes:

[0028] A 3D frame is obtained according to the 3D detection result, and the coordinates of each coordinate point after projection are calculated using a projection formula according to the coordinates and internal parameters of the coordinate points of the 3D frame, and the coordinates of each coordinate point after projection form a projected 2D detection result;

[0029] The projection formula includes:

[0030]

[0031] Among them, the internal parameter is (f x , f y , p x , p y ), the coordinates of the coordinate points of the 3D box are (X, Y, Z), and the coordinates after projection are (u, v).

[0032] In a second aspect, the present invention provides a monocular 3D detection device, comprising:

[0033] The recognition module is used to obtain the 2D detection result and 3D detection result of the target in the image to be detected;

[0034] A projection module, used for projecting the 3D detection result onto the image to be detected to obtain a projected 2D detection result;

[0035] An adjustment module, configured to iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result;

[0036] The processing module is used to perform 2D frame annotation and 3D frame annotation on the target on the to-be-detected screen according to the 2D detection result and the adjusted 3D detection result.

[0037] In a third aspect, the present invention provides an electronic device, comprising a memory and a memory storing a computer program, wherein when the processor executes the program, the steps of the monocular 3D detection method described in the first aspect are implemented.

[0038] In a fourth aspect, the present invention provides a processor-readable storage medium, wherein the processor-readable storage medium stores a computer program, and the computer program is used to enable the processor to execute the steps of the monocular 3D detection method described in the first aspect.

[0039] The monocular 3D detection method, system, electronic device and storage medium provided by the present invention project the 3D detection result onto the screen to be detected to obtain the projected 2D detection result, and iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result, thereby performing 2D frame annotation and 3D frame annotation on the target on the screen to be detected according to the 2D detection result and the adjusted 3D detection result, thereby achieving effective optimization of the output 3D detection result, while improving the rationality and stability of the 3D detection result, and improving the overall accuracy of monocular 3D detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0041] Figure 1 It is a schematic diagram of the process of the monocular 3D detection method provided by the present invention;

[0042] Figure 2 It is a schematic diagram of a frame marking of a target in a picture to be detected provided by the present invention;

[0043] Figure 3 It is a structural schematic diagram of a monocular 3D detection device provided by the present invention;

[0044] Figure 4 is a schematic diagram of the structure of an electronic device provided by the present invention; DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] Combine the following Figure 1-Figure 4 The present invention describes a monocular 3D detection method, system, electronic device and storage medium.

[0047] Figure 1 A schematic diagram of a monocular 3D detection method of the present invention is shown in FIG. Figure 1 , the method comprising:

[0048] 11. Obtain 2D detection results and 3D detection results of the target in the image to be detected;

[0049] 12. Project the 3D detection results onto the screen to be detected to obtain the projected 2D detection results;

[0050] 13. Iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result;

[0051] 14. Perform 2D and 3D box annotations on the target on the screen to be detected based on the 2D detection results and the adjusted 3D detection results.

[0052] Regarding steps 11 to 14, it should be noted that, in the present invention, the trained screen object detection model is used to detect the object in the screen to be detected, and a 2D detection result and a 3D detection result are obtained.

[0053] The screen target detection model is a model that uses the screen features of the target and the detection results corresponding to the screen features as input, obtained through machine learning training, and is used to locate the target in the video screen captured by a monocular camera.

[0054] In the present invention, the target in the video picture can be marked in the form of a frame (2D frame or 3D frame) according to the 2D detection result and the 3D detection result, so as to realize the positioning of the target in the video picture.

[0055] However, there are cases where the 3D detection results are not accurate enough, so the 3D detection results need to be adjusted.

[0056] In the present invention, the 3D detection result needs to be projected onto the screen to be detected to obtain the projected 2D detection result. To this end, the projected 2D detection result can also obtain a frame (2D frame) on the screen.

[0057] Then, the 3D detection result is iteratively adjusted again with the help of the constraints between the 2D detection result and the projected 2D detection result. Each adjusted 3D detection result regenerates the corresponding projected 2D detection result, and then the constraint processing continues. After the iterative algorithm ends, the adjusted 3D detection result can be obtained.

[0058] Finally, according to the 2D detection results and the adjusted 3D detection results, the target on the screen to be detected is annotated with 2D boxes and 3D boxes, so as to achieve accurate annotation of the target in the screen.

[0059] The monocular 3D detection method provided by the present invention projects the 3D detection result onto the screen to be detected to obtain the projected 2D detection result, and iteratively adjusts the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result, thereby performing 2D frame annotation and 3D frame annotation on the target on the screen to be detected according to the 2D detection result and the adjusted 3D detection result, thereby achieving effective optimization of the output 3D detection result, while improving the rationality and stability of the 3D detection result, and improving the overall accuracy of monocular 3D detection.

[0060] In the further description of the above method, the processing process of iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result is mainly explained as follows:

[0061] Determine loss information based on the 2D detection result and the projected 2D detection result;

[0062] The target parameters in the 3D detection result are adjusted according to the loss information and the preset iteration step size and termination step size to obtain the adjusted 3D detection result.

[0063] In this regard, it should be noted that in the present invention, the loss information is determined based on the 2D detection results and the 2D detection results after projection, that is, the loss value of the 2D frame corresponding to the 2D detection result and the projected 2D frame corresponding to the 2D detection result after projection, and the loss value can characterize the degree of fit between the 2D frame and the projected 2D frame.

[0064] When the loss value of the 2D frame corresponding to the 2D detection result and the projected 2D frame corresponding to the projected 2D detection result is large, it means that the 2D frame and the projected 2D frame are not consistent enough. At this time, the target parameters in the 3D detection result need to be adjusted. In the present invention, the parameters are adjusted by iterative adjustment, and the loss value is compared in each iteration.

[0065] During the iteration process, the iteration step and the termination step are configured. After the iteration is completed, the adjusted 3D detection result can be determined. According to the 3D detection result, the target can be marked with a 3D frame on the screen.

[0066] See also Figure 2 , 21 is the target on the screen to be detected, 22 is the 2D frame of the target, 23 is the 3D frame of the target, and 24 is the projected 2D frame of the target.

[0067] In the present invention, the 3D detection result requires a center point (Amodal Cneter offset), a depth value (ObjectDepth), an orientation angle (Object Orientation) and an object size (Object Dimension), which are all parameters that can be further optimized. However, based on the actual performance, Amodal Cneter offset and Object Dimension are generally predicted to be more accurate, so no additional optimization is required. Object Depth and Object Orientation mainly need to be optimized. For this reason, the target parameters in the 3D detection result include a depth value and an orientation angle.

[0068] A further method of the present invention adjusts and limits the target parameters by configuring the iteration step size and the termination step size, thereby ensuring a gradual adjustment and making the adjusted 3D detection result more reasonable and stable.

[0069] In the further description of the above method, the processing of determining the loss information based on the 2D detection result and the projected 2D detection result is mainly explained as follows:

[0070] Determine a 2D frame and a projected 2D frame according to the 2D detection result and the projected 2D detection result;

[0071] The loss value is determined based on the coordinate points of the 2D box and the coordinate points of the projected 2D box.

[0072] In this regard, it should be noted that in the present invention, the 2D frame and the projected 2D frame are determined based on the 2D detection result and the projected 2D detection result. Assume that the coordinates of the 2D frame are (x1, y1, x2, y2), and the coordinates of the projected 2D frame are (x1', y1', x2', y2'), and calculate the loss value L = -(abs(x1-x1')+abs(y1-y1')+abs(x2-x2')+abs(y2-y2')). Where abs is the absolute value.

[0073] In the further description of the above method, it mainly explains the process of adjusting the target parameters in the 3D detection result according to the loss information, the preset iteration step size, and the termination step size to obtain the adjusted 3D detection result, which is specifically as follows:

[0074] Set the initial iteration step size step_d corresponding to the depth value depth (such as 10m), the iteration termination step size step_d_end (such as 0.1m), the initial iteration step size step_r corresponding to the orientation angle rot (such as 0.3*π), the iteration termination step size step_r_end (such as 0.01 radians), and set the iteration step size attenuation coefficient η (such as 0.5); L is the loss value determined by the 2D box and the initial projected 2D box.

[0075] 1) If step_d > step_d_end, let depth_neg = depth - step_d, depth_pos = depth + step_d, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos of the projected 2D box and the 2D box.

[0076] 2) If L_neg <= L and L_pos < L, let step_d = step_d * η, jump to 4), otherwise go to 3).

[0077] 3) If L_pos > L and L_pos > L_neg, let depth = depth + step_d, L = L_pos; otherwise let depth = depth - step_d, L = L_neg.

[0078] 4) If step_r > step_r_end, let rot_neg = rot - step_r, rot_pos = rot + step_r, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos of the projected 2D box and the 2D box.

[0079] 5) If L_neg <= L and L_pos < L, let step_r = step_r * η, jump to 1), otherwise go to 6).

[0080] 6) If L_pos > L and L_pos > L_neg, let rot = rot + step_r, L = L_pos; otherwise let rot = rot - step_r, L = L_neg.

[0081] In this regard, it should be noted that depth_neg and depth_pos are respectively the adjustment values ​​of the depth value in two adjustment directions, rot_neg and rot_pos are respectively the adjustment values ​​of the orientation angle in two adjustment directions. L_neg and L_pos are respectively the loss values ​​between the 2D frame corresponding to the adjustment values ​​of the depth value or the orientation value in two adjustment directions and the projected 2D frame.

[0082] A further method of the present invention adjusts and limits the target parameters by configuring the iteration step size and the termination step size, thereby ensuring a gradual adjustment and making the adjusted 3D detection result more reasonable and stable.

[0083] In the further description of the above method, the 3D detection result is mainly projected onto the screen to be detected to obtain the projected 2D detection result, including:

[0084] The internal parameters of the image acquisition device are obtained, and the 3D detection results are projected onto the image to be detected according to the internal parameters to obtain the projected 2D detection results.

[0085] In this regard, it should be noted that the internal parameters of the image acquisition device include the length of the focal length in the x-axis direction, the length of the focal length in the y-axis direction, and the coordinates of the midpoint.

[0086] In the present invention, the 3D detection result is projected onto the screen to be detected according to the internal reference to obtain the projected 2D detection result, including:

[0087] A 3D frame is obtained according to the 3D detection result, and the coordinates of each coordinate point after projection are calculated using a projection formula according to the coordinates and internal parameters of the coordinate points of the 3D frame, and the coordinates of each coordinate point after projection form a projected 2D detection result;

[0088] The projection formula includes:

[0089]

[0090] Among them, the internal parameter is (f x , f y , p x , p y ), the coordinates of the coordinate points of the 3D box are (X, Y, Z), and the coordinates after projection are (u, v).

[0091] The monocular 3D detection device provided by the present invention is described below. The monocular 3D detection device described below and the monocular 3D detection method described above can be referred to each other.

[0092] Figure 3 A schematic diagram of the structure of a monocular 3D detection device provided by the present invention is shown in FIG. Figure 3The device comprises a recognition module 31, a projection module 32, an adjustment module 33 and a processing module 34, wherein:

[0093] The recognition module 31 is used to obtain the 2D detection result and the 3D detection result of the target in the image to be detected;

[0094] The projection module 32 is used to project the 3D detection result onto the image to be detected to obtain the projected 2D detection result;

[0095] An adjustment module 33, configured to iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result;

[0096] The processing module 34 is used to perform 2D frame marking and 3D frame marking on the target on the to-be-detected screen according to the 2D detection result and the adjusted 3D detection result.

[0097] In a further description of the above device, the adjustment module is specifically used for:

[0098] Determine loss information based on the 2D detection result and the projected 2D detection result;

[0099] The target parameters in the 3D detection result are adjusted according to the loss information and the preset iteration step size and termination step size to obtain the adjusted 3D detection result.

[0100] In a further description of the above device, the adjustment module is specifically used in the process of determining the loss information according to the 2D detection result and the projected 2D detection result:

[0101] Determine a 2D frame and a projected 2D frame according to the 2D detection result and the projected 2D detection result;

[0102] The loss value is determined based on the coordinate points of the 2D box and the coordinate points of the projected 2D box.

[0103] In a further description of the above device, the target parameters in the 3D detection result include a depth value and an orientation angle. Accordingly, the adjustment module adjusts the target parameters in the 3D detection result according to the loss information and the preset iteration step size and the termination step size to obtain the adjusted 3D detection result. Specifically used in the process:

[0104] Set the initial iteration step size step_d corresponding to the depth value depth, the iteration termination step size step_d_end, the initial iteration step size step_r corresponding to the orientation angle rot, the iteration termination step size step_r_end, and set the iteration step size attenuation coefficient η; L is the loss value determined by the 2D box and the initial projected 2D box;

[0105] 1) If step_d > step_d_end, let depth_neg = depth - step_d, depth_pos = depth + step_d, recalculate the projection of the 3D detection result, and calculate the loss values L_neg and L_pos between the projected 2D box and the 2D box;

[0106] 2) If L_neg <= L and L_pos < L, let step_d = step_d * η, jump to 4), otherwise go to 3);

[0107] 3) If L_pos > L and L_pos > L_neg, let depth = depth + step_d, L = L_pos; otherwise let depth = depth - step_d, L = L_neg;

[0108] 4) If step_r > step_r_end, let rot_neg = rot - step_r, rot_pos = rot + step_r, recalculate the projection of the 3D detection result, and calculate the loss values L_neg and L_pos between the projected 2D box and the 2D box;

[0109] 5) If L_neg <= L and L_pos < L, let step_r = step_r * η, jump to 1), otherwise go to 6);

[0110] 6) If L_pos > L and L_pos > L_neg, let rot = rot + step_r, L = L_pos; otherwise let rot = rot - step_r, L = L_neg;

[0111] Among them, depth_neg and depth_pos are the adjustment values of the depth value in two adjustment directions respectively, rot_neg and rot_pos are the adjustment values of the orientation angle in two adjustment directions respectively; L_neg and L_pos are the loss values between the 2D box and the projected 2D box corresponding to the adjustment values of the depth value or the orientation value in two adjustment directions respectively.

[0112] In the further description of the above device, the projection module is specifically used for:

[0113] Obtain the internal parameters of the image acquisition device, project the 3D detection result onto the image to be detected according to the internal parameters, and obtain the projected 2D detection result.

[0114] In the further description of the above device, during the process of projecting the 3D detection result onto the image to be detected according to the internal parameters and obtaining the projected 2D detection result, it is specifically used for:

[0115] The 3D frame is obtained according to the 3D detection result. The coordinates of each coordinate point after projection are calculated using the projection formula according to the coordinates and internal parameters of the coordinate points of the 3D frame. The coordinates of each coordinate point after projection form the 2D detection result after projection.

[0116] The projection formula includes:

[0117]

[0118] Among them, the internal parameter is (f x , f y , p x , p y ), the coordinates of the coordinate points of the 3D box are (X, Y, Z), and the coordinates after projection are (u, v).

[0119] Since the principle of the device described in the embodiment of the present invention is the same as that of the method described in the above embodiment, a more detailed explanation will not be repeated here.

[0120] It should be noted that, in the embodiment of the present invention, the relevant functional modules can be implemented by a hardware processor.

[0121] The monocular 3D detection device provided by the present invention projects the 3D detection result onto the screen to be detected to obtain the projected 2D detection result, and iteratively adjusts the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result, thereby performing 2D frame annotation and 3D frame annotation on the target on the screen to be detected according to the 2D detection result and the adjusted 3D detection result, thereby achieving effective optimization of the output 3D detection result, while improving the rationality and stability of the 3D detection result, and improving the overall accuracy of monocular 3D detection.

[0122] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown, the electronic device may include: a processor 41, a communication interface 42, a memory 43 and a communication bus 44, wherein the processor 41, the communication interface 42 and the memory 43 communicate with each other through the communication bus 44. The processor 41 may call a computer program in the memory 43 to execute the steps of the monocular 3D detection method, for example, including: obtaining a 2D detection result and a 3D detection result of a target in a to-be-detected screen; projecting the 3D detection result onto the to-be-detected screen to obtain a projected 2D detection result; iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result; and performing 2D frame annotation and 3D frame annotation on the target on the to-be-detected screen according to the 2D detection result and the adjusted 3D detection result.

[0123] In addition, the logic instructions in the above-mentioned memory 43 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0124] On the other hand, the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the monocular 3D detection method provided by the above methods, and the method includes: obtaining a 2D detection result and a 3D detection result of a target in a picture to be detected; projecting the 3D detection result onto the picture to be detected to obtain a projected 2D detection result; iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result; and performing 2D box annotation and 3D box annotation on the target on the picture to be detected according to the 2D detection result and the adjusted 3D detection result.

[0125] On the other hand, an embodiment of the present application also provides a processor-readable storage medium, which stores a computer program, and the computer program is used to enable the processor to execute the monocular 3D detection method provided by the above-mentioned embodiments, for example, including: obtaining 2D detection results and 3D detection results of the target in the picture to be detected; projecting the 3D detection results onto the picture to be detected to obtain the projected 2D detection results; iteratively adjusting the 3D detection results according to the 2D detection results and the projected 2D detection results to obtain the adjusted 3D detection results; and performing 2D box annotation and 3D box annotation on the target on the picture to be detected according to the 2D detection results and the adjusted 3D detection results.

[0126] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.

[0127] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0128] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A monocular 3D detection method, characterized in that: Including: Obtain the 2D detection result and 3D detection result of the target in the to-be-detected image; Project the 3D detection result onto the to-be-detected image to obtain the projected 2D detection result; Iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result; Perform 2D bounding box annotation and 3D bounding box annotation on the target in the to-be-detected image according to the 2D detection result and the adjusted 3D detection result; The iteratively adjusting the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain the adjusted 3D detection result includes: Determine the loss information according to the 2D detection result and the projected 2D detection result; the loss information can characterize the degree of coincidence between the 2D detection result and the projected 2D detection result; Adjust the target parameters in the 3D detection result according to the loss information, the preset iteration step size and termination step size to obtain the adjusted 3D detection result; The target parameters in the 3D detection result include the depth value and the orientation angle. Correspondingly, adjusting the target parameters in the 3D detection result according to the loss information, the preset iteration step size and termination step size to obtain the adjusted 3D detection result includes: Set the initial iteration step size step_d corresponding to the depth value depth, the iteration termination step size step_d_end, the initial iteration step size step_r corresponding to the orientation angle rot, the iteration termination step size step_r_end, and set the iteration step size attenuation coefficient η; L is the loss value determined by the 2D bounding box and the initial projected 2D bounding box; 1) If step_d > step_d_end, let depth_neg = depth – step_d, depth_pos = depth + step_d, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos of the projected 2D bounding box and the 2D bounding box; 2) If L_neg <= L and L_pos < L, let step_d = step_d * η, jump to 4), otherwise go to 3); 3) If L_pos > L and L_pos > L_neg, let depth = depth + step_d, L = L_pos; otherwise let depth = depth - step_d, L = L_neg; 4) If step_r > step_r_end, let rot_neg = rot – step_r, rot_pos = rot + step_r, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos of the projected 2D bounding box and the 2D bounding box; 5) If L_neg <= L and L_pos < L, let step_r = step_r * η, jump to 1), otherwise go to 6); 6) If L_pos > L and L_pos > L_neg, let rot = rot + step_r, L = L_pos; otherwise let rot = rot - step_r, L = L_neg; Among them, depth_neg and depth_pos are the adjustment values ​​of the depth value in the two adjustment directions, rot_neg and rot_pos are the adjustment values ​​of the orientation angle in the two adjustment directions; L_neg and L_pos are the loss values ​​between the 2D box corresponding to the adjustment values ​​of the depth value or the orientation value in the two adjustment directions and the projected 2D box.

2. The monocular 3D detection method according to claim 1, characterized in that: Determining loss information according to the 2D detection result and the projected 2D detection result includes: Determine a 2D frame and a projected 2D frame according to the 2D detection result and the projected 2D detection result; A loss value is determined according to the coordinate points of the 2D box and the coordinate points of the projected 2D box.

3. The monocular 3D detection method according to claim 1, characterized in that: The step of projecting the 3D detection result onto the image to be detected to obtain a projected 2D detection result includes: The internal parameters of the image acquisition device are obtained, and the 3D detection result is projected onto the image to be detected according to the internal parameters to obtain the projected 2D detection result.

4. The monocular 3D detection method according to claim 3, characterized in that: The step of projecting the 3D detection result onto a screen to be detected according to the internal reference to obtain a projected 2D detection result includes: A 3D frame is obtained according to the 3D detection result, and the coordinates of each coordinate point after projection are calculated using a projection formula according to the coordinates and internal parameters of the coordinate points of the 3D frame, and the coordinates of each coordinate point after projection form a projected 2D detection result; The projection formula includes: Among them, the internal parameter is (f x , f y , p x , p y ), the coordinates of the coordinate points of the 3D box are (X, Y, Z), and the coordinates after projection are (u, v).

5. A monocular 3D detection device, characterized in that: include: The recognition module is used to obtain the 2D detection result and 3D detection result of the target in the image to be detected; A projection module, used for projecting the 3D detection result onto the image to be detected to obtain a projected 2D detection result; An adjustment module, configured to iteratively adjust the 3D detection result according to the 2D detection result and the projected 2D detection result to obtain an adjusted 3D detection result; A processing module, used for performing 2D frame annotation and 3D frame annotation on the target on the to-be-detected screen according to the 2D detection result and the adjusted 3D detection result; The adjustment module is specifically used for: Determining loss information according to the 2D detection result and the projected 2D detection result; the loss information can characterize the degree of consistency between the 2D detection result and the projected 2D detection result; Adjusting target parameters in the 3D detection result according to the loss information and the preset iteration step size and termination step size to obtain an adjusted 3D detection result; The target parameters in the 3D detection result include a depth value and an orientation angle. Accordingly, the target parameters in the 3D detection result are adjusted according to the loss information and the preset iteration step size and the termination step size to obtain the adjusted 3D detection result, including: Set the initial iteration step size step_d corresponding to the depth value depth, the iteration termination step size step_d_end, the initial iteration step size step_r corresponding to the orientation angle rot, the iteration termination step size step_r_end, and set the iteration step size attenuation coefficient η; L is the loss value determined by the 2D box and the initial projected 2D box; 1) If step_d > step_d_end, set depth_neg = depth – step_d, depth_pos = depth + step_d, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos between the projected 2D box and the 2D box; 2) If L_neg <= L and L_pos < L, set step_d = step_d * η, jump to 4); otherwise, go to 3); 3) If L_pos > L and L_pos > L_neg, set depth = depth + step_d, L = L_pos; otherwise, set depth = depth - step_d, L = L_neg; 4) If step_r > step_r_end, set rot_neg = rot – step_r, rot_pos = rot + step_r, recalculate the 3D detection result projection, and calculate the loss values L_neg and L_pos between the projected 2D box and the 2D box; 5) If L_neg <= L and L_pos < L, set step_r = step_r * η, jump to 1); otherwise, go to 6); 6) If L_pos > L and L_pos > L_neg, set rot = rot + step_r, L = L_pos; otherwise, set rot = rot - step_r, L = L_neg; Wherein, depth_neg and depth_pos are the adjustment values of the depth value in two adjustment directions respectively, rot_neg and rot_pos are the adjustment values of the orientation angle in two adjustment directions respectively; L_neg and L_pos are the loss values between the 2D box and the projected 2D box corresponding to the adjustment values of the depth value or the orientation value in two adjustment directions respectively.

6. An electronic device comprising a processor and a memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the monocular 3D detection method according to any one of claims 1 to 4.

7. A processor-readable storage medium, characterized in that: The processor-readable storage medium stores a computer program, and the computer program is used to cause the processor to execute the steps of the monocular 3D detection method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Three-dimensional pose determination method and device, electronic equipment and storage medium

    CN112767489A

  • Three-dimensional reconstruction and related interaction and measurement methods and related devices and equipment

    CN112767538A

  • Method for performing 3D target detection on monocular RGB image

    CN113128434A